Responsible AI
How HireFlow-AI uses AI, what its scores can and cannot tell you, and the steps we take to be fair and honest about it.
Last updated: 6 October 2026
1. How we use AI
- A large language model (Google's Gemini, used through its API; we have not trained or fine-tuned it) acts as the interviewer or examiner, writes questions, reads answers and writes feedback.
- Two small models we trained ourselves help with scoring. One estimates facial expression from your camera inside your browser. The other predicts the score the language model would give an answer, and is blended with the language model's own score.
- Other parts are ordinary software: timing, limits, keyword counts, the arithmetic of the Confidence Score, and the rules that stop an off-topic or empty answer from earning a high mark.
2. What the scores mean
Scores, feedback and expression estimates are generated by automated systems for practice. They can be wrong. They are not a measure of you as a person and they do not predict whether any employer or teacher will hire or pass you. Use them to decide what to practise next, not to judge yourself.
The Confidence Score combines speech clarity (30%), expression steadiness (25%), content relevance (25%) and eye contact (20%). Each part is explained in your report. When our two ways of scoring an answer disagree by a lot, the report says the number is less certain.
3. Fairness
- Your name, gender, age, nationality and photo are not inputs to any score. In the Viva Specialist the examiner does not see a student's name.
- The AI is told to judge what you say, not your accent, grammar or fluency, and answers are copy-edited for language before scoring so that non-native speakers are not marked down for their English. In our own tests this roughly halved the gap, but did not remove it.
- Off-topic, empty and brush-off answers are capped by fixed rules that the AI cannot override, so a lenient or harsh response cannot move them far.
- Without a camera or microphone you are never penalised: eye-contact and expression results simply do not appear.
4. Limits we know about
- The expression model was trained on still photographs and is only moderately accurate on a public test set. Facial expressions are not reliable evidence of how someone feels, so the label is coaching feedback only.
- The answer-scoring model learned from a language model's grades of synthetic answers, not from human interviewers. It measures agreement with that model, not with people.
- We have not yet tested how scoring behaves across different accents, skin tones, lighting, glasses or disabilities. Language models and image datasets can carry biases we have not found.
- Interviews are in English only, and speech recognition varies by browser and microphone.
5. People stay in charge
HireFlow-AI is a practice tool and is not used to make hiring decisions. In the Viva Specialist, scores are a starting point for the teacher: teachers can read every transcript, change any rubric score, and decide the final grade. Teachers and students should treat an automatic score as one input, never as the whole assessment.
6. Your data and the AI
We send our AI provider the text it needs to do its job, and we do not use your interviews to train our own models. The Privacy Policy lists exactly what is sent, and the Security page explains how we protect it.
Questions about this page?
Send us a message through the contact page and we will get back to you.
See also: Terms · Privacy · Security · Cookies · Responsible AI
