AIA
Back to Blog

From Score-Based to Evidence-Based Interview Reports

Why a universal AI interview score creates false precision, and what we show instead.

July 27, 20263 min read
AI in HiringOriginal ResearchEvidence-Based HiringRecruitmentHuman-Led AI

A red apple and a green apple side by side

Is this apple tasty?

You can answer, but the number is meaningless until we agree on a scale. Sweetness? Acidity? Texture? The kind of apple you personally like?

Put two apples next to each other and the task becomes easier. You can say which is sweeter, firmer, or closer to what you want.

Recruitment has the same problem. There is no universal 8/10 candidate. Every company, role, and hiring team is looking for something different. An AI score hides that context behind a precise-looking number.

That is why we stopped using an overall interview score as the report.

From a single AI score to a criteria-based report with source evidence

The score was not a reliable scale

We tested GPT-5.4 on 600 synthetic interview transcripts across four roles, five interview types, three interview lengths, and ten controlled quality levels.

The model never scored an interview above 8.3. Candidates designed at levels 8, 9, and 10 collapsed into almost the same range. The decimal places looked precise, but they did not represent a stable universal scale.

Direct comparison worked better because it asked a narrower question: which candidate demonstrated more of the criteria required for this specific role?

We also tested the opposite extreme. We removed numbers from the internal pipeline and used only broad evidence labels. On 100 evaluations and 20 comparison groups, broad-tier agreement fell from 89% to 76%, while top-two comparison accuracy fell from 95% to 75%.

Benchmark results comparing the previous pipeline with the score-free evidence prototype

The lesson was not "all numbers are bad". Relative signals can still help compare closely matched candidates internally. The problem was presenting an AI score as if it were an objective hiring result.

What the report shows instead

The customer-facing report is now organized around the criteria defined for the role:

  • What requirement was assessed
  • Whether the interview confirmed it, partially confirmed it, contradicted it, or did not discuss it
  • How strong the evidence was
  • The required and demonstrated levels
  • The exact answer behind the conclusion

Criteria confirmation matrix from the sample interview analysis

Not discussed is not a negative score. It means the interview did not produce enough evidence. Partially confirmed means there is relevant evidence, but its depth or scope is limited.

Most importantly, every significant conclusion can lead back to the source passage. The recruiter can read the answer in context and decide whether the interpretation is justified.

A highlighted answer in the sample transcript linked to the report conclusion

Comparison without pretending there is a universal winner

Candidates are compared against the same role criteria and required levels. The report shows where each person is stronger, where evidence is missing, and what should be verified next.

Competency profile comparing both candidates against the same required levels

The order helps a team navigate the report. The evidence and trade-offs are the decision material.

Humans decide.

A narrower, more honest claim

Our tests used synthetic transcripts. They helped us improve the product, but they do not prove real-world job performance or hiring validity.

What we can say is simpler:

  • A universal AI interview score creates false precision.
  • Evidence tied to role criteria is easier to review.
  • Source passages make AI conclusions traceable.
  • The final decision belongs to the hiring team.

Open the sample reports to inspect the fictional candidate comparison, interview analysis, transcript, CV ranking, and downloadable PDFs.

Try AI Interview Analyzer

50 free credits. No credit card required.

Start Free
From Score-Based to Evidence-Based Interview Reports — AI Interview Analyzer