Veridict

Improve

Calibration

Agreement between the engine and your QA lead, per criterion. This is the number that decides whether the programme is trustworthy.

Overall agreement
85.9%
target ≥ 80%
Criteria measured
12
Below threshold
5
guidance likely ambiguous
Judgements compared
377

5 criteria are below 80% agreement

Low agreement almost always means the guidance text is ambiguous, not that the model is wrong — two humans would disagree on these too. Fix the wording before changing anything else.

Edit guidance in Scorecards →
CriterionAgreementAI pass / human failAI fail / human pass
Was correct hold procedure followed?hold_procedure
57.1%
03guidance ambiguous
Did the agent demonstrate active listening?active_listening
70.6%
91guidance ambiguous
Were required disclosures read where applicable?compliance_disclosure
71.4%
33guidance ambiguous
Did the agent acknowledge the customer's situation appropriately?empathy
72%
70guidance ambiguous
Did the agent maintain a professional tone throughout?tone_professional
73.5%
81guidance ambiguous
Did the agent avoid collecting personal data beyond what the task required?data_minimisation
85.7%
31calibrated
Was the information the agent provided accurate and complete?accurate_info
87.5%
50calibrated
Was the issue resolved, or a clear next step and owner given?resolution
87.5%
50calibrated
Did the agent close correctly?closing
95%
20calibrated
Did the agent open with the approved greeting and identify themselves?greeting
97.5%
10calibrated
Did the agent request or accept NRIC digits as a means of verifying identity?nric_auth
97.5%
10calibrated
Was the caller's identity verified before account details were discussed?identity_verification
100%
00calibrated

Why the soft criteria score worst

Empathy, tone and active listening sit at the bottom here, and that is expected — they are judgement calls where two experienced reviewers also disagree. The objective criteria (NRIC use, greeting, verification) sit above 90% because they describe observable events. When a soft criterion falls below threshold, the fix is to rewrite its guidance in observable terms — “paraphrased the issue back” rather than “showed empathy”.