Improve
Calibration
Agreement between the engine and your QA lead, per criterion. This is the number that decides whether the programme is trustworthy.
5 criteria are below 80% agreement
Low agreement almost always means the guidance text is ambiguous, not that the model is wrong — two humans would disagree on these too. Fix the wording before changing anything else.
Edit guidance in Scorecards →| Criterion | Agreement | AI pass / human fail | AI fail / human pass | |
|---|---|---|---|---|
| Was correct hold procedure followed?hold_procedure | 57.1% | 0 | 3 | guidance ambiguous |
| Did the agent demonstrate active listening?active_listening | 70.6% | 9 | 1 | guidance ambiguous |
| Were required disclosures read where applicable?compliance_disclosure | 71.4% | 3 | 3 | guidance ambiguous |
| Did the agent acknowledge the customer's situation appropriately?empathy | 72% | 7 | 0 | guidance ambiguous |
| Did the agent maintain a professional tone throughout?tone_professional | 73.5% | 8 | 1 | guidance ambiguous |
| Did the agent avoid collecting personal data beyond what the task required?data_minimisation | 85.7% | 3 | 1 | calibrated |
| Was the information the agent provided accurate and complete?accurate_info | 87.5% | 5 | 0 | calibrated |
| Was the issue resolved, or a clear next step and owner given?resolution | 87.5% | 5 | 0 | calibrated |
| Did the agent close correctly?closing | 95% | 2 | 0 | calibrated |
| Did the agent open with the approved greeting and identify themselves?greeting | 97.5% | 1 | 0 | calibrated |
| Did the agent request or accept NRIC digits as a means of verifying identity?nric_auth | 97.5% | 1 | 0 | calibrated |
| Was the caller's identity verified before account details were discussed?identity_verification | 100% | 0 | 0 | calibrated |
Why the soft criteria score worst
Empathy, tone and active listening sit at the bottom here, and that is expected — they are judgement calls where two experienced reviewers also disagree. The objective criteria (NRIC use, greeting, verification) sit above 90% because they describe observable events. When a soft criterion falls below threshold, the fix is to rewrite its guidance in observable terms — “paraphrased the issue back” rather than “showed empathy”.