The problem this demo shows: automated QA vendors score 100% of calls, but the scoring judge is itself unaudited. Its flag rate is not the violation rate. Below, the same synthetic corpus is scored by an intentionally imperfect judge: the vendor-dashboard number (gray) and the truth (green dashed) disagree on every rule. The calibrated estimate (blue) recovers the truth with an honest interval, using only a small gold-review sample — the same method as our PII prevalence audits.
Judge-the-Judge Compliance Audit
Contact-center compliance QA with a calibrated judge — seeded synthetic run
f145bd172fee7d22, 500 calls, 120-call gold audit.
Calls scored
500
100% coverage (synthetic)
Gold-audited
120
judge vs. gold labels
Rules checked
4
debt-collection pack
Largest vendor gap
5.2 pp
Right-party contact verification
Violation-rate estimates, per rule
Vendor judge flag rate (uncorrected)
Calibrated estimate ± 95% CI
Injected true rate (known because synthetic)
Mini-Miranda disclosure
mini_miranda
Judge on gold sample — sensitivity 66.7%
(35.4%–87.9%),
specificity 100.0%
(96.7%–100.0%)
· misses 3 · false flags 0
Right-party contact verification
right_party
Judge on gold sample — sensitivity 82.4%
(59.0%–93.8%),
specificity 92.2%
(85.4%–96.0%)
· misses 3 · false flags 8
Prohibited language
prohibited_language
Judge on gold sample — sensitivity 87.5%
(52.9%–97.8%),
specificity 99.1%
(95.1%–99.8%)
· misses 1 · false flags 1
Recording/company disclosure script
disclosure_script
Judge on gold sample — sensitivity 69.2%
(42.4%–87.3%),
specificity 91.6%
(84.8%–95.5%)
· misses 4 · false flags 9
Judge error decomposition (gold sample)
| Rule | Sens. | Spec. | Misses (FN) | False flags (FP) | Vendor rate | Corrected | 95% CI |
|---|---|---|---|---|---|---|---|
| Mini-Miranda disclosure | 66.7% | 100.0% | 3 | 0 | 7.6% | 11.4% | 6.9%–22.2% |
| Right-party contact verification | 82.4% | 92.2% | 3 | 8 | 15.2% | 10.0% | 2.0%–18.4% |
| Prohibited language | 87.5% | 99.1% | 1 | 1 | 7.8% | 8.0% | 4.3%–12.8% |
| Recording/company disclosure script | 69.2% | 91.6% | 4 | 9 | 17.6% | 15.1% | 4.8%–30.6% |
Calibration freshness
Run generated 2026-08-11 14:04 UTC · judge mock-judge/1.0
· corpus digest f145bd172fee7d22 · seed 7.
This calibration is valid for this judge version on this rule pack only. Any judge model change, prompt change, rule-pack edit, or drift in call mix invalidates the sensitivity/specificity estimates above and requires a fresh gold-sample audit — a drift sentinel should re-run this audit on a schedule and alarm when calibration moves.