SYNTHETIC DATA — DEMONSTRATION ONLY. Every transcript, verdict, and number on this page is generated. No real calls, consumers, agents, clients, or vendors.

Judge-the-Judge Compliance Audit

Contact-center compliance QA with a calibrated judge — seeded synthetic run f145bd172fee7d22, 500 calls, 120-call gold audit.

The problem this demo shows: automated QA vendors score 100% of calls, but the scoring judge is itself unaudited. Its flag rate is not the violation rate. Below, the same synthetic corpus is scored by an intentionally imperfect judge: the vendor-dashboard number (gray) and the truth (green dashed) disagree on every rule. The calibrated estimate (blue) recovers the truth with an honest interval, using only a small gold-review sample — the same method as our PII prevalence audits.

Calls scored
500
100% coverage (synthetic)
Gold-audited
120
judge vs. gold labels
Rules checked
4
debt-collection pack
Largest vendor gap
5.2 pp
Right-party contact verification

Violation-rate estimates, per rule

Vendor judge flag rate (uncorrected) Calibrated estimate ± 95% CI Injected true rate (known because synthetic)
Mini-Miranda disclosure mini_miranda
7.6% vendor 11.4% corrected true 8.0%
Judge on gold sample — sensitivity 66.7% (35.4%–87.9%), specificity 100.0% (96.7%–100.0%) · misses 3 · false flags 0
Right-party contact verification right_party
15.2% vendor 10.0% corrected true 12.0%
Judge on gold sample — sensitivity 82.4% (59.0%–93.8%), specificity 92.2% (85.4%–96.0%) · misses 3 · false flags 8
Prohibited language prohibited_language
7.8% vendor 8.0% corrected true 5.0%
Judge on gold sample — sensitivity 87.5% (52.9%–97.8%), specificity 99.1% (95.1%–99.8%) · misses 1 · false flags 1
Recording/company disclosure script disclosure_script
17.6% vendor 15.1% corrected true 15.0%
Judge on gold sample — sensitivity 69.2% (42.4%–87.3%), specificity 91.6% (84.8%–95.5%) · misses 4 · false flags 9

Judge error decomposition (gold sample)

RuleSens.Spec.Misses (FN)False flags (FP)Vendor rateCorrected95% CI
Mini-Miranda disclosure66.7%100.0%307.6%11.4%6.9%–22.2%
Right-party contact verification82.4%92.2%3815.2%10.0%2.0%–18.4%
Prohibited language87.5%99.1%117.8%8.0%4.3%–12.8%
Recording/company disclosure script69.2%91.6%4917.6%15.1%4.8%–30.6%

Calibration freshness

Run generated 2026-08-11 14:04 UTC · judge mock-judge/1.0 · corpus digest f145bd172fee7d22 · seed 7.

This calibration is valid for this judge version on this rule pack only. Any judge model change, prompt change, rule-pack edit, or drift in call mix invalidates the sensitivity/specificity estimates above and requires a fresh gold-sample audit — a drift sentinel should re-run this audit on a schedule and alarm when calibration moves.