AI clinical documentation tools are proliferating, but comparable evaluation data is not. We designed this benchmark to give clinicians, procurement teams, and health services a transparent basis for comparison.
The problem with vague marketing claims
Every AI scribe vendor claims high accuracy. Without a standardised test protocol, clinicians and procurement teams have no way to compare platforms on the dimensions that actually matter — hallucination risk, MBS coding, or medico-legal defensibility. We built this benchmark to replace marketing language with measurable outcomes.
What ACT Health clinicians raised
During the Canberra Health Services pilot, clinicians consistently flagged the same concerns: fabricated clinical details in generated notes, poor handling of multi-speaker consultations, incorrect Medicare item attribution, and notes that would not hold up under medico-legal review. These became the design requirements for every category in this benchmark.
Why independent evaluation
Scores were assigned by clinical reviewers who did not know which platform produced each note. Each platform received identical synthetic consultation audio — the same 20 cases, the same specialty mix, the same complexity profile. No vendor was involved in scoring or category weighting.
What the seven categories measure
Each category targets a specific failure mode observed in real-world AI documentation: hallucination safety (fabricated facts), multi-speaker accuracy (attribution errors), SOAP quality (structure and completeness), specialist terminology (domain vocabulary), referral quality (clinical handover), MBS handling (Medicare item coding), and medico-legal defensibility (audit-ready documentation).