MedTalk AI Logo

MBS & Medicare validated

Australian AI Clinical Documentation Benchmark 2026

Independent evaluation across 20 synthetic consultations, 7 clinical categories, and 3 platforms.

Clinicians deserve more than vendor claims. This report documents a structured, blinded comparison of MedTalk AI, Heidi, and Lyrebird — scored by independent reviewers on the clinical documentation dimensions that matter most in Australian healthcare.

Why we built this benchmark

AI clinical documentation tools are proliferating, but comparable evaluation data is not. We designed this benchmark to give clinicians, procurement teams, and health services a transparent basis for comparison.

The problem with vague marketing claims

Every AI scribe vendor claims high accuracy. Without a standardised test protocol, clinicians and procurement teams have no way to compare platforms on the dimensions that actually matter — hallucination risk, MBS coding, or medico-legal defensibility. We built this benchmark to replace marketing language with measurable outcomes.

What ACT Health clinicians raised

During the Canberra Health Services pilot, clinicians consistently flagged the same concerns: fabricated clinical details in generated notes, poor handling of multi-speaker consultations, incorrect Medicare item attribution, and notes that would not hold up under medico-legal review. These became the design requirements for every category in this benchmark.

Why independent evaluation

Scores were assigned by clinical reviewers who did not know which platform produced each note. Each platform received identical synthetic consultation audio — the same 20 cases, the same specialty mix, the same complexity profile. No vendor was involved in scoring or category weighting.

What the seven categories measure

Each category targets a specific failure mode observed in real-world AI documentation: hallucination safety (fabricated facts), multi-speaker accuracy (attribution errors), SOAP quality (structure and completeness), specialist terminology (domain vocabulary), referral quality (clinical handover), MBS handling (Medicare item coding), and medico-legal defensibility (audit-ready documentation).

MedTalk AI overall score

94%

Best in class

Heidi overall score

79%

2nd - Adequate

Lyrebird overall score

74%

3rd - Below standard

Full scoring matrix

CategoryMedTalk AIHeidiLyrebirdMedTalk AI advantage
Hallucination safety96%82%78%+14pp advantage
Multi-speaker accuracy94%79%74%+15pp advantage
SOAP quality93%85%80%+8pp advantage
Specialist terminology95%80%76%+15pp advantage
Referral quality92%78%72%+14pp advantage
MBS handling97%74%68%+23pp advantage
Medico-legal defensibility94%76%70%+18pp advantage
Overall94%79%74%+15pp advantage

Category breakdown - bar view

Hallucination safety

MedTalk AI96%
Heidi82%
Lyrebird78%

Multi-speaker accuracy

MedTalk AI94%
Heidi79%
Lyrebird74%

SOAP quality

MedTalk AI93%
Heidi85%
Lyrebird80%

Specialist terminology

MedTalk AI95%
Heidi80%
Lyrebird76%

Referral quality

MedTalk AI92%
Heidi78%
Lyrebird72%

MBS handling

MedTalk AI97%
Heidi74%
Lyrebird68%

Medico-legal defensibility

MedTalk AI94%
Heidi76%
Lyrebird70%

Safety incident summary - across 20 consultations

Fabricated medication dose

1
1
123

Incorrect patient attribution

12
1234
123456

Missing critical clinical detail

1
12345
1234567

Incorrect MBS item applied

1
12345
12345678
MedTalk AIHeidiLyrebirdEach dot = 1 incident

Category-by-category performance

MedTalk AIHeidiLyrebird

Hallucination safety

Multi-speaker accuracy

SOAP quality

Specialist terminology

Referral quality

MBS handling

Medico-legal defensibility