What We Benchmark Instead

What we benchmark instead

Every instrument ever built for medical AI aims at that 12.5% — the answer at the moment of the note. Here they are, era by era, and here is the same chart recolored to ask who each one is for and who gets to see what it produces.

Four eras of measuring medical AI

2019–24Exam era“Are you booksmart?”

Multiple-choice recall — integrated retrieval, not practice. Saturates at 84–90%, at or above physician level; falls to 45–69% on clinical tasks. Ends as marketing.

2023–Rubric era“Can it converse safely?”

Physician-written rubrics over multi-turn conversations — a curated vignette stands in for the patient.

2024–Task era“Can it do clinical work?”

Real EHR data and agentic workflows — graded on output quality, never on adoption.

2025–Deployment era“What does deployment show — and did it change the patient?”

Real use and real outcomes, one era: benchmarks curated from live telemetry, while the outcome question runs daily as private telemetry.

patient-chat long tail, not charted: K Health · Ada · Babylon · HealthTap · Woebot · symptom checkers & wellness bots
JAN 2026
The instruments, as they arrived
January 2026 — where the exam era actually endsPatients had been asking for three years — 230 million health questions a week — before any instrument measured it. Then the labs stopped selling scores and built for that instead, in a single week: ChatGPT Health on the 7th, Claude for Healthcare on the 12th. The question stops being can it pass? and becomes what is it doing?

Two questions, one chart