Every instrument ever built for medical AI aims at that 12.5% — the answer at the moment of the note. Here they are, era by era, and here is the same chart recolored to ask who each one is for and who gets to see what it produces.
Multiple-choice recall — integrated retrieval, not practice. Saturates at 84–90%, at or above physician level; falls to 45–69% on clinical tasks. Ends as marketing.
Physician-written rubrics over multi-turn conversations — a curated vignette stands in for the patient.
Real EHR data and agentic workflows — graded on output quality, never on adoption.
Real use and real outcomes, one era: benchmarks curated from live telemetry, while the outcome question runs daily as private telemetry.