TODAY
Exam eramultiple-choice knowledge · saturated at 84–90%
Rubric eraphysician-written rubrics over multi-turn chats
Task erareal EHR data · agentic workflows
Deployment erareal use → real outcomes · measured privately
key eval timeline — the instruments, by steward

OpenAI

Google
industry
academia
HealthBench
len-adjusted
HB Pro
Med-PaLM
Med-PaLM LF
AMIE
MedHELM
MedAgentBench
ARISE·MAST
GPT-4 USMLE
MEDEC
OpenEvidence
PubMedQA
MedQA
MMLU-Med
MedMCQA
Abaluck et al.
unlabeled: Med-Gemini · AgentClinic · BiasMedQA · MedQA-CS · ClinicBench · MedXpertQA · LLMEval-Med · CSEDB
product launches — the usage the instruments chase, by brand

OpenAI

Anthropic

Google

Epic

Doximity
ChatGPT '22
GPT-4 '23
ChatGPT Health '26
Claude '23
Claude for Healthcare '26
MedGemma '25
Art · Emmie '25
Art '25
Emmie '25
Doximity GPT '23

OpenEvidence

Microsoft

Abridge
NEJM '25
Visits '25
EvidenceGrade
DAX Copilot '23
Abridge in Epic '23
Scribe '25
AI Suite '26
NOHARM v2