One chart, then the argument
Contents
Healthcare AI benchmarks are saturated, and the evaluation that answers “does this help?” looks like product analytics — in private. One argument, in three acts and an appendix.
Part I

The Encounter →

Anatomy of a visit: the question stream, the scope, the same patient twice, and the question that never gets asked.
Part II

What We Benchmark Instead →

The four eras of measuring medical AI, and the one chart with two questions.
Part III

The Argument →

Medicine runs on tickets. The ticket nobody closes. No one grades the outcome.
Appendix

References & Prior Art →

Every card, sparkline, and source behind the chart.