HealthBench Professional

Compare models

updated August 16, 2026

The pairings that decide real deployments get full head-to-head pages: score gap, list pricing, and the cost of a concrete consult workload on each model. Every other pairing is in the score-difference matrix below.

Head-to-head pages

Score-difference matrix

Row score minus column score, so a positive number means the row model scores higher. Columns are numbered by rank.

model123456789
1. Claude Fable 5·+0.055+0.062+0.082+0.083+0.102+0.103+0.276+0.310
2. GPT-5.6 Sol-0.055·+0.007+0.027+0.028+0.047+0.048+0.221+0.255
3. Claude Opus 5-0.062-0.007·+0.020+0.021+0.040+0.041+0.214+0.248
4. Claude Sonnet 5-0.082-0.027-0.020·+0.001+0.020+0.021+0.194+0.228
5. GPT-5.6 Terra-0.083-0.028-0.021-0.001·+0.019+0.020+0.193+0.227
6. Claude Opus 4.8-0.102-0.047-0.040-0.020-0.019·+0.001+0.174+0.208
7. GPT-5.6 Luna-0.103-0.048-0.041-0.021-0.020-0.001·+0.173+0.207
8. GPT-5.5 Instant-0.276-0.221-0.214-0.194-0.193-0.174-0.173·+0.034
9. MAI-Thinking-1-0.310-0.255-0.248-0.228-0.227-0.208-0.207-0.034·

Differences of 0.01 or less on a 525-task set are within plausible re-run variation and should be read as ties.

Per-model detail, including pricing and context windows, is on the models page.