Claude Opus 4.8 on HealthBench Professional
rank 6 of 9 · updated August 16, 2026
Claude Opus 4.8 scores 0.558 on HealthBench Professional, rank 6 of 9 evaluated models. The last release of the Opus 4 series, superseded by Claude Opus 5 in July 2026. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 6 of 9 |
|---|---|
| score | 0.558 |
| lab | Anthropic |
| context window | 1.0M tokens |
| API price per 1M tokens | $5.00 in / $25.00 out |
| license | proprietary |
| released | 2026-05-28 |
Position in the field
The gap to the leader, Claude Fable 5 at 0.660, is 0.102. Directly above sits GPT-5.6 Terra at 0.577. Directly below sits GPT-5.6 Luna at 0.557. Scores on this page come from the same evaluation run, so differences between models are differences on identical tasks, not across configurations.
What does Claude Opus 4.8 score on HealthBench Professional?
Claude Opus 4.8 scores 0.558 on HealthBench Professional, which places it at rank 6 of 9 evaluated models as of August 16, 2026.
How much does Claude Opus 4.8 cost to run?
Claude Opus 4.8 is priced at $5.00 per million input tokens and $25.00 per million output tokens through Anthropic's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Claude Opus 4.8 vs Claude Fable 50.558 vs 0.660 · Claude Fable 5 by 0.102
- Claude Opus 4.8 vs GPT-5.6 Sol0.558 vs 0.605 · GPT-5.6 Sol by 0.047
- 0.558 vs 0.598 · Claude Opus 5 by 0.040
- Claude Opus 4.8 vs Claude Sonnet 50.558 vs 0.578 · Claude Sonnet 5 by 0.020
- Claude Opus 4.8 vs GPT-5.6 Terra0.558 vs 0.577 · GPT-5.6 Terra by 0.019
- Claude Opus 4.8 vs GPT-5.6 Luna0.558 vs 0.557 · Claude Opus 4.8 by 0.001
- Claude Opus 4.8 vs GPT-5.5 Instant0.558 vs 0.384 · Claude Opus 4.8 by 0.174
- Claude Opus 4.8 vs MAI-Thinking-10.558 vs 0.350 · Claude Opus 4.8 by 0.208
How tasks are selected and graded is on the methodology page. The full ranking is on the leaderboard.