HealthBench Professional

Claude Opus 5 vs Claude Opus 4.8 on HealthBench Professional

updated August 16, 2026

Claude Opus 5 scores 0.598 to Claude Opus 4.8's 0.558, a gap of 0.040 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.

Side by side

Anthropic logoClaude Opus 5Anthropic logoClaude Opus 4.8
score0.5980.558
rank3 of 96 of 9
context window1.0M1.0M
price per 1M tokens, in / out$5.00 / $25.00$5.00 / $25.00
1,000 consult exchanges$27.50$27.50
released2026-07-242026-05-28
licenseproprietaryproprietary

Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.

Reading this pairing

A clean generational read, because the price did not move: Opus 5 scores 0.040 higher than Opus 4.8 at identical list rates. Anthropic shipped the upgrade 57 days after 4.8's release. For this workload there is no remaining reason to pin the older model; the only argument for 4.8 is an existing eval baseline you are not ready to re-run.

Which scores higher on HealthBench Professional, Claude Opus 5 or Claude Opus 4.8?

Claude Opus 5 scores higher: 0.598 against Claude Opus 4.8's 0.558, a difference of 0.040 on the 525-task set, as of August 16, 2026.

Which is cheaper to run, Claude Opus 5 or Claude Opus 4.8?

They cost the same at list rates: 1,000 typical consult exchanges (2,000 input and 700 output tokens each) run $27.50 on either model.

Related comparisons

Full results for both models: Claude Opus 5 and Claude Opus 4.8. The complete score-difference matrix is on the compare page.