HealthBench Professional

Claude Fable 5 vs GPT-5.6 Sol on HealthBench Professional

updated August 16, 2026

Claude Fable 5 scores 0.660 to GPT-5.6 Sol's 0.605, a gap of 0.055 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.

Side by side

Anthropic logoClaude Fable 5OpenAI logoGPT-5.6 Sol
score0.6600.605
rank1 of 92 of 9
context window1.0M1.1M
price per 1M tokens, in / out$10.00 / $50.00$5.00 / $30.00
1,000 consult exchanges$55.00$31.00
released2026-06-092026-07-09
licenseproprietaryproprietary

Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.

Reading this pairing

This is the flagship pairing, and it is not close by this board's standards: 0.055 is the widest gap between neighboring models in the top seven, and Claude Fable 5 is the only model above 0.61. The lead is priced accordingly. Fable 5 lists at twice Sol's input rate and $20 more per million output tokens, so a high-volume deployment pays roughly double for the last increment of rubric adherence. Where single hard consults matter more than volume, the gap is the point; where volume dominates, Sol is the sensible flagship.

Which scores higher on HealthBench Professional, Claude Fable 5 or GPT-5.6 Sol?

Claude Fable 5 scores higher: 0.660 against GPT-5.6 Sol's 0.605, a difference of 0.055 on the 525-task set, as of August 16, 2026.

Which is cheaper to run, Claude Fable 5 or GPT-5.6 Sol?

GPT-5.6 Sol. 1,000 typical consult exchanges (2,000 input and 700 output tokens each) cost $31.00 against $55.00 at list rates.

Related comparisons

Full results for both models: Claude Fable 5 and GPT-5.6 Sol. The complete score-difference matrix is on the compare page.