HealthBench Professional

GPT-5.6 Luna vs GPT-5.5 Instant on HealthBench Professional

updated August 16, 2026

GPT-5.6 Luna scores 0.557 to GPT-5.5 Instant's 0.384, a gap of 0.173 on the 525 physician-graded tasks of HealthBench Professional. The table below puts the scores next to what each model costs to actually run.

Side by side

OpenAI logoGPT-5.6 LunaOpenAI logoGPT-5.5 Instant
score0.5570.384
rank7 of 98 of 9
context window1.1M400K
price per 1M tokens, in / out$0.20 / $1.20$5.00 / $30.00
1,000 consult exchanges$1.24$31.00
released2026-07-092026-05-05
licenseproprietaryproprietary

Consult exchange: 2,000 input and 700 output tokens, priced at list rates as of August 16, 2026. GPT-5.6 models charge higher rates above 272K input tokens. MAI-Thinking-1 is in public preview on Microsoft Foundry without final list pricing.

Reading this pairing

The most lopsided pairing on the board: Luna scores 0.173 higher at a twenty-fifth of the input price. GPT-5.5 Instant predates the 5.6 family and its July 2026 repricing, and this table shows what that generation gap means for healthcare work. Luna is also the cheapest model within 0.11 of the leader, which makes it the default answer for high-volume clinical text at low cost.

Which scores higher on HealthBench Professional, GPT-5.6 Luna or GPT-5.5 Instant?

GPT-5.6 Luna scores higher: 0.557 against GPT-5.5 Instant's 0.384, a difference of 0.173 on the 525-task set, as of August 16, 2026.

Which is cheaper to run, GPT-5.6 Luna or GPT-5.5 Instant?

GPT-5.6 Luna. 1,000 typical consult exchanges (2,000 input and 700 output tokens each) cost $1.24 against $31.00 at list rates.

Full results for both models: GPT-5.6 Luna and GPT-5.5 Instant. The complete score-difference matrix is on the compare page.