Muse Spark 1.1 on HealthBench Professional
rank 12 of 31 · updated September 30, 2026
Muse Spark 1.1 scores 0.593 on HealthBench Professional, rank 12 of 31 models on the board. Meta's July 2026 update to Muse Spark, sold through the Meta Model API at $1.25 per million input tokens. HealthBench Professional scores models on 525 tasks drawn from real clinician conversations, graded criterion by criterion against physician-written rubrics on a 0 to 1 scale.
Result and API facts
| rank | 12 of 31 |
|---|---|
| score | 0.593 |
| lab | Meta |
| context window | 1.0M tokens |
| API price per 1M tokens | $1.25 in / $4.25 out |
| license | proprietary |
| source | Muse Spark 1.1 Evaluation Report (model card) |
| released | 2026-07-09 |
Position in the field
The gap to the leader, GPT-6 Astra (Anthropic run) at 0.703, is 0.110. Directly above sits Claude Opus 5 at 0.598. Directly below sits Claude Sonnet 5 at 0.578. The scores on this page are compiled from published documents rather than from one controlled run, so a small gap between 2 models can reflect a difference in grader version, reasoning effort, or deployment setting as well as a difference in capability.
Source of this score
Read from Muse Spark 1.1 Evaluation Report (model card, Meta, 2026-07-09). Vendor-reported. Confidence: verified. Configuration: length-normalized, GPT-5.4 low-reasoning grader, xhigh reasoning via Meta Model API (Muse Spark 1.1 Evaluation Report Figure 44).
Health | HealthBench Professional | 59.3 | 54.1 | 41.6 | 55.8 | 51.8 (Figure 44 image table row; columns Muse Spark 1.1, Muse Spark, Gemini 3.1 Pro (high), Opus 4.8 (max), GPT 5.5 (xhigh))
p. 101, Figure 44 'General capability benchmark results' (image), row HealthBench Professional, column Muse Spark 1.1; protocol p. 104 (printed 103): HealthBench Pro comprises 525 evaluation data points graded by rubrics. We use GPT-5.4 with low reasoning effort as the grader and report the length-normalized rubric score as done in their paper. · full entry on the sources page
What does Muse Spark 1.1 score on HealthBench Professional?
Muse Spark 1.1 scores 0.593 on HealthBench Professional, which places it at rank 12 of 31 models on the board as of September 30, 2026. The number was read from Muse Spark 1.1 Evaluation Report, listed on the sources page.
How much does Muse Spark 1.1 cost to run?
Muse Spark 1.1 is priced at $1.25 per million input tokens and $4.25 per million output tokens through Meta's API.
Head to head
Pairings with a dedicated comparison page are linked; every other difference is in the score-difference matrix.
- Muse Spark 1.1 vs GPT-6 Astra (Anthropic run)0.593 vs 0.703 · GPT-6 Astra (Anthropic run) by 0.110
- Muse Spark 1.1 vs Claude Sonnet 5.50.593 vs 0.692 · Claude Sonnet 5.5 by 0.099
- Muse Spark 1.1 vs Claude Fable 50.593 vs 0.660 · Claude Fable 5 by 0.067
- Muse Spark 1.1 vs Claude Opus 5.50.593 vs 0.656 · Claude Opus 5.5 by 0.063
- Muse Spark 1.1 vs GPT-6 Astra0.593 vs 0.647 · GPT-6 Astra by 0.054
- Muse Spark 1.1 vs Claude Fable 5 (September card)0.593 vs 0.633 · Claude Fable 5 (September card) by 0.040
- Muse Spark 1.1 vs Claude Fable 5.10.593 vs 0.621 · Claude Fable 5.1 by 0.028
- Muse Spark 1.1 vs GPT-6 Luna0.593 vs 0.608 · GPT-6 Luna by 0.015
- Muse Spark 1.1 vs GPT-6 Sol0.593 vs 0.608 · GPT-6 Sol by 0.015
- Muse Spark 1.1 vs GPT-5.6 Sol0.593 vs 0.605 · GPT-5.6 Sol by 0.012
- Muse Spark 1.1 vs Claude Opus 50.593 vs 0.598 · Claude Opus 5 by 0.005
- Muse Spark 1.1 vs Claude Sonnet 50.593 vs 0.578 · Muse Spark 1.1 by 0.015
- Muse Spark 1.1 vs GPT-5.6 Terra0.593 vs 0.577 · Muse Spark 1.1 by 0.016
- Muse Spark 1.1 vs Claude Opus 4.8 (Opus 4.8 grader)0.593 vs 0.574 · Muse Spark 1.1 by 0.019
- Muse Spark 1.1 vs Grok 4.70.593 vs 0.567 · Muse Spark 1.1 by 0.026
- Muse Spark 1.1 vs Claude Opus 4.80.593 vs 0.558 · Muse Spark 1.1 by 0.035
- Muse Spark 1.1 vs GPT-5.6 Luna0.593 vs 0.557 · Muse Spark 1.1 by 0.036
- Muse Spark 1.1 vs Muse Spark0.593 vs 0.541 · Muse Spark 1.1 by 0.052
- Muse Spark 1.1 vs GPT-5.6 Sol (August)0.593 vs 0.540 · Muse Spark 1.1 by 0.053
- Muse Spark 1.1 vs Claude Opus 4.70.593 vs 0.519 · Muse Spark 1.1 by 0.074
- Muse Spark 1.1 vs GPT-5.50.593 vs 0.518 · Muse Spark 1.1 by 0.075
- Muse Spark 1.1 vs Grok 4.60.593 vs 0.485 · Muse Spark 1.1 by 0.108
- Muse Spark 1.1 vs GPT-5.40.593 vs 0.481 · Muse Spark 1.1 by 0.112
- Muse Spark 1.1 vs GPT-50.593 vs 0.462 · Muse Spark 1.1 by 0.131
- Muse Spark 1.1 vs GPT-5.20.593 vs 0.459 · Muse Spark 1.1 by 0.134
- Muse Spark 1.1 vs Claude Sonnet 4.60.593 vs 0.442 · Muse Spark 1.1 by 0.151
- Muse Spark 1.1 vs GPT-5.6 Luna (August)0.593 vs 0.441 · Muse Spark 1.1 by 0.152
- Muse Spark 1.1 vs GPT-5.10.593 vs 0.396 · Muse Spark 1.1 by 0.197
- Muse Spark 1.1 vs GPT-5.5 Instant0.593 vs 0.384 · Muse Spark 1.1 by 0.209
- Muse Spark 1.1 vs MAI-Thinking-10.593 vs 0.350 · Muse Spark 1.1 by 0.243
Where the scores come from and how they are read is on the methodology page. The full ranking is on the leaderboard.