Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.78
In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.78. OpenAI’s GPT-5.6 Luna Pro followed at 0.77 with the lowest reported cost among the top three, at $0.0005 per task.
By RuntimeWire Staff · Published

Leaderboard
| Rank | Model | Mean score | Est. cost/task |
|---|---|---|---|
| 1 | SpaceXAI: Grok 4.6 | 0.78 | $0.0056 |
| 2 | OpenAI: GPT-5.6 Luna Pro | 0.77 | $0.0005 |
| 3 | OpenAI: GPT-5.6 Sol Pro | 0.76 | $0.0215 |
| 4 | Meta: Muse Spark 1.2 | 0.73 | $0.0044 |
| 5 | Anthropic: Claude Opus 4.8 | 0.71 | $0.0211 |
| 6 | Google: Gemini 3.7 Flash | 0.67 | $0.0014 |
How we scored it
Every model answered the same 50-task battery from Newsroom Reliability v0.2, one task at a time, with no tools and no retries on content.
50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.
Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.
A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.
Explore every prompt, answer, and per-task grade in the interactive leaderboard.