Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79

In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.79 at an estimated $0.0056 per task. OpenAI’s GPT-5.6 Sol followed at 0.77, while several GPT-5.6 variants and Anthropic’s Claude Opus 4.8 clustered close behind.

By · Published

Newsroom Reliability v0.2 benchmark leaderboard chart

Leaderboard

Rank Model Mean score Est. cost/task
1 SpaceXAI: Grok 4.6 0.79 $0.0056
2 OpenAI: GPT-5.6 Sol 0.77 $0.0217
3 OpenAI: GPT-5.6 Sol Pro 0.76 $0.0217
4 OpenAI: GPT-5.6 Luna Pro 0.76 $0.0005
5 OpenAI: GPT-5.6 Luna 0.76 $0.0005
6 OpenAI: GPT-5.6 Terra 0.75 $0.0045
7 Anthropic: Claude Opus 4.8 0.74 $0.0223
8 Inkling Small 0.73
9 OpenAI: GPT-5.6 Terra Pro 0.72 $0.0045
10 Anthropic: Claude Opus 4.8 (Fast) 0.72 $0.0436
11 Kimi K3 0.71
12 Inkling FP4 0.70
13 GLM 5.2 0.69
14 Google: Gemini 3.6 Flash 0.67 $0.0056
15 Gemini 3.7 Flash 0.66
16 DeepSeek-V4-Pro 0.63
17 Qwen: Qwen3.7 Max 0.62 $0.0040
18 DeepSeek-V4-Flash 0.61

How we scored it

Every model answered the same 50-task battery from Newsroom Reliability v0.2, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.


Explore every prompt, answer, and per-task grade in the interactive leaderboard.

Reader comments

Conversation for this story loads after sign-in.