GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score

In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models.

By · Published

Editorial Craft benchmark leaderboard chart

Leaderboard

Rank Model Mean score Est. cost/task
1 OpenAI: GPT-5.6 Sol Pro 0.97 $0.0091
2 OpenAI: GPT-5.6 Luna 0.94 $0.0007
3 OpenAI: GPT-5.6 Terra Pro 0.94 $0.0073
4 OpenAI: GPT-5.6 Luna Pro 0.93 $0.0007
5 SpaceXAI: Grok 4.6 0.93 $0.0038
6 OpenAI: GPT-5.6 Terra 0.93 $0.0075
7 Anthropic: Claude Opus 4.8 0.93 $0.0160
8 OpenAI: GPT-5.6 Sol 0.92 $0.0091
9 Anthropic: Claude Sonnet 5 0.88 $0.0063
10 Anthropic: Claude Fable 5 0.80 $0.0297
11 Google: Gemini 3.7 Flash 0.78 $0.0011
12 Phi-4-reasoning 0.47

How we scored it

Every model answered the same 12-task battery from Editorial Craft, one task at a time, with no tools and no retries on content.

2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.


Explore every prompt, answer, and per-task grade in the interactive leaderboard.

Reader comments

Conversation for this story loads after sign-in.