RuntimeWire ranks nine AI models on news research, summaries and writing
GPT-6.1 Sol Pro leads the equal-weight index, which combines three task suites with samples ranging from 50 to 100 tasks.
By Ryan Merket · Published · Updated
Primary source: RuntimeWire
Why it matters
A newsroom-specific composite can help teams compare models on editorial work, but its equal weighting and small samples make the ranking a starting point, not a universal measure of model quality.

RuntimeWire has published a combined leaderboard ranking nine AI models across news research, grounded summarization and constrained story writing, with OpenAI's GPT-6.1 Sol Pro at the top on a score of 97.6 out of 100. The benchmark took 48 hours to complete. Founder and editor in chief Ryan Merket (@ryanmerket) is extending RuntimeWire's work to examine a question relevant to the newsroom's operation: how models perform on editorial tasks, rather than on general-purpose chat alone.

The next two places also go to OpenAI models: GPT-6 Astra scored 96.8, and Anthropic's Claude Opus 5.5 scored 93.5. Claude Sonnet 5.5 followed at 93.2. The rest of the list is xAI's Grok Latest at 88.3, Qwen3.8 Max Prime at 82.6, Z.ai's GLM 5.3 Prime at 81.4, Google's Gemini 3.8 Flash at 76.0, and StepFun's Step 3.7 Flash at 61.2. RuntimeWire labels the table a frozen leaderboard.
The combined score is the equal-weight mean of each model's score on three suites: a 50-task sample from NEWSAGENT, a 100-task public-v1 sample from Vectara's Grounded Summarization evaluation, and a 50-task sample for Mazur Constrained Story Writing. That design gives each suite one-third of the composite, even though the Vectara sample contains twice as many tasks as either of the other two. The composite is an average across three types of work, not a simple pooled accuracy rate across all tasks.
The 97.6 is a composite score, not a claim that GPT-6.1 Sol Pro answered 97.6% of questions correctly. The leaderboard also shows the top two separated by 0.8 points, while Claude Opus 5.5 is 3.3 points behind GPT-6 Astra. Those gaps describe this index's scoring system; by themselves, they do not establish how meaningful the differences would be on another task set or in a newsroom's day-to-day workflow.
The suites cover a narrower use case than broad model rankings. NEWSAGENT evaluates work involved in producing a news article from source material, including finding and selecting information. Vectara's suite focuses on grounded summarization. Mazur's tests constrained story writing. Together, the three scores assess whether a model can gather relevant material, summarize against supplied information and write within requirements. The index does not rank coding, reasoning, multimodal ability or overall model quality.
The page gives the suite names, sample sizes, combined scores and averaging rule. The composite alone does not show which suite drove a particular model's placement, so the leaderboard is a first comparison rather than a definitive purchasing guide. Model selection for a production newsroom would also depend on factors the composite does not measure, such as latency, cost, reliability over repeated runs and the amount of human correction required. RuntimeWire's table does not claim to settle those operational questions.
In a May 2026 account of building RuntimeWire, Merket described the bet as using automation to keep a one-person newsroom publishing across the day. RuntimeWire's own author profile says Merket held early product roles at Reddit and Facebook before founding the publication. The benchmark applies that automation approach to model evaluation by comparing performance on work closer to reporting and editing, rather than using a single general score.
The comparison also sits alongside a different model-ranking tradition. Arena's leaderboard draws on people comparing anonymous model responses and voting for the one they prefer. RuntimeWire's composite instead averages reported performance across fixed editorial and summarization suites. The results answer different questions: preference voting reflects which answer people choose in a given comparison, while this index is structured around specific task sets and a declared averaging method.
For teams considering AI in editorial workflows, the leaderboard offers a short list of models to examine. On RuntimeWire's three selected suites and equal-weight scoring, GPT-6.1 Sol Pro ranked first, followed closely by GPT-6 Astra. The score does not show that either will be the best model for every newsroom, or that a small lead on this composite will hold across a different mix of tasks.