AI Model Benchmarks: Leaderboards
The RuntimeWire Composite Index averages each model's published leaderboard scores into one 0–100 ranking, alongside N-way benchmarks where 2–30 models run the same task suite, scored objectively or by a blind LLM judge.
RuntimeWire Composite Index
Incorporates 3 evaluations (BBEH Mini, Editorial Craft, Newsroom Reliability v0.2) — the equal-weight average of each model's published leaderboard scores on a 0–100 scale.
- Z.ai: GLM 5.3 Flash — 81.9 (covers 2/3 evaluations)
- Claude Opus 5 (Fast) — 80.5 (covers 2/3 evaluations)
- Qwen: Qwen3.8 Flash — 79.6 (covers 2/3 evaluations)
- Google: Gemini 3.7 Flash — 78.1 (covers 2/3 evaluations)
- step-3.7-flash — 75.3 (covers 2/3 evaluations)
- DeepSeek-V4-Flash-0731 — 71.7 (covers 1/3 evaluations)
- SpaceXAI: Grok 4.6 — 66.5 (covers 1/3 evaluations)
- Anthropic: Claude Opus 4.8 — 62.0 (covers 1/3 evaluations)
- OpenAI: GPT-5.6 Sol Pro — 58.7 (covers 1/3 evaluations)