AI Model Benchmarks: Leaderboards

The RuntimeWire Composite Index averages each model's published leaderboard scores into one 0–100 ranking, alongside N-way benchmarks where 2–30 models run the same task suite, scored objectively or by a blind LLM judge.

RuntimeWire Composite Index

Incorporates 3 evaluations (BBEH Mini, Editorial Craft, Newsroom Reliability v0.2) — the equal-weight average of each model's published leaderboard scores on a 0–100 scale.

  1. Z.ai: GLM 5.3 Flash — 81.9 (covers 2/3 evaluations)
  2. Claude Opus 5 (Fast) — 80.5 (covers 2/3 evaluations)
  3. Qwen: Qwen3.8 Flash — 79.6 (covers 2/3 evaluations)
  4. Google: Gemini 3.7 Flash — 78.1 (covers 2/3 evaluations)
  5. step-3.7-flash — 75.3 (covers 2/3 evaluations)
  6. DeepSeek-V4-Flash-0731 — 71.7 (covers 1/3 evaluations)
  7. SpaceXAI: Grok 4.6 — 66.5 (covers 1/3 evaluations)
  8. Anthropic: Claude Opus 4.8 — 62.0 (covers 1/3 evaluations)
  9. OpenAI: GPT-5.6 Sol Pro — 58.7 (covers 1/3 evaluations)