RuntimeWire Combined Benchmark

9 models · 3 suites · frozen combined leaderboard

RankModelCombined score (0-100)Suite coverage
1OpenAI: GPT-6.1 Sol Pro97.63/3
2OpenAI: GPT-6 Astra96.83/3
3Anthropic: Claude Opus 5.593.53/3
4Anthropic: Claude Sonnet 5.593.23/3
5xAI: Grok Latest88.33/3
6Qwen: Qwen3.8 Max Prime82.63/3
7Z.ai: GLM 5.3 Prime81.43/3
8Google: Gemini 3.8 Flash76.03/3
9StepFun: Step 3.7 Flash61.23/3

Benchmark suites

Combined scores are the equal-weight mean of each model’s suite scores.

See every leaderboard on the Benchmarks index.