Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks

In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59.

By · Published

BBEH Mini benchmark leaderboard chart

Leaderboard

Rank Model Mean score Est. cost/task
1 SpaceXAI: Grok 4.6 0.67 $0.0034
2 Anthropic: Claude Opus 4.8 0.62 $0.0485
3 OpenAI: GPT-5.6 Sol Pro 0.59 $0.0089

How we scored it

Every model answered the same 460-task battery from BBEH Mini (BBEH — Google DeepMind (Apache-2.0)), one task at a time, with no tools and no retries on content.

460 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.


Explore every prompt, answer, and per-task grade in the interactive leaderboard.

Reader comments

Conversation for this story loads after sign-in.