Head to head: diffusiongemma-26b-a4b-it vs gpt-oss-20b
diffusiongemma-26b-a4b-it vs gpt-oss-20b
By Ryan Merket · Published
This was a low-scoring, narrow matchup in practice, but diffusiongemma-26b-a4b-it still finished ahead by being the only model to convert a task win. The result is a lean verdict, not a rout: one correct politics answer separated two models that otherwise mostly failed together.
On the raw scoreboard, diffusiongemma-26b-a4b-it posts a perfect aggregate edge, 10.0 to 0.0. But the task-level picture is much less dramatic: out of six SimpleQA prompts, five were ties and only one produced separation. That matters, because this is not a broad-based demolition so much as a narrow but real advantage. The decisive moment came in **Politics**, where diffusiongemma-26b-a4b-it got the benchmark answer exactly right and gpt-oss-20b missed. Everywhere else, the two models moved in lockstep — and not in a flattering way. Both were wrong on the other politics item, both missed both geography questions, both missed sports, and both missed science and technology. So the editorial read is straightforward: diffusiongemma-26b-a4b-it was the only model here that showed any ability to break out of the pack. gpt-oss-20b never managed a single task win, which is hard to defend even in a tiny sample. At the same time, with five ties driven by shared errors, this matchup says more about one model finding a single opening than about either model demonstrating consistent factual command. The statistical verdict gets the tone right: **75% confidence, lean**. That's enough to call a winner, but not enough to pretend the gap is settled beyond debate. **Final call: diffusiongemma-26b-a4b-it wins on the strength of the only clean hit in an otherwise error-heavy contest.**
SimpleQA · Politics
What was P. V. Sanjay Kumar's position just before being appointed as a judge of the Supreme Court of India? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.
SimpleQA · Geography
During which years was Otto Schlüter a professor of geography at the University of Halle? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.
SimpleQA · Sports
What was the halftime score between Liverpool and Leicester in the game from December 30, 2022? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.
SimpleQA · Geography
On what month, day, and year was the cave known as Ursa Minor first discovered in Sequoia National Park, California, United States? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.
SimpleQA · Politics
What is the initial name of the political party that Emmanuel Macron founded? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was correct, Model B was incorrect.
SimpleQA · Science and technology
Who originally recorded the scientific name of the white-headed duck as *Anas leucocephala* in 1769? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>
Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.
Matchup powered by OpenRouter.