Head to head: OpenAI: GPT-5.6 Luna vs Grok 4.20 0309 Non Reasoning

OpenAI: GPT-5.6 Luna vs Grok 4.20 0309 Non Reasoning

By · Published

RuntimeWire Head-to-Head: Head to head: OpenAI: GPT-5.6 Luna vs Grok 4.20 0309 Non Reasoning
RuntimeWire Head-to-Head matchup

This was a low-separation matchup, but OpenAI’s GPT-5.6 Luna did enough to edge Grok 4.20 0309 Non Reasoning where it counted. The scoreline flatters neither model: one clean task win, five ties, and a statistical verdict that says Luna is ahead—but only by a lean margin.

Most of this contest was defined by mutual non-performance. Of the six tasks, five were ties, and four of those ties came because both models missed the question. That matters: this was not Luna dominating the board, and it was not Grok hanging tough through strong parity. It was largely a stalemate built on shared failures. The only real separation came in **SimpleQA · Politics**, where **GPT-5.6 Luna got the exact-match answer right and Grok missed**. That single hit is the entire practical basis for Luna’s aggregate lead, and in a benchmark this small, that’s enough to matter. On **Science and technology**, both models split their performance the same way—one correct tie, one incorrect tie—so Grok never found a category where it could claw back ground. The aggregate score, **20.0 to 10.0**, makes Luna look comfortably ahead, but the task-level record tells the truer story: **1 win, 0 losses, 5 ties**. That’s why the statistical read is appropriately restrained at **75% confidence (lean)**. Luna was better in this sample, but not by the kind of margin that justifies chest-thumping. **Final call: GPT-5.6 Luna wins, narrowly but legitimately.** It was the only model here to convert a contested task into a point on the board, while Grok 4.20 0309 Non Reasoning never produced a category win of its own. That’s enough for the verdict—even if this head-to-head says as much about both models’ limits as it does about Luna’s edge.

SimpleQA · Science and technology

Who originally recorded the scientific name of the white-headed duck as *Anas leucocephala* in 1769? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: The name *Anas leucocephala* was originally recorded by Italian naturalist Giovanni Antonio Scopoli in 1769. Final answer: Giovanni Antonio Scopoli
Grok 4.20 0309 Non Reasoning: Final answer: Giovanni Antonio Scopoli

Scored against the benchmark's reference answer (exact match): Model A was correct, Model B was correct.

SimpleQA · Politics

What was P. V. Sanjay Kumar's position just before being appointed as a judge of the Supreme Court of India? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: P. V. Sanjay Kumar was the Chief Justice of the Manipur High Court. Final answer: Chief Justice of the Manipur High Court.
Grok 4.20 0309 Non Reasoning: **Chief Justice of the High Court of Jammu and Kashmir and Ladakh** Final answer: Chief Justice of the High Court of Jammu and Kashmir and Ladakh

Scored against the benchmark's reference answer (exact match): Model A was correct, Model B was incorrect.

SimpleQA · Sports

What was the halftime score between Liverpool and Leicester in the game from December 30, 2022? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: Liverpool led Leicester 2–1 at halftime. Final answer: 2–1 to Liverpool
Grok 4.20 0309 Non Reasoning: Halftime score was Liverpool 1, Leicester 0. Final answer: 1-0

Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.

SimpleQA · Sports

On what day, month, and year did the first Bennett Cup auto race, for a prize sponsored by New York Herald publisher James Gordon Bennett Jr., begin as five entrants departed from the Parc de Saint-Cloud, near Paris, on a 566-kilometer (352 miles) trip to Lyon? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: The first Bennett Cup auto race began on 14 June 1900. Final answer: 14 June 1900
Grok 4.20 0309 Non Reasoning: The first Bennett Cup auto race began on 26 May 1900. Final answer: 26 May 1900

Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.

SimpleQA · Geography

According to the 2011 Nepal census, what percentage of the population of Nepal speaks Urdu? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: According to Nepal’s 2011 census, Urdu was spoken by **2.61%** of the population. Final answer: 2.61%
Grok 4.20 0309 Non Reasoning: According to the 2011 Nepal census, 0.01% of the population speaks Urdu as their first language. Final answer: 0.01%

Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.

SimpleQA · Science and technology

In which year did Pramod Kale (an Indian engineer) win the Shri Hari Om Ashram Prerit Vikram Sarabhai Award for System Analysis and Management Problems? Answer concisely, then give ONLY the short factual answer on the last line as: Final answer: <answer>

OpenAI: GPT-5.6 Luna: He won the award in **1970**. Final answer: 1970
Grok 4.20 0309 Non Reasoning: **1984** Pramod Kale, an Indian engineer known for his work in satellite technology and systems engineering at ISRO, received the Shri Hari Om Ashram Prerit Vikram Sarabhai Award for System Analysis and Management Problems in 1984 for his contributions to space systems management and project analysis. Final answer: 1984

Scored against the benchmark's reference answer (exact match): Model A was incorrect, Model B was incorrect.

Matchup powered by OpenRouter.