Meta claims gold-level AI results across five STEM Olympiads
Meta reported perfect physics theory scores, while leaving the models and evaluation setup unspecified.
By Ryan Merket · Published
Why it matters
Olympiad exams offer harder, fresher reasoning tests than saturated benchmarks, but Meta's scores cannot be compared cleanly until it identifies the models and test-time setup.

Meta said in an August 6th post that its AI models achieved gold-medal-level performance across five STEM Olympiad competitions, including perfect scores on the theoretical exams at the Asian Physics Olympiad and International Physics Olympiad.
https://x.com/AIatMeta/status/2085388945148297322
The announcement packages the competitions as a test of whether Meta's models are making "genuine progress on reasoning." That is a higher bar than another benchmark score. Olympiad problems are written for each year's competition, require long derivations and can include diagrams or other visual information. The answers are graded for intermediate work, rather than a final multiple-choice selection.
Meta's accompanying text named the two physics results but did not identify which models produced them. It also left out the inference configuration, number of attempts, tool access, compute budget and whether Meta selected answers from multiple agents or runs. Those details determine how closely the result measures a single model's reasoning ability rather than the performance of a larger test-time system.
Perfect theory scores cover part of the physics contests
Meta specified that its perfect results applied to the theory exams. Both physics Olympiads also include experimental components, making the company's scores narrower than a complete human competition result.
The Asian Physics Olympiad follows the structure of the International Physics Olympiad, with a five-hour theoretical examination and separate laboratory work. The 2026 Asian competition was held in Busan, South Korea, according to the event's official history.
Meta previously said its Asian Physics Olympiad submission scored 30 out of 30 and tied the top three students. The August 6th post extended the company's account to the International Physics Olympiad, where Meta said a model also earned a perfect theory score.
That IPhO result matched a standard reached by human contestants this year. Two members of Vietnam's 2026 team earned 30 out of 30 on the theory exam, according to the country's Ministry of Education and Training as reported by VnExpress. The relevant comparison is therefore with the strongest human theory performances, without covering the laboratory exam that competitors also completed.
The unnamed model is central to the claim
Meta released Muse Spark 1.1 on July 9th, describing it as a multimodal reasoning model for coding, tool use and agentic tasks. The Olympiad announcement did not say whether Muse Spark 1.1, another public model or an internal research system generated the reported solutions.
That distinction matters because Meta has explicitly developed reasoning systems that scale inference across several agents. When Meta introduced the original Muse Spark in April, it described a "Contemplating" mode that has multiple agents reason in parallel. A result produced by such orchestration can be useful evidence about the overall system, but it is not directly comparable with a single model answering once under a fixed token budget.
The strongest version of Meta's case would pair the scorecard with the exact prompts, model checkpoint, time limits, tool permissions, full solutions and grading records. The August 6th announcement offered the headline results instead.
Olympiad tests are replacing saturated benchmarks
AI laboratories have increasingly turned to recent academic competitions because commonly cited benchmark suites can saturate or leak into training data. Live or newly released Olympiad papers reduce that contamination risk and test whether a system can sustain a chain of reasoning across unfamiliar problems.
The difficulty is still sensitive to methodology. A model can generate several candidate solutions, use external software, search the web or rely on a second model to critique its work. Human contestants operate under fixed time, tool and submission rules. Medal labels are meaningful only when the evaluation states which of those constraints were preserved.
Research published with the HiPhO physics benchmark illustrates how quickly performance has moved. Its evaluation of 30 language and multimodal models found that closed reasoning systems could reach gold-medal thresholds on several recent physics exams, although most models remained short of full marks. Meta is now claiming perfect theory results at two live 2026 competitions and gold-level performance across five Olympiads.
The scores put Meta in the front rank of scientific reasoning systems by the company's account. Establishing how much progress belongs to the underlying models will require the test configuration behind them.