Meta、5つのSTEMオリンピアードで金メダル相当のAI成果を主張
Metaは、モデルや評価の設定を明示せずに、物理学理論のスコアが満点であると報告した。
By Ryan Merket · Published
Primary source: AI at Meta on X
Why it matters
Olympiad exams offer harder, fresher reasoning tests than saturated benchmarks, but Meta's scores cannot be compared cleanly until it identifies the models and test-time setup.

Meta said in an August 6th post that its AI models achieved gold-medal-level performance across five STEM Olympiad competitions, including perfect scores on the theoretical exams at the Asian Physics Olympiad and International Physics Olympiad.
https://x.com/AIatMeta/status/2085388945148297322
The announcement packages the competitions as a test of whether Meta's models are making "genuine progress on reasoning." That is a higher bar than another benchmark score. Olympiad problems are written for each year's competition, require long derivations and can include diagrams or other visual information. The answers are graded for intermediate work, rather than a final multiple-choice selection.
Metaの発表は、これらの競技を同社のモデルが「推論において真の進歩を遂げているか」を試すテストとして位置づけている。これは単なるベンチマークスコアよりも高いハードルだ。オリンピアードの問題は各年の競技のために新たに作成され、長い導出を要し、図やその他の視覚情報を含むことがある。採点は最終的な択一式の選択ではなく、中間作業に対して行われる。
Meta's accompanying text named the two physics results but did not identify which models produced them. It also left out the inference configuration, number of attempts, tool access, compute budget and whether Meta selected answers from multiple agents or runs. Those details determine how closely the result measures a single model's reasoning ability rather than the performance of a larger test-time system.
Metaの添付文では2件の物理の結果が示されたが、それらを生成したモデルがどれであるかは特定していない。また、推論の設定、試行回数、ツールの利用可否、計算予算、Metaが複数のエージェントや複数回の実行から回答を選んだかどうかといった点も記載されていない。これらの詳細は、結果が単一のモデルの推論能力をどれだけ直接的に測っているか、あるいはテスト時のより大きなシステムの性能を示しているかを決定する。
理論試験の満点は物理コンテストの一部をカバーしている
Meta specified that its perfect results applied to the theory exams. Both physics Olympiads also include experimental components, making the company's scores narrower than a complete human competition result.
Metaは、満点の結果が理論試験に適用されたものであると明示した。両方の物理オリンピアードには実験の要素も含まれており、同社のスコアは人間の競技全体の結果よりも範囲が限定される。
The Asian Physics Olympiad follows the structure of the International Physics Olympiad, with a five-hour theoretical examination and separate laboratory work. The 2026 Asian competition was held in Busan, South Korea, according to the event's official history.
[Asian Physics Olympiad]はInternational Physics Olympiadの構成に従っており、5時間の理論試験と別個の実験作業がある。[event's official history]によれば、2026年のアジア大会は韓国の釜山で開催された。
Meta previously said its Asian Physics Olympiad submission scored 30 out of 30 and tied the top three students. The August 6th post extended the company's account to the International Physics Olympiad, where Meta said a model also earned a perfect theory score.
Metaは以前、自社のAsian Physics Olympiadへの提出が30点満点中30点を記録し、上位3名と並んだと述べていた。8月6日の投稿では、その説明をInternational Physics Olympiadにも拡大し、同社はモデルがそちらでも理論試験で満点を獲得したと述べた。
That IPhO result matched a standard reached by human contestants this year. Two members of Vietnam's 2026 team earned 30 out of 30 on the theory exam, according to the country's Ministry of Education and Training as reported by VnExpress. The relevant comparison is therefore with the strongest human theory performances, without covering the laboratory exam that competitors also completed.
そのIPhOの結果は、今年の人間の競技者が達成した基準と一致する。VnExpressが報じるところによれば、Vietnamの2026年チームの2人が理論試験で30点満点を獲得したという。したがって比較すべきは、競技者が同時に行った実験試験を含まない、最も優れた人間の理論試験の成績である。
名を明かさないモデルが主張の中心である
Meta released Muse Spark 1.1 on July 9th, describing it as a multimodal reasoning model for coding, tool use and agentic tasks. The Olympiad announcement did not say whether Muse Spark 1.1, another public model or an internal research system generated the reported solutions.
Metaは7月9日に[Muse Spark 1.1]をリリースし、コーディング、ツール利用、エージェンシー的タスク向けのマルチモーダル推論モデルであると説明した。オリンピアードの発表では、報告された解答がMuse Spark 1.1によるものなのか、他の公開モデルなのか、あるいは内部の研究システムによるものなのかは述べられていない。
That distinction matters because Meta has explicitly developed reasoning systems that scale inference across several agents. When Meta introduced the original Muse Spark in April, it described a "Contemplating" mode that has multiple agents reason in parallel. A result produced by such orchestration can be useful evidence about the overall system, but it is not directly comparable with a single model answering once under a fixed token budget.
その区別は重要だ。というのもMetaは明示的に複数のエージェントに推論をスケールさせる推論システムを開発しているからである。Metaが4月にオリジナルの[Muse Spark]を紹介した際には、複数のエージェントが並列に推論する「Contemplating」モードが説明されていた。そのようなオーケストレーションによって生成された結果は全体システムに関する有用な証拠になり得るが、固定されたトークン予算の下で単一モデルが一回回答した場合と直接比較できるものではない。
The strongest version of Meta's case would pair the scorecard with the exact prompts, model checkpoint, time limits, tool permissions, full solutions and grading records. The August 6th announcement offered the headline results instead.
Metaの主張を最も強くするには、得点表とともに正確なプロンプト、モデルのチェックポイント、制限時間、ツールの許可、完全な解答および採点記録を提示することだろう。8月6日の発表は代わりに主要な結果のみを示した。
オリンピアード試験は飽和したベンチマークに代わりつつある
AI laboratories have increasingly turned to recent academic competitions because commonly cited benchmark suites can saturate or leak into training data. Live or newly released Olympiad papers reduce that contamination risk and test whether a system can sustain a chain of reasoning across unfamiliar problems.
AI研究所は、一般に引用されるベンチマークスイートが飽和したりトレーニングデータに流出したりするため、最近の学術競技にますます目を向けている。公開中の、あるいは新たに公開されたオリンピアード論文はその汚染リスクを減らし、システムが馴染みのない問題にわたって推論の連鎖を維持できるかを検証する。
The difficulty is still sensitive to methodology. A model can generate several candidate solutions, use external software, search the web or rely on a second model to critique its work. Human contestants operate under fixed time, tool and submission rules. Medal labels are meaningful only when the evaluation states which of those constraints were preserved.
難易度は依然として方法論に敏感だ。モデルは複数の候補解を生成したり、外部ソフトウェアを使用したり、ウェブ検索を行ったり、別のモデルに批評させたりできる。人間の競技者は固定された時間、ツール、提出ルールの下で競う。メダルのラベルが意味を持つのは、どの制約が保持されたかを評価が明示した場合に限られる。
Research published with the HiPhO physics benchmark illustrates how quickly performance has moved. Its evaluation of 30 language and multimodal models found that closed reasoning systems could reach gold-medal thresholds on several recent physics exams, although most models remained short of full marks. Meta is now claiming perfect theory results at two live 2026 competitions and gold-level performance across five Olympiads.
[HiPhO physics benchmark]とともに発表された研究は、性能がいかに急速に進んだかを示している。30の言語モデルおよびマルチモーダルモデルの評価では、クローズドな推論システムがいくつかの最近の物理学試験で金メダル基準に到達できることがわかったが、ほとんどのモデルは満点には達していなかった。Metaは現在、2026年の2つの現地競技で理論試験の満点を主張し、5つのオリンピアードで金メダル級の性能を示したとしている。
The scores put Meta in the front rank of scientific reasoning systems by the company's account. Establishing how much progress belongs to the underlying models will require the test configuration behind them.
これらのスコアは、同社の説明によればMetaを科学的推論システムの最前列に置くものだ。どれだけの進歩が基礎となるモデル自体に帰属するのかを確定するには、それらの背後にあるテスト設定が必要となる。