AlphaGo co-creator Thore Graepel seeks tens of millions for reasoning startup
Bloomberg says Graepel is seeking tens of millions from initial backers, with a possible later tranche in the hundreds of millions; neither is reported closed.
By RuntimeWire Staff · Published
Primary source: Bloomberg Technology
Why it matters
Graepel is raising against a research thesis rooted in AlphaGo: combining learned judgment with search and planning. The reported amounts are targets, not closed funding, and the product and commercial path remain unspecified.

Thore Graepel (@ThoreG), the machine-learning researcher who helped build AlphaGo, is seeking tens of millions of dollars for a new AI reasoning startup, Bloomberg reported on September 24th. The initial financing is being sought from a small group of backers, according to people familiar with the matter. Bloomberg also reports that Graepel could later pursue hundreds of millions of dollars at a higher valuation. Neither figure describes a completed raise.
Graepel's pitch starts with a technical idea he has pursued across very different settings: give machines ways to reason through uncertainty and choose what to do next. On his personal site, he describes the venture as an effort to bring AlphaGo-style reasoning to frontier AI, so systems can plan and act under uncertainty. He connects that goal to a career spanning probabilistic models, game-playing systems, and agents that learn through self-play.
The scale of the fundraising plan is striking next to the limited detail disclosed about the venture. The startup's name, product, architecture, customers, and commercial model have not been identified in the reporting. Graepel's public description lays out a research direction, not a product specification. The question for potential backers is how a method honed in controlled environments such as games will translate into useful systems for open-ended tasks.
A career built around uncertainty
Graepel studied physics at the University of Hamburg and Imperial College London before earning a PhD in machine learning from the Technical University of Berlin in 2001, according to his biography at the London Institute for Mathematical Sciences. He joined Microsoft Research Cambridge in 2003, where he worked on TrueSkill, a Bayesian ranking system used for online gaming matchmaking. Microsoft describes the system as tracking uncertainty about players' skill, rather than assigning each player a single fixed rating.
That work gives the new venture a less obvious precedent than AlphaGo. TrueSkill was a deployed system for making decisions from incomplete evidence; it had to estimate skill from game outcomes and update those estimates as more results arrived. Graepel's personal account also cites AdPredictor, a click-through prediction system, among his Microsoft projects. These are earlier examples of probabilistic reasoning applied to practical decisions, rather than a claim that his new company already has a comparable product.
At DeepMind, Graepel worked on AlphaGo and its successors. The 2016 Nature paper introducing AlphaGo lists him among its authors and describes a system combining neural networks with search. That combination is central to Graepel's current thesis: learned models can evaluate possibilities, while search examines candidate paths before a system acts. His site also points to later work on multi-agent learning and self-play, where systems improve by interacting with other agents or versions of themselves.
After his first DeepMind stint, Graepel led machine-learning work at Altos Labs, then returned to Google's research lab in 2025, according to the timeline on his personal site. Bloomberg reports that he left the lab during summer 2026 to pursue the venture. He also holds a Chair of Machine Learning position at University College London, according to the university and his public biographies.
The technical bet behind the raise
Graepel is proposing a route from game-playing systems to broader AI: combine learned judgment with planning, then apply that approach where an agent must act under uncertainty. The boundary between that ambition and a general commercial product remains open. His public description mentions frontier AI and embodied intelligence, but does not establish a finished model, target customer, benchmark, or deployment plan.
In AI, "reasoning" covers products with very different methods and claims. Google's AlphaProof Nexus, which RuntimeWire covered in May, used agentic search in formal mathematics; Graepel's venture is a separate effort, and there is no disclosed product to compare directly. His track record gives the idea technical lineage; the startup has yet to disclose how it will measure progress outside research demonstrations or games.
Bloomberg's financing account describes a proposed sequence: tens of millions first, with a possible later tranche in the hundreds of millions. Bloomberg describes a fundraising plan; it reports no investor commitments or valuation. No backers are named in the report. The size and staging suggest the ambition could demand substantial capital, but the source does not specify what that capital would fund or when the later raise might occur.
Graepel's founder story is unusually continuous: the same broad problem, machine decision-making under uncertainty, appears in online matchmaking, Go, multi-agent research, and now his proposed work on frontier AI. The venture's test will be whether that continuity yields a system that performs beyond the settings where search and planning have already worked well. For now, the reported raise attaches a large financing ambition to a clearly stated research thesis, while leaving the company and its first product undefined.