OpenAI says Moonshot-linked operators tried to extract hidden model reasoning

OpenAI says a July campaign generated 16,000 extraction-pattern requests over two days, but those figures measure attempts, not confirmed successes.

By · Published

Primary source: OpenAI

Why it matters

The dispute exposes a security problem in the API mechanics that let hidden reasoning persist between requests. It also shows why attribution needs care: OpenAI reports attempted extraction and a Moonshot-linked core cluster, not a confirmed success rate or a finding about Moonshot's leadership.

Branching filaments remain inside a sealed model enclosure as several fine probes press against its surface.

OpenAI says people associated with Moonshot AI were behind a core cluster of attempts to extract hidden reasoning from its models, using encrypted reasoning blocks as inputs to other model conversations. The activity ran from July 1st through July 28th; OpenAI disclosed it on September 30th. OpenAI says it disrupted the campaign and closed a replay pathway, while acknowledging it cannot attribute every operator to one actor or establish that every attempt worked. Tom's Hardware reported the disclosure on October 1st.

Moonshot's co-founder and CEO Yang Zhilin previously worked as an AI researcher at Google Brain and Meta AI, and founded the company in 2023, TechCrunch reported. Moonshot's Kimi models compete with OpenAI's products, as TechCrunch reported. Moonshot's own account of its founding says the company was driven by the pursuit of AGI. That competitive context makes OpenAI's attribution consequential. It does not establish that Yang directed, knew about, or participated in the activity; OpenAI's post attributes a core cluster to individuals associated with Moonshot, not to the founder personally.

The numbers describe attempts

OpenAI recorded 16,000 requests using an extraction pattern across more than 4,000 users during spikes on July 24th and 25th. OpenAI also identified related prompt-pattern activity involving a broader cluster of more than 15,000 users, which it says it disrupted by July 28th. OpenAI's footnote is important: the figures count attempted extractions, not confirmed successful ones. OpenAI did not identify the models targeted or say how many of the users it linked to Moonshot.

The method OpenAI described took advantage of how reasoning data moves through some model APIs. A model can return an encrypted block representing its hidden reasoning; a client passes that block back in later requests, so the provider need not retain it as stored conversation history. OpenAI says operators copied such a block from one conversation into another and prompted a model to decrypt and transcribe it. The encryption itself was not broken, OpenAI said, and the activity did not involve direct access to stored user conversations or a database compromise.

OpenAI says it closed a path that could let someone who already had another user's encrypted reasoning replay it and recover its contents. OpenAI also added checks to hold streamed output that might expose reasoning, strengthened protections across users and organizations, and worked with third-party providers to disrupt accounts using their services.

A weakness researchers found across providers

In an August 10th paper, researchers reported that they could feed encrypted reasoning from a stronger model into a weaker model from the same provider and get the latter to reproduce the trace in plain text. Their tests covered OpenAI, Anthropic, and Google systems. OpenAI says researchers separately disclosed related cross-model and conversation-compaction vulnerabilities and that it confirmed the attack paths were real. Tom's Hardware reported that the researchers could not launch the same attacks after providers acknowledged their reports.

That work describes a design tension for model providers: reasoning blocks need to remain usable across successive API calls, but that portability can create opportunities for replay across conversations or models. OpenAI's response now includes both technical controls and account enforcement. OpenAI also says partner-hosted deployments need comparable protections, and that tool-output attacks require further defenses.

Anthropic's September report describes a separate investigation into Moonshot. Anthropic says Moonshot routed almost 300,000 customer requests to Claude over one ten-day period through a network of 5,380 fraudulent accounts, and saved Claude's reasoning signatures for use in new sessions to recover reasoning transcripts. Those are Anthropic's findings about its own service, not independent confirmation of OpenAI's July attribution. Taken together, the two labs' disclosures show how providers are treating model outputs and hidden reasoning as valuable training material that can be targeted through API access.

Moonshot's financial trajectory gives that contest added context. A May 7th TechCrunch report said Moonshot raised about $2 billion at a $20 billion valuation, with Long-Z Investments leading and Tsinghua Capital, China Mobile, and CPE Yuanfeng participating. The report also said Moonshot's annual recurring revenue topped $200 million in April, citing a financial adviser. Those figures describe the commercial scale of the race around Kimi; they do not establish a motive for the alleged extraction campaign.

For Yang and Moonshot, the allegation lands as Moonshot competes on models that have drawn substantial investor backing and demand. OpenAI says it shared findings through the Frontier Model Forum and government information-sharing channels, and expects extraction attempts to grow more sophisticated as models improve. OpenAI says protections for partner-hosted deployments and tool outputs remain part of the work.

Reader comments

Conversation for this story loads after sign-in.