Investor-backed AI startups are reselling frontier-model traces, Nomic CEO alleges
Mulyar's screenshots made the accused founders and startups identifiable; RuntimeWire cropped their names because the thread offered no transaction records or other corroborating evidence.
By Ryan Merket · Published
Primary source: X - Andriy Mulyar
Why it matters
Startup credits and priority rate limits are strategic distribution tools for AI labs. If recipients resell the resulting traces, those programs can subsidize rival-model training.

Andriy Mulyar (@andriy_mulyar), founder and CEO of Nomic, alleged on September 14th that startups with ties to OpenAI, Anthropic and Y Combinator are selling valuable frontier-model traces to unidentified buyers, including companies from China.

Mulyar made the accusation in a four-post thread on X, writing that he had heard of "so many companies" using elevated model privileges to generate and resell traces during the past month. He said founders of companies that build reinforcement-learning environments were receiving substantial inbound interest from Chinese buyers.
The screenshots attached to the thread made the accused founders and startups readily identifiable. RuntimeWire cropped their names from the image above because the allegations are unverified. Mulyar provided no transaction records, datasets or buyer identities, so the thread does not establish that any OpenAI-, Anthropic- or Y Combinator-backed startup sold model output. OpenAI, Anthropic and Y Combinator were cited because of their alleged connections to the startups, rather than accused of authorizing the sales.
Mulyar specifically invoked Kimi, the model family developed by Moonshot AI, and Z.ai's GLM models. His argument was direct: access intended to help portfolio companies build products may also give intermediaries a cheap supply of reasoning data that rival labs can use for training.
Startup access creates an obvious resale target
OpenAI and Anthropic openly offer venture-backed startups more favorable access than an ordinary API customer may receive. The OpenAI for Startups program advertises free API credits, rate-limit upgrades and technical support for eligible companies in its investor network. Anthropic's startup program offers credits and priority rate limits, with additional benefits available to startups backed by partner investors and accelerators.
Those programs are designed to recruit developers and help young companies absorb the cost of building on frontier models. They also create accounts capable of producing larger volumes of output at lower effective prices. A startup that converts subsidized access into training data for an outside buyer would be monetizing the lab's distribution budget against the lab itself.
The traces at issue are records of a model working through tasks, using tools, grading answers or producing step-by-step reasoning. They are particularly useful for distillation, where outputs from a stronger model are used to train or improve another model. Environment builders sit close to that process because they create the software tasks, simulations and scoring systems used to train agents through reinforcement learning.
An Epoch AI survey of the RL environment market described those environments as a central input for frontier-model training, spanning coding, computer use and enterprise workflows. The work can generate large collections of prompts, responses, tool calls and reward signals. Those records can become training material if retained and transferred.
OpenAI and Anthropic have documented the wider pipeline
Mulyar's allegation lands after both major US labs described organized efforts to acquire their outputs through intermediaries.
On February 23rd, Anthropic said it had identified distillation campaigns attributed to DeepSeek, Moonshot AI and MiniMax. According to Anthropic, the three labs generated more than 16 million exchanges with Claude through approximately 24,000 fraudulent accounts.
Anthropic said proxy services assembled networks of accounts, mixed distillation traffic with ordinary customer requests and replaced accounts as they were banned. Moonshot allegedly generated more than 3.4 million exchanges focused on agentic reasoning, coding, computer use and vision, including attempts to reconstruct Claude's reasoning traces. Anthropic attributed that activity using request metadata, infrastructure indicators and information from industry partners.
OpenAI described a similar market in a February 12th submission to a US House committee. OpenAI said Chinese companies had used unauthorized resellers and third-party routers to conceal the origin of requests, and that some usage patterns were consistent with developing competing models through distillation.
OpenAI's business agreement prohibits customers from transferring API keys and, outside permitted exceptions, using output to develop competing AI models. Anthropic likewise says customers may not use Claude outputs to train models that compete with its own.
Those disclosures establish that a market for concealed frontier-model access exists. They do not corroborate Mulyar's narrower allegation that legitimate startups backed by the labs or Y Combinator are supplying it. The distinction matters because the enforcement problem changes if approved portfolio accounts, rather than fabricated identities and stolen credentials, are generating the data.
Mulyar has built with distilled data before
Mulyar speaks from direct experience with model distillation. He was part of the Nomic team that created GPT4All, an open-source project for running language models locally. The project's 2023 paper said its original training dataset contained roughly one million prompt-response pairs collected through OpenAI's GPT-3.5 Turbo API over six days.
GPT4All's repository described the project as an assistant model trained through "large scale data distillation from GPT-3.5-Turbo." Nomic later expanded into embedding models, data tooling and AI agents for architecture, engineering and construction. Mulyar remains Nomic's CEO.
That history gives Mulyar firsthand knowledge of how quickly API output can become a training dataset. It does not supply evidence for the transactions alleged in his thread.
The commercial incentive is nevertheless clear. Frontier traces are expensive to generate, valuable to model developers and difficult to source directly when a lab blocks access by region or intended use. Startups with credits, high rate limits and trusted accounts occupy a useful position between the model provider and buyers seeking data at scale.
If portfolio companies are exploiting that position, startup access programs have become a supply-chain vulnerability for the labs funding them. The immediate test for OpenAI and Anthropic is whether monitoring built to catch fraudulent account clusters can also detect approved customers whose legitimate workloads are being redirected into a gray market for model training data.