Cognition puts Nvidia's Vera Rubin to work, reports up to 4.8x token throughput

Cognition, Scott Wu's coding-agent company, ran its SWE-2 workload on CoreWeave's first production Vera Rubin cluster, comparing it with GB200 in a vendor-led test.

By · Published

Primary source: Cognition

Why it matters

Cognition is tying its coding-agent strategy to the compute needed for long, multi-step tasks. Its 4.8x figure is an early, company-run throughput result, not proof of cheaper inference or more completed engineering work.

Cognition puts Nvidia's Vera Rubin to work, reports up to 4.8x token throughput — Scott Wu's coding-agent company ran its SWE-2 workload on CoreWeave's first production Vera Rubin cluster, comparing it with GB200 in a vendor-led test.

Cognition, the company founded by Scott Wu to build an AI software engineer, says its SWE-2 model produced up to 4.8 times more total token throughput on Nvidia's Vera Rubin NVL72 than on a GB200 NVL72 system. The test ran on CoreWeave, which identified Cognition as the first customer running production workloads on Vera Rubin. The announcement came on September 30th.

https://x.com/cognition/status/2105408461701824732

poster=/api/storage/public-objects/video-posters/29af8c91a06ddd14bb28dde81c9ff97fa8e95593ef5640599d90611ef7f725f4.webp|Video from @cognition on X
Video from the original post on X.

Wu's route to the infrastructure question began with competitive programming, not data centers. He won three International Olympiad in Informatics gold medals, worked on software performance optimization at Addepar, attended Harvard for about two years, then left to co-found Lunchclub. In an interview with Colossus, Wu described how he and former collaborators were exploring AI after ChatGPT arrived in November 2022. Their interest in coding, he said, made software engineering a natural place to focus.

That focus became Devin, Cognition's autonomous software engineer, which takes on multi-step coding tasks in a team's existing environment. Cognition's co-founders Steven Hao and Walden Yan also came from competitive programming; the group is betting that engineers will increasingly set goals while agents carry out longer sequences of work. Faster inference capacity goes directly to that bet: an agent can make repeated model calls as it plans, edits, tests and revises code.

The benchmark is about throughput, not completed work

Cognition says its engineers compared Vera Rubin and GB200 using SWE-2 inference workloads built from a sample of tasks in its FrontierCode benchmark. The Nvidia account describes the workload as real-world software-engineering tasks. Cognition reported up to 4.8 times higher total token throughput, while CoreWeave said the improvement came with no loss in generation speed. CoreWeave's release also reported 3.8 times higher output-token throughput on reinforcement-learning workloads.

Those are vendor-associated results, not an independent replication. CoreWeave says Cognition's engineers ran the benchmark against a GB200 NVL72 baseline. The published materials do not provide absolute throughput figures, system pricing, energy use or enough configuration detail to reproduce the comparison. The 4.8x result is an upper-end throughput claim for this SWE-2 test; it does not establish a 4.8x increase in completed coding tasks, nor does it show lower total cost across workloads.

That qualification matters for Cognition's business model. An agent that runs through many steps can consume far more inference than a single prompt-and-answer exchange. More tokens per unit of time could allow Cognition to run more sessions on a given deployment, shorten some agent cycles or use available compute more intensively. Which benefit dominates depends on how the capacity is priced and on whether the agent's work remains useful as throughput rises. The published comparison does not answer those questions.

A new cluster for a fast-growing workload

Cognition worked with CoreWeave to stand up its Vera Rubin cluster in early September, according to the cloud provider. CoreWeave said Cognition had expanded to thousands of GPUs on its platform in less than nine months, using it for Devin training, reinforcement learning and production inference. That makes the announcement a customer deployment as well as a chip comparison: the new hardware is being tested inside a system Cognition already uses to build and serve its agent.

The rollout follows Cognition's September 8th announcement of a Series E that it said raised more than $2 billion at a $48 billion valuation. Andreessen Horowitz and Accel led the round, alongside existing investors Founders Fund, General Catalyst and Avenir, according to Cognition's announcement. Cognition said then that run-rate revenue had grown from $492 million in May to almost $900 million. That is company-reported revenue, and its announcement did not define the calculation in detail.

RuntimeWire previously reported on Cognition's SWE-2 launch, where the company compared the model's coding performance and rollout cost with a rival model. The Vera Rubin test extends that same operating question to the hardware layer: how much agent activity can Cognition serve as it pushes software work from short code suggestions toward longer autonomous tasks?

Nvidia and CoreWeave have an incentive to spotlight an early customer benchmark as Vera Rubin enters production. Cognition has its own reason to publish the result: it is showing investors and customers that infrastructure capacity is part of how it plans to scale Devin. The throughput number is a useful early signal for a workload built around repeated model calls. Establishing what it means for cost, code quality and task completion will require metrics this announcement does not provide.

Reader comments

Conversation for this story loads after sign-in.