Gimlet and Cerebras plan 100 MW of inference capacity

In announcements published September 28th, Gimlet Labs and Cerebras said their integrated system was already serving inference traffic in private deployments. The first Cerebras-powered Gimlet Cloud data center is expected later in 2026, with customer access to Cerebras' CS-4 system planned for 2027.

By · Published

Primary source: Cerebras Newsroom

Why it matters

Gimlet is moving its multi-chip inference thesis toward a planned 100 MW deployment of Cerebras-powered capacity. Whether that architecture can deliver measurable speed and efficiency at production scale remains the key test.

A detailed view of advanced computing hardware, featuring a Cerebras wafer-scale engine integrated alongside multiple graphics processing units within a high-tech server cabinet.

On September 28th, 2026, Gimlet Labs co-founder and CEO Zain Asgar and Cerebras announced a collaboration to bring Cerebras wafer-scale compute into Gimlet Cloud, Gimlet's inference service. In its September 28th announcement, Gimlet said the companies plan to deploy 100 megawatts of Cerebras-powered inference capacity; Cerebras' same-day release also confirmed the collaboration and said an integrated solution was already serving tokens in private deployments. The first Cerebras-powered Gimlet Cloud data center is expected later in 2026. Cerebras named Gimlet a launch partner for its CS-4 system, with Gimlet Cloud customers expected to get access in 2027. The companies describe speeds of up to 3,000 tokens per second as a target, not a published benchmark result.

The partnership puts Asgar's central argument about AI infrastructure into a larger operational test: inference workloads have different computing needs at different stages, so a cloud built around one kind of accelerator can leave performance or capacity on the table. Gimlet says its software will assign those stages to the silicon best suited to each. Cerebras brings wafer-scale systems; Gimlet contributes orchestration, infrastructure and the developer-facing cloud.

That thesis has followed Asgar across two startups. Before founding Gimlet, he co-founded Pixie Labs, which built Kubernetes observability software. New Relic announced an agreement to acquire Pixie in 2020, describing its product as a way to collect telemetry and debug Kubernetes applications without manually adding code instrumentation. Asgar said then that joining New Relic would help Pixie scale its platform. At Gimlet, the challenge is different: coordinate computing hardware across a data center to serve models quickly and at scale.

From orchestration software to infrastructure

Gimlet's September 28th announcement describes a vertically integrated platform spanning data centers, systems software, workload orchestration and developer APIs. The software can split inference into phases such as prefill, decode, attention and feed-forward processing, then schedule them across GPUs and specialized accelerators. Developers are meant to use a standard inference API rather than manage the underlying hardware themselves.

Diagram of Gimlet’s standard inference API, workload orchestration, inference phases, and GPU and specialized-accelerator scheduling.
Gimlet says its software can split inference into phases and schedule them across GPUs and specialized accelerators — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

Cerebras is a notable addition because its wafer-scale systems are built differently from conventional GPU clusters. Gimlet says the Wafer Scale Engine's on-chip memory and bandwidth suit memory-intensive inference, including token generation. The intended design pairs that capability with GPUs, which Cerebras co-founder and CTO Sean Lie described as delivering high throughput. Gimlet and Cerebras plan to combine the architectures within one workload.

Gimlet and Cerebras say their collaboration grows out of customer work underway since 2025 and an integrated system already serving tokens in private deployments. The Cerebras release also names Gimlet as a launch partner for Cerebras' CS-4 system. Gimlet says customers should get access to that next-generation technology in 2027.

Asgar said the combined system is planned to reach speeds of up to 3,000 tokens per second. That is a company-stated target; the announcement provides no benchmark result. The materials do not identify the model, prompt length, batch size, concurrency or measurement method behind the figure. Tokens per second also needs a denominator to make a useful comparison: the claim does not say whether it describes output for one user, an entire system or another configuration.

Gimlet's blog post says Gimlet and Cerebras plan 100 megawatts of Cerebras-powered inference capacity. That is a capacity target; the sources do not say a data center of that size is already operating. Gimlet and Cerebras expect the first Cerebras-powered Gimlet Cloud data center to come online later in 2026. The announcement does not specify its location or initial capacity.

A larger bet after a large round

Gimlet has already raised substantial capital to pursue the infrastructure thesis. Its March 23rd Series A announcement disclosed an $80 million round led by Menlo Ventures. Gimlet announced a $300 million Series B on September 4th, led by Andreessen Horowitz. In that Series B post, Gimlet said it had added billions in contracted revenue and gigawatts of data-center pipeline since March, and was scaling toward hundreds of megawatts in managed capacity. Those rounds are separate from the Cerebras collaboration, which Gimlet and Cerebras did not describe as a financing deal. The funding announcements show investors backing Gimlet's effort to build and operate a hardware-diverse cloud; the Cerebras partnership adds a named accelerator supplier and a deployment plan to that effort.

Gimlet emerged publicly in October 2025 with a pitch for making AI workloads more efficient by distributing work across available hardware. Its founding post described an orchestrator that breaks agent workloads into compute fragments, a hardware-agnostic compiler and autonomous kernel generation. Gimlet's team page lists Asgar as CEO alongside co-founders Michelle Nguyen, Omid Azizi, Natalie Serrino and James Bartlett. That range of systems, compiler and hardware work is relevant to the deal: the product depends on software translating specialized chips into a cloud developers can use without needing to manage each system themselves.

The commercial test will be whether the combination can deliver its promised responsiveness while serving enough simultaneous workloads to justify the infrastructure. The announcements emphasize real-time applications and agents, where delays can accumulate across sequential model calls. They give no prices, utilization, customer names for the private deployments or public comparison against GPU-only systems under matched conditions. Those details will determine whether the 3,000-token figure and the 100-megawatt plan describe a competitive service, rather than two large targets attached to a partnership announcement.

Reader comments

Conversation for this story loads after sign-in.