Wafer claims a 12T monthly token run rate and keeps hiring engineers
The San Francisco inference startup is offering up to $300K for engineers who will own kernels, clusters and customer workloads.
By Ryan Merket · Published
Primary source: X
Why it matters
Wafer's claimed traffic growth gives its $40 million Series A a concrete use: hiring scarce performance engineers while trying to automate enough of their work to protect margins.

Emilio Andere (@gpuemi) and Steven Arellano (@gpusteve)'s Wafer says its inference platform has reached a 12 trillion-token monthly run rate, a company-supplied milestone that arrives three weeks after the startup raised a $40 million Series A.
Eugene Ye (@gpugene), who works on operations and compute at Wafer, disclosed the figure in a September 21st post on X. His post paired the number with another recruiting pitch for members of technical staff. Ye joined Wafer after leading operations at Mercor, according to Wafer's team page. He studied applied math, computer science and economics at Harvard while pursuing classical music, a field in which he was named to a Canadian 30-under-30 list.
The chart attached to Ye's post shows Wafer's reported token volume accelerating sharply between June 1st and September 21st. It provides no intermediate values, customer breakdown or methodology. At the stated run rate, Wafer would be processing roughly 400 billion tokens a day over a 30-day month.
Token volume is a useful measure of infrastructure demand, though it leaves the economics unanswered. The 12 trillion figure does not disclose how much traffic comes from paid production customers, how much capacity those workloads consume, or the revenue and gross margin attached to each token. Wafer has not paired the chart with updated revenue, customer or utilization figures.
The hiring plan shows where Wafer expects that volume to create pressure. Its member of technical staff opening carries a base salary of $200,000 to $300,000 plus equity. The role is based in San Francisco and requires five days a week in the office. Wafer also offers a post-tax housing stipend of $1,000 a month to employees who live within half a mile of the office.
The job extends well beyond writing GPU kernels. Wafer expects each engineer to optimize inference engines, operate clusters spanning multiple chip vendors, develop autonomous performance-engineering agents and take direct responsibility for customer accounts. Wafer says there is no separate solutions organization between the engineer and the workload. The same person who wins a technical evaluation may also diagnose latency changes, keep the deployment running and explain the result to the customer's engineering team.
That structure reflects the difficult part of selling optimized inference. Benchmark gains have to survive production traffic, changing model architectures and different accelerators. Wafer's pitch is that software agents can automate work usually handled by a scarce group of engineers who understand kernels, compilers, runtimes and serving systems. Its technical staff still have to prove those improvements inside customer deployments.
Andere and Arellano developed that thesis while studying at the University of Chicago, where they met as freshmen and later became roommates. Andere studied mathematics, researched weather models at Argonne National Laboratory and published machine-learning security work. Arellano studied computer science and previously worked on high-performance computing and AI infrastructure at Two Sigma, Google and Sei Labs.
In Wafer's April 14th seed announcement, Andere wrote that increasingly diverse AI hardware was creating an optimization problem too large and hardware-specific for manual engineering to keep pace. Wafer initially raised $4 million, led by Fifty Years, to build an agent that performs that work across accelerator architectures.
On September 1st, Wafer followed with a $40 million Series A co-led by Marathon and Chemistry. Wing, AMD Ventures and Outset Capital participated, while existing backers Fifty Years and Y Combinator invested again. Wafer said the capital would fund further automation of the inference optimization loop.
Wafer currently lists four openings: member of technical staff, founding growth, go-to-market and chief of staff. The concentration of roles around engineering, customer acquisition and operations gives the financing a clear immediate purpose. Wafer has to add enough technical capacity to support the traffic represented by its chart while continuing to lower the human effort required for each deployment.
The 12 trillion-token claim offers evidence that Wafer is finding workloads for its infrastructure. The next test sits inside the job description: whether a small group of engineers and Wafer's optimization agents can turn that volume into repeatable deployments without recreating the labor-heavy consulting model the founders set out to automate.