Nebius buys Inferize to add GPU snapshots to Token Factory
Nebius acquired Inferize on October 1st, adding its inference-snapshot technology to Token Factory. Terms were undisclosed; CTech and Globes estimated $100M-$150M and $100M-$130M, respectively.
By RuntimeWire Staff · Published
Primary source: Ctech
Why it matters
Nebius is adding a way to bring inference capacity online without keeping as many GPUs idle between demand spikes. The deal's practical value will depend on whether Inferize's company-reported speed and cost gains hold up in production.

Nebius acquired Inferize on October 1st, bringing co-founder and CEO Guy Bortnikov and Inferize's technology and engineering team into Nebius Token Factory, Nebius's managed inference platform. CTech published its report that day, and Globes reported the deal later the same day. CTech reported that Inferize had 17 employees in Tel Aviv and estimated the deal at $100M-$150M; Globes estimated $100M-$130M. Nebius said the terms were not disclosed.
Inferize's system captures a warmed engine's CPU and GPU state, including compiled kernels and CUDA graphs, then restores it on GPUs in seconds, according to Inferize. The snapshot is taken after the model is serving. Inferize says its software works below the serving engine, without requiring changes to the engine or application. The intended uses include scaling replicas up when traffic rises, releasing capacity when it falls, and recovering serving capacity after a spot or preemptible node is reclaimed.

Those functions target a familiar infrastructure trade-off. Operators can pay to keep spare GPUs running so they can handle sudden demand, or scale closer to actual demand and risk waiting through model loading and engine initialization when traffic spikes. Nebius says cold starts also leave GPUs idle during launches and model-weight updates, including updates during reinforcement-learning workloads. Inferize co-founder and CEO Guy Bortnikov told CTech, "Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform."
Nebius says Inferize was founded in January 2026 and had a working prototype within three months. CTech reported that Inferize had 17 employees in Tel Aviv and raised a seed round led by TLV Partners and angel investors. Globes reported that TLV Partners was the company's sole investor. Inferize's site offers benchmark results and a way to request a demo.
Inferize's founders had previously worked at Granulate, an Israeli infrastructure-optimization company acquired by Intel in 2022. TechCrunch reported that the Granulate deal was reportedly worth as much as $650 million; Intel did not disclose its price. TLV Partners had also backed Granulate, according to its portfolio, a connection CTech cited in its account of Inferize's financing.
For Nebius, Inferize adds a runtime and capacity-management layer to Token Factory. Nebius launched the platform in November 2025 to deploy and optimize open-source and custom models in production. Nebius says Eigen AI contributes model, kernel and system optimization, while Clarifai's core team and licensed technology add system-level inference and compute orchestration. Nebius says Inferize's engineers will work across Token Factory, starting with the integration of the snapshot technology.

The approach sits alongside other efforts to shorten inference startup. Modal's engineering account describes combining cloud buffers, a custom filesystem and CPU- and CUDA-level checkpoint and restore; Modal reported reducing a sample inference-server startup from about 2,000 seconds to 50 seconds. vLLM's initialized engine snapshots are an experimental feature documented for a single GPU on the same machine, with specific environment and configuration requirements. Inferize describes restoring a complete multi-GPU serving engine onto new GPUs. These systems address related delays through different designs, so their reported results are not directly comparable.
Inferize's site claims GPU-cost reductions of 10% to 30%. In one Inferize-published DeepSeek V4 Pro benchmark, a snapshot starts in 20 seconds, compared with a cold start of about 23 minutes, a roughly 70-times speedup. Those figures are Inferize's benchmarks, not independent measurements. Nebius disclosed no acquisition terms; CTech estimated the deal at $100M-$150M, while Globes estimated $100M-$130M.