Rysana says its V2 AI runs up to 10,000 times faster
Founder John is pitching a multimodal model family around inference speed and structured outputs; Rysana's comparisons use its own benchmark and mid-August API measurements.
By Ryan Merket · Published
Primary source: X
Why it matters
Rysana is selling inference efficiency as a way to make AI products faster and cheaper to run. Its 10,000-times claim is company-reported, and V2 access is gated, so customers will need to test the model on their own tasks before the headline multiplier has practical meaning.

Rysana says its V2 model family can run up to 10,000 times faster and use 1,000 times less energy than conventional AI models, a claim that puts inference cost and latency at the center of the startup's pitch. The company posted the announcement on September 29th, but the measurements it cites on its website compare a July 2026 V2 checkpoint with other models measured in mid-August.
https://x.com/Rysana/status/2104950406589898811
Rysana is led by CEO John (@jrysana), a technically oriented founder whose earlier public work focused on structured language-model outputs. In a Latent Space podcast discussion, Instructor founder Jason Liu referred to John as part of the University of Waterloo's machine-learning community. That is a connection, not evidence of a particular degree or role at the university.
The company describes V2 as a multimodal model family for text, audio, images, code and documents. Rysana also emphasizes structured outputs, including schemas, grammars and programs, and says developers can control sampling and trade off speed, cost and effort. Those features point toward applications where a model must respond quickly and return machine-usable results, rather than simply generate long conversational answers.
The performance numbers remain Rysana's claims. Its site says the comparison covers more than 20,000 varied real-world tasks without chain-of-thought, and plots task scores against peak end-to-end tokens per second per request. It compares V2 with GPT, Claude, GLM and Llama, using what Rysana calls the fastest available conventional API for each. That setup is useful as the company's chosen comparison, but it is not an independent evaluation: the site does not identify the individual models and API configurations in the comparison text, or explain enough about the test mix to show how the result transfers to a particular customer's workload.
The headline multiplier also needs a denominator. "Up to" describes a best-case result, while speed and efficiency can vary with the task, output format, hardware and service configuration. Rysana's website says V2 can generate millions of tokens per second and reach sub-millisecond latency in some configurations. Those are company-reported capabilities, not independently verified results. The site describes a model family, but the announcement and benchmark framing do not establish that every supported modality or task reaches the maximum speed.
Access is another open question for prospective users. Rysana's site says it is opening access and asks interested customers to request it; the login page says accounts are invite-only. A previously indexed pricing page lists a $25 monthly plan and usage charges for Inversion models, but that page describes the earlier product, so those figures should not be read as V2 pricing. The V2 page's FAQ asks what pricing Rysana offers without displaying a price in the captured page.
Rysana's approach extends a thesis it has pursued through Inversion, its earlier family of models for structured tasks. John previously described Inversion as focused on speed, latency, reliability and typed JSON output. V2 widens the pitch from structured generation toward general, multimodal AI while retaining the same emphasis on controlled outputs and fast execution. The shift matters commercially because inference speed can reduce waiting time and computing expense in high-volume products, if customers can reproduce the gains on their own workloads.
The public materials leave the practical test to customers: whether V2 maintains useful task quality at the claimed operating speeds, and whether access, deployment requirements and pricing make those gains usable in production. Rysana points to its benchmark and invites customers building for large volumes or low latency to seek access. Until those results are tested outside the company's own comparison, the 10,000-times figure is best read as the upper bound of Rysana's pitch, not a general performance guarantee.