Runway says Solaris generates interfaces frame by frame in real time

The early-access research model turns clicks and drags into synthesized 720p frames, while text, trust and accessibility remain open problems.

By · Published

Primary source: Runway on X

Why it matters

Solaris extends generative video into software execution, giving Runway a route beyond media tools. Its core obstacles are the basic requirements conventional interfaces already satisfy.

Runway says Solaris generates interfaces frame by frame in real time

Runway unveiled Solaris on August 31st, describing the research model as an operating layer that generates interactive interfaces frame by frame instead of rendering a conventional app or website from code. In a post on X, Runway called Solaris its first "Interface World Model."

https://x.com/runwayml/status/2094463070466646019

poster=/api/storage/public-objects/tweet-videos/runway-solaris-interface-world-model-real-time-poster-0d490532.jpg|Video from @runwayml on X

The New York AI lab was founded in 2018 by Cristobal Valenzuela (@c_valenzuelab), Anastasis Germanidis (@agermanidis) and Alejandro Matamala Ortiz (@matamalaortiz), who met at New York University's Interactive Telecommunications Program. Runway began with machine-learning tools for artists. Solaris carries that visual-first thesis into software itself: the interface becomes a continuously generated image that changes as a user clicks, drags or types.

Solaris remains a research preview rather than a publicly available operating system. Runway is accepting early-access requests and says it is working with partners toward a public launch. Runway has not attached a release date, API, pricing structure or hardware requirements to the announcement.

How Solaris generates an interface

Solaris builds on Runway's Gen-4.5 video model and follows the broader world-model research program that Germanidis introduced in December 2023. Runway adapted the video model to treat mouse and keyboard activity as conditioning data for each subsequent frame.

Runway says it first trained Solaris to generate frames autoregressively, meaning each frame depends on those already produced. Runway then distilled the video model's multi-step denoising process into fewer steps and trained the faster model on its own outputs to improve stability during longer interactions. The resulting system targets interactive response times and 720p visual quality, although Runway did not publish a latency figure.

A separate language model handles reasoning. It interprets a user's request, decides whether the interface should alter the current scene or move to another one, and sends instructions to Solaris for rendering. Text prompts also define what particular clicks or drags mean inside a scene.

That architecture complicates Runway's claim that Solaris uses "no code." Developers do not have to construct each screen and interaction through a conventional interface framework, according to Runway, but the system still relies on prompts, a language model and predefined starting material to determine what happens.

Runway demonstrated potential uses including virtual stores where shoppers drag clothes onto an image of themselves, product scenes that respond to spoken editing requests, generated tutorials and simulations that react to direct manipulation. Runway also argues that constantly changing interfaces could provide training environments for computer-using AI agents, reducing their dependence on memorized website layouts.

Runway's benchmark tests a narrow advantage

Runway compared Solaris with a coded interface generated by Claude Opus 5 using the same starting image and interaction request. In a Runway-run study involving 250 participants, 30 examples and nearly 7,500 pairwise judgments, participants preferred Solaris for following instructions in 61% of comparisons, compared with 24% for the coded result. Solaris received 71% of preferences for natural behavior, versus 21% for the coded interface. Runway classified most remaining judgments as equivalent.

A separate 30-interface reconstruction test examined how multimodal models including GPT-4o, Gemini 2.5 Pro and Claude Fable 5 reproduced interfaces from screenshots. Runway measured visual preservation using structural similarity and DINOv3 features, finding that reconstruction quality declined as the source images became more visually complex.

Those tests support a specific argument: a generated visual scene can preserve motion and appearance better than screenshot-to-code reconstruction on Runway's examples. They do not establish Solaris as a replacement for conventional software across reliability, cost or accessibility.

The operating-system pitch meets operating-system requirements

Runway acknowledges that Solaris struggles with stable, legible text, extended-session coherence and grounding responses in verified information. Those limitations cut directly into the commercial uses Runway proposes. Stores need accurate product details, tutorials need reliable instructions, and most productivity software depends on dense text.

Accessibility presents another constraint. A synthesized frame does not automatically expose the document structure, labels and controls used by screen readers and other assistive technologies. Runway says generated interfaces will need integration with accessibility APIs and the wider software stack.

Runway also says generating each frame remains more expensive than serving a page built once, even after lowering inference costs relative to a standard video diffusion model.

Solaris gives Runway a product-shaped result from a world-model strategy it has pursued for nearly three years. In February, Runway raised a $315 million Series E led by General Atlantic, with NVIDIA, Adobe Ventures, AllianceBernstein, AMD Ventures, Fidelity Management & Research, Mirae Asset, Emphatic Capital, Felicis and Premji Invest participating. Runway said the capital would fund the pre-training of world models and their expansion into new industries.

Solaris is the clearest expression of that expansion so far. Runway is taking technology developed for generating video and positioning it as the runtime for software. Whether that becomes an operating layer will depend on less cinematic work: readable text, predictable behavior, affordable inference and interfaces that every user can navigate.

Reader comments

Conversation for this story loads after sign-in.