Fal launches H3 Max Director for live, steerable AI video

Fal's H3 Max Director turns short generated clips into a WebRTC session that developers can steer with live prompts. The API supports 24 fps output and remembers up to 50 earlier prompt segments. Fal launched it with a [75% generation discount](https://x.com/fal/status/2095599871449342288) for the first two weeks.

By · Published

Primary source: Fal on X

Why it matters

Director moves generative video from exported clips toward live application infrastructure. Its alpha client, 768p ceiling and bounded 50-segment memory keep the initial market experimental.

Fal launches H3 Max Director for AI video that keeps taking direction

Burkay Gur and Gorkem Yurtseven's generative-media infrastructure company Fal announced H3 Max Director on September 3rd, giving developers an API for generating a running video-and-audio stream that can change course as new prompts arrive. Fal laid out the product in its launch thread on X.

Fal launch thread on X

Gur previously led machine-learning development at Coinbase after working at Oracle, while Yurtseven spent seven years as a software engineer at Amazon. The longtime friends began exploring startup ideas during the pandemic and founded Fal in 2021 around AI infrastructure, TechCrunch reported.

Their early thesis was broad. Gur and Yurtseven narrowed Fal's focus to image, video and audio models as demand for generative media grew. H3 Max Director shows how that focus is moving Fal beyond hosting other labs' models. Fal post-trained MiniMax's H3 model, optimized the serving stack around it and has now packaged that work as an interactive runtime.

From faster clips to a running session

Fal released the underlying H3 Max model on August 27th. In its H3 Max launch article, Fal said it generated a five-second video in under three seconds and described its inference team as having spent four years optimizing diffusion and generative-media workloads. Those are Fal's own performance and engineering claims. Director applies that work to a persistent WebRTC session where developers can steer the output while it runs.

Developers open H3 Max Director through WebRTC, receive video and audio tracks, and send new prompts while the session is active. According to the API documentation, Fal offers 480p and 768p output at 24 frames per second in landscape, portrait and square formats, with five- to 15-second chunks and memory for as many as 50 earlier prompt segments.

That architecture gives developers something closer to a stateful video application. A viewer could vote on what a character does next, a game could generate a setting around player choices, or a virtual presenter could move through a program without waiting for a conventional render queue after every instruction.

Fal is demonstrating the model on fal.live, an experimental collection of AI-generated channels spanning anime, sitcoms, fantasy, cooking and horror. The launch thread says Director is powering several continuous channels, with audience input influencing the scenes that follow. Fal lists interactive video, persistent worlds and continuous livestreams among the intended uses.

The word "persistent" needs a practical qualifier. Director carries characters, settings and prompts forward inside a session, while the documented maximum of 50 remembered prompt segments puts a boundary around that memory. Developers building a world expected to survive across long sessions, returning users or multiple devices will still need their own state-management layer.

The founders' infrastructure bet reaches the screen

Gur and Yurtseven focused Fal on generative media because images and video created a distinct infrastructure problem: heavy GPU workloads, fast-changing model architectures and applications where latency directly shapes the user experience.

That choice is visible in the product design. The Director API documentation exposes buffering, chunk timing and prompt-version controls rather than hiding Director behind a consumer video editor. It labels fal.realtime.open an experimental alpha API that "may change in a minor version." Director is aimed first at developers willing to build the surrounding product themselves.

Fal has the financing to test whether that developer audience can produce a new application category. In December 2025, Fal raised a $140 million Series D led by Sequoia, with Kleiner Perkins and Nvidia joining existing investors. Fal said the round would fund global infrastructure and new generative-media capabilities. Director is the clearest expression yet of where some of that capital is going: research and inference engineering developed together, then sold through Fal's existing API distribution.

The founders are also moving Fal into a more exposed position. Serving outside models lets an infrastructure provider benefit regardless of which model wins. H3 Max and Director ask customers to choose Fal's own post-training work and trust Fal to maintain the stream, preserve continuity and keep latency low under sustained demand.

The alpha designation and bounded prompt memory define the work ahead for Gur and Yurtseven. Director must remain stable over longer sessions and consistent enough that audience choices feel immediate rather than queued between generated clips.

Fal has cleared an initial technical hurdle: developers can steer a generated audiovisual scene without ending the session and starting another render. That changes the unit of creation from a clip into a running program. Gur and Yurtseven are betting that developers will find businesses inside that difference.

Reader comments

Conversation for this story loads after sign-in.