Fal brings H3 Max reference-to-video to real time with four references
The GA endpoint accepts images, video and audio; Fal's API allows 12 files, though the company ties its real-time performance claim to four references.
By RuntimeWire Staff · Published
Primary source: Fal
Why it matters
Gur and Yurtseven are pushing Fal beyond model hosting by tuning open weights and inference together, giving developers a faster production loop while deepening Fal's control of the stack.

Burkay Gur and Gorkem Yurtseven's Fal made H3 Max reference-to-video generally available on September 6, extending the founders' speed-first infrastructure thesis to video generations conditioned on existing characters, objects, movement and sound.
Fal says the endpoint is up to twice as fast as its preview and supports real-time generation with as many as four references. The supplied materials do not establish the benchmark's hardware, resolution, clip duration, queue time or file-delivery conditions. Developers will still have to measure end-to-end latency on their own workloads.
Four references at real time, 12 references in the API
Fal's announcement covers real-time generation with up to four references, while the API documentation describes a broader limit of up to 12 reference files. The supplied materials do not establish whether the performance claim extends to that larger limit.
Images can define subjects or style. Video clips can supply movement, while audio can guide sound. Developers identify each asset inside the prompt as "Image 1," "Video 1" or "Audio 1." Individual video and audio references can run from two to 15 seconds, with combined video duration capped at 15 seconds. Audio cannot serve as the only reference; the request also needs an image or video.
The wording around real time can blur two distinct products. The reference-to-video release generates a complete clip faster than that clip takes to play under the stated conditions. H3 Max Director, by comparison, opens a WebRTC session that can be steered with new prompts while video is playing. Both reduce the distance between an instruction and visible output, though they serve different application designs.
Gur and Yurtseven move beyond serving other labs' models
Gur and Yurtseven founded Fal in 2021 after encountering slow and expensive AI infrastructure while working at Coinbase and Amazon. Gur had led machine-learning and platform work at Coinbase after engineering roles at Oracle. Yurtseven worked as an Amazon software developer with experience around AWS machine-learning infrastructure.
The longtime friends began with infrastructure for scaling Python workloads before focusing on generative media. Gur described their founding thesis to TechCrunch in 2024: "The big bet was that the nascent space of generative media was about to change all media consumed."
Fal built its position by hosting and accelerating models from outside labs through one API. With H3 Max, Gur and Yurtseven are also post-training an open-weight model and optimizing it alongside Fal's inference stack. The resulting product sits between a raw MiniMax release and a conventional hosting endpoint.
Fal introduced H3 Max on August 26 as a post-trained version of MiniMax H3, tuned for prompt adherence, aesthetics and speed. RuntimeWire reported on September 1 that H3 Max could produce a five-second video in roughly three seconds under Fal's stated test conditions. General availability for reference-to-video brings that latency target to a heavier workflow where the model must interpret and preserve several supplied assets.
This pace also reflects what investors funded. Fal raised a $140 million Series D on December 9, 2025, led by Sequoia Capital, with Kleiner Perkins and NVIDIA participating. Fal said at the time that it had 70 employees. Sequoia's account singled out the founders' depth across model-provider relationships, kernels, compilers and developer tooling, the same combination Fal is applying to H3 Max.
Reference inputs change the bill
The endpoint costs $0.05 per generated second at 480p and $0.08 at 768p. A five-second 768p output starts at $0.40 before reference charges.
According to Fal's reference-to-video model page, the first 4,096 reference tokens are free, and tokens above that allowance cost $0.02 per 1,000.
Those economics put a sharper edge on the latency claim. Reference-video applications usually generate several candidates before choosing one, so faster inference can shorten an editing loop while multiplying the number of paid attempts a user can make in the same session. Stronger preservation can reduce discarded generations, which may matter as much as raw speed when reference assets cost more than the output clip itself.
Gur and Yurtseven are using H3 Max to tighten Fal's control over both sides of that equation. Post-training determines whether the model follows the supplied material. Fal's serving stack determines how quickly and cheaply developers can test the result. The founders' wager is that production generative media will reward the infrastructure provider that can optimize those layers together, then expose the work through an API before application builders have time to assemble the stack themselves.