Head to head: Flux 3 Image To Video Draft vs Infinity Star

This one isn’t a blowout, but the edge is real: Flux 3 Image To Video Draft wins on prompt execution and scene choreography, while Infinity Star’s best case is steadier frame-to-frame consistency. Across the four tests, Flux is the model more likely to give you the shot you actually asked for.

By · Published

Visual comparison of AI-generated video sequence quality from competing models (Claymation still — hand-modelled plasticine scene, visible fingerprints, warm studio lighting)

Flux 3 Image To Video Draft takes this matchup on both the scoreboard and the substance. The aggregate gap — 34.4 to 28.7 — is meaningful rather than massive, and the stats back that up: Flux wins with 81% confidence, which is a lean, not a landslide. Still, the task count is decisive enough: Flux wins 3 of 4, with no ties.

What separates Flux is prompt fidelity in motion-heavy, spatially specific scenes. In Rush Through Platform 11, it’s the only model that really delivers the brief: low-angle start from the blocks, crowded station-concourse energy, side-tracking movement, kiosks, reflections, and believable speed blur. Infinity Star looks cleaner, but it drifts into the wrong scene grammar entirely — more indoor track than frantic commuter sprint. The same pattern shows up in Blue Hour Waltz Under Glass, where Flux better captures the wet cobblestones, bundled crowd, street-sweeper, and blue-hour amber atmosphere; Infinity Star again feels tidy but underdirected, too empty and too static for the prompt.

Flux also wins the single continuous shot test for a simple reason: it understands the difference between a push-in and a traversal. Its cathedral clip reads like an actual glide from the entrance down the nave toward the altar, with depth and dusty warm light carrying the shot. Infinity Star is attractive and stable, but it feels staged closer to the destination, which undercuts the whole point of the prompt.

Infinity Star’s lone win, temporal consistency, is legitimate. It keeps the man, umbrella, coat, and centered tracking notably more stable across frames, while Flux shows small shifts in facial visibility, hood presentation, and framing. If your top priority is subject stability above all else, Infinity Star has a real argument. But in this head-to-head, that strength doesn’t outweigh Flux’s broader advantage in actually following complex camera language and scene instructions.

Final call: Flux 3 Image To Video Draft wins. Not by knockout, but by being the more dependable director’s tool — the model that more often turns a detailed prompt into the right shot, not just a polished approximation.

How they were tested

We ran 4 fresh video tasks, generated on the fly for this matchup so neither model could prepare in advance, and had gpt-5.4 score each one. To cancel position bias, every task was judged twice — once in each presentation order — and every number reported here, including the headline totals, is the average of both passes. Flux 3 Image To Video Draft scored 34.4 to Infinity Star's 28.7.

1. Rush Through Platform 11

A single continuous 16:9 shot of a neon-orange track sprinter in lane markings painted for a promotional race blasting through a crowded commuter concourse labeled Platform 11, exploding off the line and weaving past flaring ad kiosks as loose ticket stubs skitter in the slipstream; the camera starts low at ankle height beside the blocks, then accelerates in a tight side-tracking move that keeps pace with the runner before arcing slightly forward for a breathless near-profile close-up, with harsh white station LEDs and streaks of cyan timetable light reflecting off the polished floor, strong momentum, rhythmic footfall vibration, and natural motion blur in the background creating an urgent, electric mood.

Winner: Flux 3 Image To Video Draft — Model A matches the prompt far better, showing a low-angle start from the blocks in a crowded Platform 11-style concourse with side-tracking motion, ad kiosks, polished floor reflections, and convincing speed blur. Model B is cleaner but misses the specified camera path and commuter-concourse chaos, instead resembling a centered frontal run on a conventional indoor track with weaker prompt alignment. (Second judge pass, order swapped — scores are the average of both: Model A matches the prompt far more closely with the low ankle-height start at the blocks, a crowded Platform 11-style concourse, side-tracking sprint motion, ad kiosks, ticket stubs, and convincing motion blur. Model B is cleaner but misses the specified camera language and station-concourse action, reading more like a centered frontal run on an indoor track than a breathless rush through commuters.)

2. Temporal consistency

A man in a yellow raincoat walking toward camera down a rainy street; his face, coat, and umbrella must stay perfectly consistent with no morphing or flicker from the first frame to the last, steady tracking shot, 16:9.

Winner: Infinity Star — Model B is more temporally consistent: the man's face, coat, umbrella, and centered tracking remain very stable across frames with no visible morphing, while Model A shows slight changes in facial visibility/hood presentation and a bit more variation in framing. Model A has a moodier, more cinematic rainy street, but Model B better satisfies the core consistency requirement while still looking clean and coherent. (Second judge pass, order swapped — scores are the average of both: Model B adheres more closely to the prompt with a bright yellow raincoat, clear face visibility, centered steady tracking, and very strong consistency in the man, coat, and umbrella across frames. Model A is atmospheric and visually appealing, but the hood obscures the face, the coat reads more mustard than bright yellow, and there is slightly less exact consistency in framing and subject appearance.)

3. Blue Hour Waltz Under Glass

A single continuous 16:9 shot in a covered winter market arcade where a young municipal street-sweeper slowly dances with her broom through the thinning evening crowd, her movement calm and unhurried as she spins once around a puddle and glides forward between bundled strangers who subtly part for her; the camera begins in a gentle backward dolly at chest height, then eases into a slow circular move around her while staying close, letting the vaulted glass roof catch the last blue dusk and warm amber stall lights gradually brighten across drifting steam, with soft reflections on wet cobblestones and a serene, quietly joyful mood built through the lingering pace and evolving light.

Winner: Flux 3 Image To Video Draft — Model A better matches the prompt’s winter market mood, wet reflective cobblestones, bundled crowd, and street-sweeper with broom, while maintaining a more cinematic blue-hour/amber-light atmosphere. Model B is cleaner and more stable, but it feels too empty and static, with weaker evidence of the dance-like interaction, drifting steam, puddle/reflection emphasis, and evolving close circular camera movement described in the prompt. (Second judge pass, order swapped — scores are the average of both: Model A better matches the prompt’s municipal street-sweeper in a covered winter market arcade, with wetter cobblestones, stronger blue-hour/amber-light atmosphere, and clearer interaction with the crowd and broom. Model B is clean and stable, but it feels more like a centered walk toward camera than a calm dance, with less evident circular camera motion and less of the evolving reflective mood described.)

4. Single continuous shot

One unbroken take gliding slowly through a candlelit cathedral from the entrance toward the altar, no cuts, jumps, or transitions, dust and warm light in the air, 16:9.

Winner: Flux 3 Image To Video Draft — Model A better matches the prompt of a slow unbroken glide through a cathedral from the entrance toward the altar, with clear forward progression down the nave, strong depth, and visible dusty warm light. Model B is visually pleasing and temporally stable, but it feels more like a shorter push-in already near the altar rather than a traversal from the entrance through the cathedral space. (Second judge pass, order swapped — scores are the average of both: Model B matches the candlelit cathedral mood and shows a smooth forward glide, but it feels more like a short push-in near the altar than a journey from the entrance, with a somewhat staged, less airy look. Model A better conveys a single continuous move down the nave toward the altar, with stronger depth, visible dust-filled warm light, and more convincing temporal continuity and overall realism.)


See every prompt and the full side-by-side outputs in the interactive Head-to-Head.

Reader comments

Conversation for this story loads after sign-in.