Head to head: Bytedance Seedance V1.5 Pro Image To Video vs Flux 3 Text to Video
This matchup splits cleanly down the middle: Seedance looks stronger when camera motion, atmosphere, and cinematic polish matter most, while Flux is better at obeying fiddly action and occlusion instructions. On the aggregate, though, the gap is small enough to be meaningless — this is effectively a tie.
By RuntimeWire · Published

Seedance and Flux trade punches in a way that makes the overall result impossible to spin into anything cleaner than a draw. The aggregate scores — 30.9 for Bytedance Seedance V1.5 Pro Image To Video versus 33.5 for Flux 3 Text to Video — land in statistical dead-heat territory, with only 50% confidence that either model is genuinely better. In plain English: they are effectively even here.
Where Seedance earns its keep is in shots that live or die on cinematic execution. It won Neon ferry sprint by delivering the more convincing low-angle, urgent lateral run: stronger blue-hour reflections, clearer motion, and a more polished sense of speed. It also took Single continuous shot, where its cathedral sequence better sold the brief’s warm candlelit glide toward the altar, with cleaner forward movement, stronger volumetric light, and a more evocative dust-filled atmosphere.
Flux, meanwhile, was better at the kind of prompt adherence that exposes a model’s discipline. In Balloon behind columns, it handled the actual assignment: a recognizable balloon moving through believable partial and full occlusion behind columns and a fern, while preserving the winter-garden market context. And in Subject action, it produced the more credible overhead latte-art pour, with motion that resolves into something rosetta-like instead of drifting into the wrong pattern.
That split tells you what this comparison really is. Seedance is the more persuasive stylist when the brief rewards mood, camera feel, and visual coherence. Flux is the more reliable instruction-follower when the prompt contains specific object interactions or narrowly defined actions. Neither model established enough separation across the set to claim the crown.
Final call: Too close to call. Seedance wins on cinematic polish; Flux wins on prompt fidelity; overall, this head-to-head is a genuine tie.
How they were tested
We ran 4 fresh video tasks, generated on the fly for this matchup so neither model could prepare in advance, and had gpt-5.4 score each one. To cancel position bias, every task was judged twice — once in each presentation order — and every number reported here, including the headline totals, is the average of both passes. Bytedance Seedance V1.5 Pro Image To Video scored 30.9 to Flux 3 Text to Video's 33.5.
1. Neon ferry sprint
A single continuous 16:9 shot at blue-hour on the wet public promenade outside the Västerkajen ferry terminal: a courier in a saffron windbreaker explodes into a full sprint through a stream of commuters, clutching a dented silver document tube under one arm, shoes kicking up fine spray from the rain-slick paving; the camera starts low beside his trailing foot, then accelerates in a fast lateral tracking move that arcs slightly forward to keep pace with him as LED timetable boards, umbrellas, and kiosk lights streak into tasteful motion blur, emphasizing raw momentum; cool cyan reflections from the harbor mix with warm sodium streetlamps on faces and puddles, and the mood is urgent, electric, and breathless.
Winner: Bytedance Seedance V1.5 Pro Image To Video — Model A matches the prompt more closely with the low opening angle by the runner’s foot, stronger sense of explosive sprinting, clearer lateral tracking, and richer blue-hour reflections with tasteful motion blur; Model B captures the ferry-terminal promenade and document tube well, but the courier styling is less accurate, the motion feels softer and less urgent, and the imagery is less polished and consistent. (Second judge pass, order swapped — scores are the average of both: Model B captures the blue-hour ferry-promenade setting, low start, and lateral sprint well, but the courier’s windbreaker reads more mustard than saffron and the motion blur/subject clarity feel a bit softer and less polished. Model A is more visually striking and coherent, with stronger reflections, clearer commuter context, and a more convincing urgent run with the silver tube, though it is slightly less faithful in wardrobe styling and exact camera-start framing.)
2. Single continuous shot
One unbroken take gliding slowly through a candlelit cathedral from the entrance toward the altar, no cuts, jumps, or transitions, dust and warm light in the air, 16:9.
Winner: Bytedance Seedance V1.5 Pro Image To Video — Model A better matches the prompt with a clear slow forward glide from the entrance toward the altar in a single continuous-feeling shot, with strong warm candlelight, visible light rays, and dust-filled atmosphere. Model B is also coherent and attractive, but its darker look and pew-centered composition feel less aligned with the specified warm airy cathedral glide, and the motion reads slightly less smooth and evocative. (Second judge pass, order swapped — scores are the average of both: Model B matches the cathedral glide prompt well and appears temporally stable, but the lighting is darker and the sense of dust-filled warm atmosphere is less pronounced. Model A better captures the candlelit, warm-light mood with stronger volumetric rays and a clean forward glide toward the altar, while remaining consistent and more visually striking overall.)
3. Balloon behind columns
One continuous 16:9 shot inside the crowded municipal winter garden of Mercado de Santa Elina at noon: a little girl in a plum wool coat walks steadily along the tiled arcade holding a bright octopus-shaped turquoise helium balloon with one missing eye sticker, weaving past shoppers carrying bread and flowers; the camera performs a gentle handheld dolly backward at her eye level as she and the balloon pass behind a row of thick cream stone columns and a tall fern display, becoming partially and then fully occluded for a moment before re-emerging unchanged on the other side, with the balloon’s crooked ribbon, missing sticker eye, and bobbing height remaining perfectly consistent; soft skylight pours through the glass roof, catching dust in the air, and the mood is calm, observational, and quietly magical.
Winner: Flux 3 Text to Video — Model B matches the prompt more completely, clearly showing the winter-garden market, shoppers with bread and flowers, and the key occlusion/re-emergence behind both columns and a fern while keeping the balloon recognizable and stable. Model A looks polished, but it misses important prompt details such as the fern occlusion and the balloon’s specified imperfections, and its sampled progression suggests less complete handling of the behind-columns action. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt more closely: the girl walks through a crowded winter-garden arcade with bread and flowers visible, and the balloon becomes partially then fully occluded by cream stone columns and a fern before re-emerging in a believable way. Model A is visually polished, but it breaks key prompt details by showing the balloon dragged on the floor, changing the girl's styling, and not clearly delivering the specified behind-columns-and-fern occlusion with consistent balloon behavior.)
4. Subject action
A barista's hands pouring latte art: the milk stream forms a clean rosetta in the crema with natural, fluid wrist motion, no cuts, overhead close-up, soft café light, 16:9.
Winner: Flux 3 Text to Video — Model B better matches the prompt by producing a recognizable rosetta/tulip-like latte art pattern from an overhead close-up with cleaner progression and more believable hand-and-pitcher motion. Model A is visually sharp, but the pattern develops as concentric rings rather than a rosetta, making it less adherent despite decent consistency and aesthetics. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt more closely with an overhead close-up, soft café lighting, and a convincing progression into a clean rosetta-like latte art pattern with coherent hand positioning across frames. Model A is visually sharp, but the angle is less overhead and the pour develops into concentric rings rather than a rosetta, making it less adherent despite decent consistency and aesthetics.)
See every prompt and the full side-by-side outputs in the interactive Head-to-Head.