Head to head: LTX Video (preview) vs Seedance 2 Image to Video
LTX Video (preview) vs Seedance 2 Image to Video
By Ryan Merket · Published
This one wasn’t competitive. Across all 12 judged tasks, Seedance 2 Image to Video consistently delivered closer prompt adherence, stronger motion coherence, and far better temporal stability than LTX Video (preview).
Seedance 2 Image to Video wins this matchup outright, and the numbers leave very little room for debate: **34.5 to 19.4 overall, 12 task wins to 0, with 95% confidence**. That’s not a stylistic preference split or a narrow edge on a few cherry-picked prompts. It’s a clean sweep. What stands out is how repeatable the gap was. In **crowd motion**, Seedance actually gave the judges the shot they asked for: a recognizable overhead Tokyo scramble crossing with dense, independent pedestrian flows that stayed distinct over time. LTX could gesture at “busy crosswalk,” but it repeatedly fell apart on framing, specificity, and basic scene integrity, with blur, warping, and pedestrians smearing into each other. The same pattern showed up in the more demanding cinematic prompt, **Cliffside Water Tank Spill**. Seedance hit the brief with the right vehicle form, blue-hour mood, labeled tanker, cones, tarp motion, and water that behaved like water—spraying, arcing, puddling, and spilling over the cliff in a coherent camera move. LTX missed core prompt elements outright, often reading more like a generic van in the wrong setting, with weaker physics and shakier temporal continuity. And on the fundamentals, Seedance was simply more dependable. In **temporal consistency**, it kept the man’s face, yellow raincoat, and umbrella stable through a steady rainy tracking shot, while LTX showed obvious coat morphing and artifacting. In **Paraglider Identity Lock**, LTX sometimes held a close-up better, but Seedance more often captured the actual requested shot design and key identity cues—helmet marking, harness details, compass, sunrise aerial motion—rather than just producing a decent-looking person in the air. **Final call: Seedance 2 Image to Video is the clear winner. LTX Video (preview) isn’t losing on taste; it’s losing on prompt fidelity, physical plausibility, and frame-to-frame reliability, and it lost every single task here.**
Crowd motion
A busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9.
Model B matches the prompt much better with a clear overhead view of a Tokyo scramble crossing, dense independent pedestrian flows, and stable, realistic scene structure across frames. Model A shows crowd motion but the framing is too tight and generic, with noticeable blur/warping and less convincing separation between pedestrians. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt closely with a clear overhead view of a Tokyo scramble crossing, many pedestrians moving independently in multiple directions, and strong scene coherence across frames. Model A shows a more generic crosswalk scene with heavy blur and deformation, weaker Tokyo-specific context, and less reliable independent pedestrian motion.)
Cliffside Water Tank Spill
At blue-hour on the wind-scoured basalt rim above Lake Venshar, a dented silver water tanker labeled "Kestrel-9 Survey" slowly tips on a gravel turnout and releases a sudden sheet of water from its side hatch; the water must arc, break into droplets, slam onto the sloped rock, split around scattered orange survey cones, and pour over the cliff edge with believable gravity, momentum, spray, and puddling while a loose tarp on the tanker whips in gusts; the camera begins as an epic aerial wide over the canyon, then descends in one smooth continuous crane-and-dolly move to a medium tracking angle beside the spill, cool twilight light with sharp work-lamp highlights, tense realistic mood, 16:9
Model B matches the prompt far better with the requested aerial-to-medium descent, blue-hour canyon setting, recognizable dented silver tanker labeled "Kestrel-9 Survey," whipping tarp, cones, and a more believable water spill that arcs, sprays, puddles, and goes over the cliff. Model A has some spill interaction with cones, but it looks like a small van rather than a tanker, misses the cinematic camera progression and twilight mood, and the water behavior appears less realistic and less temporally coherent. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt far better: it shows the correct tanker on a cliffside at blue hour, includes the labeled survey vehicle, cones, tarp, and a believable water arc that puddles and spills over the edge with coherent camera progression. Model A departs heavily from the prompt with a van-like vehicle instead of a tanker, daylight lighting, weaker realism in the spill behavior, and lower temporal and visual consistency.)
Temporal consistency
A man in a yellow raincoat walking toward camera down a rainy street; his face, coat, and umbrella must stay perfectly consistent with no morphing or flicker from the first frame to the last, steady tracking shot, 16:9.
Model B better matches the prompt with a steady tracking shot of a man in a yellow raincoat walking toward camera in a rainy street, while keeping his face, coat, and umbrella highly consistent across frames. Model A follows the basic setup but shows weaker facial consistency and less convincing motion/overall polish, whereas Model B is more cinematic and temporally stable. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt closely with a steady tracking shot on a man in a yellow raincoat walking toward camera on a rainy street, while keeping his face, coat, and umbrella highly consistent across frames. Model A shows a more static, less cinematic scene with weaker street ambience and noticeable softness and inconsistency in facial detail and overall motion continuity.)
Paraglider Identity Lock
One continuous shot over the terraced green ridges of Mount Selcor at sunrise: a single female paraglider pilot with a shaved left eyebrow slit, a matte teal helmet marked "R-17", a saffron jacket with one black sleeve, white harness straps, and a tiny silver compass clipped to her chest glides steadily along the cliff line while making two gentle banking turns; her face, clothing colors, helmet markings, harness details, and body proportions must remain perfectly identical from first frame to last with no morphing, flicker, or accessory changes as the camera performs a smooth helicopter-like side follow that gradually orbits from a distant aerial wide to a closer three-quarter profile, warm low sun, crisp alpine haze, serene but exhilarating mood, 16:9
Model B better matches the requested sunrise aerial side-follow/orbit over terraced ridges and shows the helmet marking "R-17" with smoother cinematic motion, while Model A stays in a close selfie-like view and misses several specified identity details. Model A has decent character continuity, but Model B is stronger overall despite some outfit/color mismatches and less verifiable facial-detail consistency at distance. (Second judge pass, order swapped — scores are the average of both: Model B better matches the requested sunrise side-follow over terraced ridges and shows a coherent wide-to-closer progression with smoother gliding motion, though some identity details are too distant to verify and the jacket/helmet specifics are only partially clear. Model A presents the pilot closer, but it misses key prompt details such as the saffron jacket with one black sleeve, visible helmet marking and compass, and the overall shot feels less like the specified helicopter-like orbit with weaker environmental match and consistency.)
Matchup powered by OpenRouter.