Head to head: Cosmos Predict 2.5 2B vs Seedance 2.5 Reference to Video
Cosmos Predict 2.5 2B vs Seedance 2.5 Reference to Video
This one isn’t a blowout, but the scoreboard is real: Seedance 2.5 Reference to Video takes the matchup on aggregate and wins three of four tasks.
Cosmos Predict 2.5 2B does have a case in exactly one lane: **crowd motion**. In the scramble-crossing test, it produced the busier, more independently animated pedestrian field, and that mattered. Seedance’s version had the more recognizable Tokyo setting and steadier camera, but it simply wasn’t crowded enough for the prompt, so Cosmos earned the win on behavioral richness and continuity. Everywhere else, though, Seedance was the more trustworthy generator. In **Cymbal Water Burst**, it actually delivered the prompt’s dark ceremonial stage, overhead spotlight, blue reflections, slow-motion spray, and crimson-sleeve atmosphere. Cosmos looked like a cleaner studio splash study, but it missed the intended scene grammar and never really sold the finger-cymbal performance. The gap was even clearer in **Ribbon Hoop Drop** and **Subject action**. Seedance understood the mechanics: the hoop is released, tumbles, bounces, skids, and wobbles with believable ribbon lag; the latte pour resolves into a readable rosetta with fluid hand motion and stable overhead framing. Cosmos, by contrast, repeatedly drifted into near-miss territory—turning the hoop into a static prop, shifting the camera away from the requested perspective, and producing something closer to decorative swirls than the specified coffee action. That pattern explains the result better than the raw totals do. Seedance wins **32.1 to 23.6**, takes **3 of 4 tasks**, and does so with **81% confidence**, which is a lean rather than a demolition. So no, this isn’t an untouchable rout. But it is a solid editorial verdict: **Seedance 2.5 Reference to Video is the better video model here because it follows the prompt more faithfully across scene setup, motion physics, and temporal progression, while Cosmos Predict 2.5 2B looks strongest only when sheer crowd density is the main ask.**
Crowd motion
A busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9.
Model A better matches the prompt with a busy overhead scramble crossing full of independently moving pedestrians, and the crowd remains visually coherent across frames without obvious merging or warping. Model B has a more recognizable Tokyo cityscape and stable camera, but it is too sparse for a busy scramble crossing and the pedestrian motion reads less richly populated and less aligned with the requested crowd behavior. (Second judge pass, order swapped — scores are the average of both: Model B matches the Tokyo scramble crossing setting and overhead composition well, but the scene is too sparse for a busy crossing and the motion appears limited across frames. Model A shows denser, more independent pedestrian movement with stronger crowd dynamics and cleaner temporal continuity, though it is less clearly identifiable as the iconic Tokyo scramble from above.)
Cymbal Water Burst
A single continuous 16:9 ultra-slow-motion shot captures a flamenco percussionist snapping two small bronze finger cymbals together above a shallow black stage puddle, sending a razor-thin crown of water and glittering droplets upward while her crimson fringe sleeve trembles; the camera performs a deliberate close orbit from front-left to profile at chest height, preserving crisp frame-to-frame detail as ripples, spray, and sleeve fibers move fluidly in slow motion; lit by a hard white spotlight from above and faint blue footlights reflecting in the water, the mood is electric and ceremonial.
Model B matches the prompt far better with a ceremonial stage mood, hard overhead spotlight, blue reflections, visible slow-motion spray, and a flamenco-like crimson sleeve rendered with cohesive motion and atmosphere. Model A looks more like a studio tabletop splash with one oversized cymbal-like object, weaker adherence to the finger-cymbal action and orbiting composition, and less convincing temporal progression. (Second judge pass, order swapped — scores are the average of both: Model B better matches the intended dark ceremonial stage mood, spotlight-and-blue-footlight lighting, and slow-motion water spray with a more cinematic, cohesive look. Model A shows clearer cymbal contact with water, but it departs from the prompt with a bright studio setup, a drum-like surface instead of a shallow black puddle, and less elegant temporal progression.)
Ribbon Hoop Drop
A single continuous 16:9 shot in a dim rehearsal hangar shows a contemporary dancer in a copper unitard sprinting and releasing a 1.2-meter aluminum hoop wrapped with long teal ribbon tails from a 2.5-meter platform; the camera starts low beside scattered rosin dust, then makes a smooth handheld orbit around the dancer as the hoop tumbles, the ribbons lag and whip from air resistance, the hoop clips the sprung floor, bounces unevenly, skids, and wobbles to a stop under believable gravity and momentum; side-lit by cold skylight and one warm amber work lamp, the mood is tense and raw.
Model B matches the prompt far better: it shows the dancer on a platform in a dim hangar, the hoop release and floor interaction, believable ribbon lag, and a low cinematic viewpoint with strong cold/warm lighting. Model A is visually clean but largely misinterprets the action, turning the hoop into a static oversized prop around the dancer rather than a released object tumbling, bouncing, and skidding under convincing physics. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt far better: it presents the dim hangar, copper unitard, low floor-level perspective, hoop release, ribbon lag, and a plausible bounce/skid/wobble progression with coherent lighting and momentum. Model A is brighter and more staged, keeps the hoop largely static on the platform with the dancer moving through it instead of dropping and tumbling it, and misses the specified camera/motion behavior despite decent image clarity.)
Subject action
A barista's hands pouring latte art: the milk stream forms a clean rosetta in the crema with natural, fluid wrist motion, no cuts, overhead close-up, soft café light, 16:9.
Model B matches the prompt more closely by clearly forming a clean rosetta with believable progression and fluid hand motion, while maintaining strong temporal coherence and pleasing café-like lighting. Model A looks polished but the pattern reads more like concentric swirls than a rosetta, and the framing is less overhead and less faithful to the requested action. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt closely with an overhead close-up of hands pouring a clean rosetta, and the sampled frames show natural progression and stable, fluid motion. Model A is visually polished but deviates in composition and subject action: the angle is not overhead close-up, the vessel looks like a martini-style glass rather than a latte cup, and the pattern reads more like concentric swirls than a clear rosetta.)
Matchup powered by OpenRouter.