Head to head: CogVideoX-5B vs Flux 3 Image to Video
CogVideoX-5B vs Flux 3 Image to Video
One model showed up with stronger camera sense, cleaner action, and far better prompt discipline across every test. The other had occasional atmosphere, but never enough control to make this matchup competitive.
This wasn’t a split decision. **Flux 3 Image to Video wins decisively**, taking all four tasks and finishing with a **35.1 to 24.0** aggregate lead. The statistical read is even harsher for CogVideoX-5B: **97% confidence** that Flux is the better model in this matchup. No ties, no rescue category, no ambiguity. What stands out is *how* Flux won. In **Pelota Chalk Burst**, it actually rendered the jai alai setup the prompt asked for: the cesta, the yellow pelota, the blue chalk impact, and a readable progression toward contact. CogVideoX-5B, by contrast, buried the scene under an overblown blue haze, leaving only fragments of the athlete and almost no legible action. That pattern repeated elsewhere: Flux kept giving the judges coherent motion and identifiable subject matter; Cog too often gave mood without enough usable scene information. The camera-motion tests were especially one-sided. In **Single continuous shot**, Flux delivered a believable glide through a grand candlelit cathedral with depth, dust, warm light, and clear forward advancement toward the altar. CogVideoX-5B looked attractive, but more static and spatially constrained. In **Canal Row Dawn**, Flux again showed stronger shot design: low water-level tracking, old brick boathouses, mist, ripples, and a convincing ease toward the sculler’s shoulder. Cog had some nice dawn texture, but it missed the specified setting and never really found the requested tracking feel. Even in the close-up action task, where a polished image can sometimes mask weak motion logic, Flux was simply more faithful. In **Subject action**, it clearly showed the overhead barista view and the milk forming a proper rosetta with believable progression. CogVideoX-5B produced something prettier than useless, but it drifted into concentric swirls rather than the requested wrist-driven latte-art pattern. **Final call: Flux 3 Image to Video is the clear winner.** It was better on prompt adherence, temporal continuity, camera movement, and action readability in every single test. CogVideoX-5B had flashes of atmosphere, but Flux was the only model here consistently making the right video.
Pelota Chalk Burst
A single continuous 16:9 slow-motion shot captures a jai alai athlete in a white helmet and navy kit whipping a yellow pelota into the front wall of a small outdoor court dusted with blue chalk powder, the camera making a smooth lateral gimbal move from behind the player’s cesta to a three-quarter side view as the ball strikes, explodes through the chalk cloud, and ricochets back while tiny grains, fabric folds, flexing wicker, and the athlete’s calf muscles resolve in crisp fluid detail under hard late-afternoon sun, creating an intense, razor-focused mood.
Model B clearly depicts jai alai with the cesta, yellow pelota, blue chalk on the wall and court, and a coherent progression toward impact, while Model A is largely obscured by an overblown chalk cloud and only shows partial legs with little readable action. Model B also has stronger visual clarity and continuity, though it falls a bit short of the exact helmeted three-quarter slow-motion impact moment described in the prompt. (Second judge pass, order swapped — scores are the average of both: Model B matches the jai alai setup, yellow pelota, blue chalk impact on the front wall, and outdoor late-day look with clearer subject detail and a more coherent progression toward the strike. Model A is dominated by an overexposed blue haze that obscures the athlete, ball, and action, weakening both prompt adherence and visual clarity.)
Single continuous shot
One unbroken take gliding slowly through a candlelit cathedral from the entrance toward the altar, no cuts, jumps, or transitions, dust and warm light in the air, 16:9.
Model A matches the candlelit cathedral mood well, but the sampled frames show only subtle forward progress and feel comparatively static. Model B better conveys a single continuous glide from entrance toward the altar with clearer temporal advancement, strong atmospheric dust and warm light, and more convincing cathedral scale and depth. (Second judge pass, order swapped — scores are the average of both: Model B better matches the prompt with a clear slow forward glide through a grand candlelit cathedral toward the altar, with convincing depth, warm haze, and stable progression across frames. Model A is visually appealing and consistent, but it feels more like a narrower aisle shot with less atmospheric dust/light presence and a more static, less expansive cathedral impression.)
Canal Row Dawn
A single continuous 16:9 shot at dawn follows a lone sculler in a matte teal racing shell powering down a narrow training canal beside an old brick boathouse district, the camera gliding low and parallel just above the water before easing slightly forward toward the athlete’s shoulder; silver ripples peel off each oar stroke, reeds and poplar branches sway in a light breeze, thin mist lifts from the canal, laundry on a distant line flutters, and layered clouds drift steadily across a pale apricot sky, all lit by cool early sunlight with a focused, quietly determined mood.
Model B matches the prompt much more closely with the old brick boathouse district, matte teal shell, misty dawn atmosphere, and a low parallel camera that eases closer toward the sculler while maintaining strong visual coherence. Model A has decent rowing imagery and laundry/mist cues, but it misses the boathouse setting and shoulder-near tracking feel, and looks flatter and less cinematic overall. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt more closely with the old brick boathouse district, low water-level tracking feel, mist, ripples, and a clear forward easing toward the sculler’s shoulder in a continuous, coherent progression. Model A has pleasing dawn light and the laundry detail, but it misses the boathouse setting and parallel camera motion, and its framing feels more static and less aligned with the specified shot design.)
Subject action
A barista's hands pouring latte art: the milk stream forms a clean rosetta in the crema with natural, fluid wrist motion, no cuts, overhead close-up, soft café light, 16:9.
Model B matches the prompt much better by clearly showing a barista’s hands from an overhead close-up as the milk forms a clean rosetta with believable progression and pleasing café lighting. Model A has decent visual fidelity, but the pattern develops more like concentric rings than a rosetta and feels less aligned with the requested wrist-driven latte-art action. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt closely with an overhead close-up of barista hands forming a clean rosetta through believable, fluid progression and consistent soft café lighting. Model A looks visually polished, but the pattern develops more like concentric swirls than a rosetta, and the framing/motion feel less faithful to the specified natural wrist-driven latte art action.)
Matchup powered by OpenRouter.