Head to head: CogVideoX-5B vs FLUX 3 Image to Video
CogVideoX-5B vs FLUX 3 Image to Video
By Ryan Merket · Published
A demanding set of motion, composition, character, crowd, and fluid-simulation tests exposes a substantial gap between these two video models. The decisive differences emerge not merely in polish, but in whether the requested action and camera direction appear on screen at all.
FLUX 3 Image to Video wins this matchup outright: **32.4 to 20.0**, with a clean **12–0 sweep** and a statistically clear verdict at **limited confidence**. This was not a narrow aesthetic preference or a split decision—FLUX consistently translated detailed prompts into the requested shots and motion. The advantage was clearest in complex camera direction. On the basalt-ridge sprints, FLUX delivered the low, pace-matched side tracking, knife-edge terrain, storm-broken light, and ferocious momentum the prompt demanded; CogVideoX-5B repeatedly fell back to a distant trailing view. In the lighthouse scenes, FLUX produced an actual balcony walk with an orbit and push-in, while CogVideoX often left the subject nearly stationary, weakened the coastal context, or introduced continuity errors such as a duplicated navy cap. FLUX also handled scenes with many moving elements more convincingly. Its Tokyo crossings preserved the wide overhead composition, recognizable street geometry, and independently moving crowds, whereas CogVideoX showed softer figures, smearing, deformation, and tighter framing. The fluid test was even more categorical: FLUX depicted the intact balloon, rupture, expanding water sheet, and dispersing droplets as a coherent sequence; CogVideoX largely reduced the event to underexposed spray against black. FLUX was not flawless—some liquid motion looked synthetic, crowd density shifted, and a few camera moves changed scale too aggressively. But those are refinement problems. CogVideoX more often missed the fundamental staging, action, or camera instruction. **Final call: FLUX 3 Image to Video wins clearly and comprehensively.**
Basalt ridge sprint
A 6-second continuous shot, 16:9: a wiry female fell runner in a citron wind shell and black calf sleeves explodes along the knife-edge of a dark basalt sea cliff above crashing slate-blue water, her stride throwing flecks of grit as she accelerates through a gusting crosswind; the camera races low beside her on a stabilized drone, then arcs forward into a slight lead while matching her pace so the cliff face and white surf streak into natural motion blur, emphasizing extreme momentum; late-afternoon sun punches through torn storm clouds with sharp highlights and cold shadows, and the mood is ferocious, breathless, and exhilarating.
OpenAI: GPT-5.6 Sol Pro: Model B better captures the low side-tracking perspective, dramatic knife-edge cliff, slate-blue surf, storm-broken sunlight, natural foreground blur, and ferocious momentum, while maintaining strong visual consistency. Model A is coherent and clearly depicts basalt coastal running, but its rear-following camera, flatter lighting, and moderate sense of speed miss several key cinematic instructions. (Second judge pass, order swapped — scores are the average of both: Model B strongly captures the | Anthropic: Claude Fable 5.1: Model B follows the prompt much more closel…
Crowd motion
A busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9.
OpenAI: GPT-5.6 Sol Pro: Model B better matches the wide 16:9 overhead Tokyo scramble-crossing prompt, with clearer urban context, more coherent pedestrians, and stronger visual fidelity. Model A maintains a busy crowd but has softer detail, visible motion smearing and malformed figures, while Model B's sharp decrease in crowd density slightly weakens temporal consistency. (Second judge pass, order swapped — scores are the average of both: Model B better matches the requested 16:9 overhead Tokyo scramble view, with shar | Anthropic: Claude Fable 5.1: Model B delivers a convincing overhead vie…
Yellow parka lighthouse walk
A 7-second continuous shot, 16:9: a middle-aged man with a square beard, silver-rim glasses, and a mustard-yellow parka over a red fisherman's sweater walks steadily along the circular balcony of a lonely white lighthouse perched above a vast green-gray ocean, one hand on the rail and the other holding a folded navy cap, and his face, clothing, proportions, and identity must remain perfectly unchanged from first frame to last with no morphing, flicker, or wardrobe drift; the camera performs a slow, smooth clockwise orbit at chest height while gently pushing in, keeping him centered against the sweeping aerial coastal vista below; soft blue dawn light with a faint rotating beacon glow brushes the metalwork, and the mood is calm, introspective, and unwavering.
OpenAI: GPT-5.6 Sol Pro: Model A adheres closely to the character details and keeps the identity stable, but has flatter, overexposed imagery and visibly irregular railing geometry. Model B delivers a more cinematic lighthouse setting, coastal vista, dawn atmosphere, and convincing orbit/push-in progression, despite weaker centering and a somewhat aggressive change in shot scale. (Second judge pass, order swapped — scores are the average of both: Model B more convincingly delivers the lighthouse setting, sweeping coasta | Anthropic: Claude Fable 5.1: Model B delivers a far more cinematic and…
Fluid & particle dynamics
Cinematic slow-motion of a water balloon bursting, the sheet of water expanding and droplets scattering through the air in convincing detail against a dark background, hard side light, 16:9.
OpenAI: GPT-5.6 Sol Pro: Model B clearly depicts the intact balloon, expanding water sheet, and scattering droplets in a coherent slow-motion sequence, though the liquid has a somewhat synthetic CGI appearance. Model A has attractive hard-lit droplets against black but largely misses the balloon burst and expanding sheet, with motion reading mainly as fading spray. (Second judge pass, order swapped — scores are the average of both: Model B clearly depicts an intact balloon transitioning into an expanding water sheet and | Anthropic: Claude Fable 5.1: Model A shows only scattered specks on bla…
Matchup powered by OpenRouter.