Head to head: Kandinsky5 vs Wan v2.6 Image to Video

Kandinsky5 vs Wan v2.6 Image to Video

By · Published

RuntimeWire Head-to-Head: Head to head: Kandinsky5 vs Wan v2.6 Image to Video
RuntimeWire Head-to-Head matchup

Two capable video generators take sharply different routes through camera choreography, long-take continuity, and character stability. The task-level results expose meaningful strengths—and equally meaningful limitations—on both sides.

Wan v2.6 Image to Video posts the higher aggregate score, 28.3 to Kandinsky5’s 23.4, but that gap is not decisive: the analysis assigns only limited confidence that either model is genuinely better. In practical terms, this benchmark is a sample tie, not a narrow victory waiting to be spun into a headline. Wan is plainly stronger at directed motion and multi-step staging. It repeatedly executes the lateral orbit around the ramen bowl with real parallax, while Kandinsky5 often substitutes a push-in or mild drift. Wan also handles the copper-press sequence far more faithfully, including the female technician, gauge work, spinning flywheel, gloves, lid reveal, and camera progression. Its turbine attempt sometimes captures more of the requested action, too—even when wardrobe and identity details slip. Kandinsky5’s counterweight is continuity of subject and atmosphere. It more reliably preserves the turbine inspector’s shaved head, lime-green gloves, maroon scarf, and overall identity, whereas Wan can change hair, gloves, scarf, framing, or activity mid-shot. Kandinsky5 also earns stronger cathedral results when the priority is a warm, candlelit glide through a long nave, although Wan wins other versions with cleaner forward motion and sharper architectural detail. That split defines the matchup: Wan is the better choreographer, but Kandinsky5 can be the steadier continuity editor. Neither advantage survives consistently enough across the full set to separate the models with confidence. **Final call: too close to call—an effective tie.**

Temporal consistency — turbine inspector walkaround

A single continuous shot follows the same aerospace turbine inspector from first frame to last as he walks the length of a partially assembled jet engine in a test bay, and his identity must remain perfectly stable throughout: shaved head with a narrow notch in his left eyebrow, small silver hoop in the right ear, charcoal coveralls with a stitched name patch reading IVO MEREN, lime-green nitrile gloves, and a maroon scarf tucked into the collar; he rolls a borescope cart beside him, pauses to raise the flexible probe into the turbine housing, then leans in and marks one blade with a yellow paint pen before looking back toward the camera; the camera executes an intimate close-up tracking move starting on his gloved hands and gliding backward to a chest-up side profile while pacing him in one unbroken take, lit by crisp white inspection LEDs and faint amber warning beacons for a clinical yet suspenseful mood, 16:9

Kandinsky5:
Wan v2.6 Image to Video:

OpenAI: GPT-5.6 Sol Pro: Model B better captures the prescribed close-up-to-profile progression and yellow blade-marking action with strong turbine detail, though the glove and scarf colors and name patch are incorrect. Model A maintains the inspector’s appearance and requested clothing colors more consistently, but omits the clear paint-marking moment and depicts the probe as an implausibly rigid tube. (Second judge pass, order swapped — scores are the average of both: Model B better captures the requested close-up-to- | Anthropic: Claude Fable 5.1: Model B delivers the actual turbine blades…

Camera motion control

A single smooth orbit (roughly 180 degrees) around a steaming bowl of ramen on a wooden table, the camera circling at a constant speed and height in one continuous move, warm restaurant light, 16:9.

Kandinsky5:
Wan v2.6 Image to Video:

OpenAI: GPT-5.6 Sol Pro: Model B shows a much clearer, smooth lateral orbit with strong parallax and polished warm restaurant visuals, while Model A appears closer to a slow push-in with only minimal angular change. Model A is somewhat more temporally stable, but Model B better fulfills the camera-motion task despite slight changes in the toppings and chopsticks. (Second judge pass, order swapped — scores are the average of both: Model B clearly shows a smooth lateral orbit with strong background parallax, steady framin | Anthropic: Claude Fable 5.1: Model A barely moves—the spoon angle and b…

Continuous shot — copper press calibration

Inside a compact enamelware factory, a grease-smudged maintenance technician in a faded teal coverall and orange earplugs performs a live calibration on a humming 1930s copper-stamping press: she kneels to slide a brass feeler gauge under the die, rises as the flywheel gathers speed, twists a handwheel exactly one notch, then catches a freshly embossed kettle lid with insulated gloves and holds it toward the lens so the concentric rings shimmer; the camera makes one unbroken slow arc from a tight waist-level side view to a close frontal tracking position around her right shoulder, with no cuts, under warm sodium work lamps mixed with cool daylight leaking through dusty clerestory windows, creating a tense but quietly triumphant mood, 16:9

Kandinsky5:
Wan v2.6 Image to Video:

OpenAI: GPT-5.6 Sol Pro: Model B closely follows the requested sequence with a kneeling female technician, visibly spinning flywheel, press operation, and a strong final presentation of the concentric-ring lid, while Model A omits most key actions and depicts the wrong technician. Model A is reasonably consistent visually, but Model B has much stronger fidelity, staging, and overall aesthetics despite the sampled camera progression feeling less continuous. (Second judge pass, order swapped — scores are the average of bo | Anthropic: Claude Fable 5.1: Model A keeps a single continuous shot but…

Single continuous shot

One unbroken take gliding slowly through a candlelit cathedral from the entrance toward the altar, no cuts, jumps, or transitions, dust and warm light in the air, 16:9.

Kandinsky5:
Wan v2.6 Image to Video:

OpenAI: GPT-5.6 Sol Pro: Model A shows a smooth, centered advance toward the altar with highly stable architecture, warm candlelight, and atmospheric haze. Model B is more detailed aesthetically, but conspicuous changes in column carvings across frames weaken temporal consistency and the impression of one unbroken glide. (Second judge pass, order swapped — scores are the average of both: Model A more convincingly maintains a smooth forward glide toward the altar, with stable architecture, coherent parallax, warm candlel | Anthropic: Claude Fable 5.1: Model A delivers a genuinely candlelit, ha…

Matchup powered by OpenRouter.