Head to head: Kandinsky5 vs MiniMax H3 Reference to Video
Kandinsky5 vs MiniMax H3 Reference to Video
One model kept landing the brief; the other kept drifting into adjacent, less specific scenes. Across all four tests, MiniMax H3 Reference to Video separated itself on prompt fidelity and multi-subject coherence, turning this into a decisive result rather than a stylistic debate.
Kandinsky5 doesn’t lose here because it looks bad. In several clips it’s reasonably attractive and sometimes even a touch steadier. It loses because this matchup was judged on whether the model actually delivers the shot you asked for, and on that metric MiniMax H3 Reference to Video was simply in another tier. The clearest example is **Crowd motion**. The prompt called for a busy Tokyo scramble crossing from an elevated angle, with many pedestrians moving in multiple directions. MiniMax gave you that: dense, overhead, legible crowd flow. Kandinsky5 gave you a sparser, more frontal street-crossing scene that misses the defining geometry and energy of a true scramble. That pattern repeated in **Lighting transition**, where MiniMax produced the requested locked-off living room, believable dusk shift, and lamp turning on near the end, while Kandinsky5 wandered into something more like an outdoor patio and never really executed the key lighting beat. MiniMax also won the more compositionally demanding prompts. In **Moon Jelly Orbit**, it better captured the exact subject-centric orbit, the moon-jelly detail, the caustic aquarium mood, and even more of the specified staging elements. Kandinsky5’s underwater imagery was decent, but it read like a looser pan through a related scene rather than the shot described. In **Reedbed Egret Tangle**, the gap was even more practical: MiniMax actually maintained the requested cluster of birds and dragonflies with coherent independent motion, while Kandinsky5 underdelivered on animal count and showed shakier subject continuity. The numbers reflect that this wasn’t close: **34.6 to 22.4 overall, four task wins to zero, with MiniMax H3 Reference to Video ahead at 97% confidence**. That is decisive, not interpretive. Kandinsky5 showed occasional aesthetic appeal, but aesthetics can’t compensate when the model repeatedly substitutes a nearby idea for the actual prompt. **Final call: MiniMax H3 Reference to Video is the clear winner. If you care about faithful scene construction, specific camera behavior, and keeping multiple moving elements coherent, it outclasses Kandinsky5 in this head-to-head.**
Crowd motion
A busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9.
Model B matches the prompt much better with a true overhead view of a busy Tokyo scramble crossing and many pedestrians moving in multiple directions while remaining mostly distinct. Model A is temporally steadier but shows a sparse, more frontal street crossing rather than a dense scramble, so it adheres less well to the requested scene. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt much better with an elevated view of a busy Tokyo-style scramble crossing and many pedestrians moving in multiple directions while remaining mostly distinct and coherent. Model A has decent stability and overcast lighting, but it shows a simpler straight crosswalk scene rather than a true scramble crossing and has far fewer independently moving pedestrians, making it less faithful and less dynamic.)
Moon Jelly Orbit
Inside a dim cobalt aquarium hall, one continuous 16:9 shot shows a translucent moon jellyfish pulsing upward beside a rusted research buoy wrapped in fine green algae while tiny air bubbles spiral past and a lone pipefish flicks through the background; the camera performs a slow, exact orbit around the jellyfish and buoy, maintaining a steady radius and smooth circular motion with no drift into a dolly or pan, so the subject rotates in perspective as the jelly’s bell rhythmically contracts and relaxes. Soft top-lit caustic shimmer, faint amber instrument LEDs in the distance, hushed uncanny mood, no cuts.
Model B better matches the prompt’s exact orbit feel and moon-jelly focus, with cleaner jelly detail, stronger caustic-lit aquarium mood, and more convincing continuous motion around the subject. Model A has a plausible aquarium scene and buoy, but it misses key prompt details like the lone pipefish and reads more like a gentle pan/dolly with weaker subject-centric orbit and less refined temporal progression. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt more closely with a translucent moon jelly beside an algae-wrapped rusted structure, visible bubbles, a pipefish-like background fish, and a tighter sense of circular perspective change while preserving the uncanny aquarium mood. Model A has pleasing underwater imagery, but it drifts from the specified setup with a different buoy/reef composition, extra jellyfish and fish, weaker exact-orbit feel, and less faithful subject staging.)
Reedbed Egret Tangle
At first light in a misty brackish reedbed, a single continuous 16:9 shot follows a chaotic but believable burst of wildlife activity: three little egrets stalk in different directions through shallow water, a pair of moorhens zigzag between reed stems, seven dragonflies skim and dart at different heights, and a tight school of tiny silver fish flickers just below the surface whenever the birds strike; every subject should keep its own clean body shape and path without blending, duplicating, or warping into others. The camera makes a gentle lateral tracking move at knee height along the waterline, keeping the cluster of animals in view as they independently change speed and direction. Cool dawn light, pearly reflections, damp atmosphere, tense but delicate mood, one continuous shot, no cuts.
Model B matches the prompt far better by clearly showing three little egrets, two moorhens, and multiple dragonflies moving independently in a coherent reedbed scene with stable forms and pleasing dawn reflections. Model A is atmospheric, but it underdelivers on the required animal count and diversity, and its subject continuity appears less faithful to the specified multi-animal choreography. (Second judge pass, order swapped — scores are the average of both: Model B matches the prompt much more closely by showing three little egrets, two moorhens, multiple dragonflies, and a coherent reedbed waterline composition with clean silhouettes and stable spacing across frames. Model A is moodier but misses key prompt elements like the full animal cluster and independent multi-subject action, and it shows weaker temporal consistency with a bird morphing into a different species-like form.)
Lighting transition
A living room at dusk as the sun sets over a few seconds: the warm light visibly dims and shifts to cool blue while a lamp flickers on, the transition smooth and believable, locked-off frame, 16:9.
Model B matches the prompt much better: it shows a locked-off living room with a clear dusk transition from warm sunlight to cool blue ambient light, plus a believable lamp turning on near the end. Model A is visually pleasing and temporally stable, but it behaves more like an outdoor patio scene and lacks the requested lamp flicker-on transition, so its prompt adherence is notably weaker. (Second judge pass, order swapped — scores are the average of both: Model B closely matches the prompt with a locked-off living room shot where warm sunset light smoothly dims into cool blue and the floor lamp turns on believably by the final frame. Model A is aesthetically pleasing, but it appears to be an outdoor patio rather than a living room, and it lacks the clear dusk-to-blue interior lighting transition and lamp flicker-on behavior requested.)
Matchup powered by OpenRouter.