Head to head: Kling Video v3 Image to Video [Pro] vs Luma Ray 3.2 Image to Video
Kling Video v3 Image to Video [Pro] vs Luma Ray 3.2 Image to Video
By Ryan Merket · Published
This one is a genuine dead heat. Kling and Luma split the four tests 2–2, and the aggregate gap is so small that there’s only 50% confidence either model is truly ahead.
Kling Video v3 Image to Video [Pro] and Luma Ray 3.2 Image to Video finish essentially level here: 28.8 to 28.4, a margin too small to mean anything decisive. That is not a polite way of saying “close win.” It’s a statistical shrug. On this set, these models are effectively even. The split in strengths is clean. Kling is better when prompt adherence and subject identity really matter. In the **civet shrine orbit** test, it was the more disciplined model by a distance: it kept the civet near center, preserved the stacked black-stone shrine and surrounding details, and delivered the requested clockwise orbit without wandering off prompt. It also edged the **physics realism** task in the combined verdict, where its soda-and-ice behavior read as cleaner and more believable overall, even if one judge found Luma’s splashier impact more faithful to the requested drop event. Luma, meanwhile, took the tests where scene staging and camera feel carried more weight. In **reedbed shadow drift**, it did the basic editorial work Kling missed: all three numbered poles were visible, the marsh held as a static observational view, and the lighting shift into stormier tones landed more convincingly. In **single continuous shot**, Luma was simply more persuasive as a moving camera model, with a stronger forward glide through the cathedral, better continuity, and a richer candlelit atmosphere. So the practical read is straightforward: Kling is the safer pick if you need tighter object fidelity and stronger obedience to a dense image-to-video brief; Luma is the better bet if you want mood, motion, and cinematic flow. But those are tendencies, not a knockout. Across these four tasks, neither model established a durable lead. **Final call: too close to call. Kling and Luma are effectively tied in this matchup.**
Physics realism
A close-up of a single ice cube dropped into a tall glass of soda: it falls, splashes, bobs, and settles with believable buoyancy and fizzing bubbles rising naturally, soft kitchen light, 16:9.
Model A better matches the requested believable buoyancy and natural fizz: the ice cube rises and settles near the surface with consistent shape, plausible liquid motion, and fine bubbles in soft kitchen lighting. Model B has a dramatic splash, but the cube appears oversized and emerges unrealistically high above the rim, with less believable physics and weaker temporal coherence. (Second judge pass, order swapped — scores are the average of both: Model B better matches the prompt’s key physics beats: a visible drop, energetic splash, bobbing rise, and a final partially floating settle with plausible soda fizz. Model A is cleaner and attractive, but the motion progression appears gentler and less convincing for an actual drop-and-splash event, with weaker impact dynamics and buoyancy storytelling.)
Reedbed shadow drift
A one continuous 8-second shot in 16:9 of a piebald marsh harrier gliding low over a copper-colored reedbed dotted with three weathered survey poles marked 11, 14, and 19, while two teal burst upward from the water and skim away; the camera is locked off in an observational static frame from the edge of a brackish lagoon, and midway through the clip a ragged thundercloud slides over the late-afternoon sun so the whole marsh transitions smoothly from warm amber glow to muted steel-gray, with the bird and reeds staying believable in color and contrast as the wind ripples the seed heads; mood tense and hushed.
Model B better matches the prompt by clearly showing all three survey poles labeled 11, 14, and 19 in a static marsh scene with a convincing warm-to-cool stormy lighting shift, while Model A misses pole 11 and the bird reads less like a piebald marsh harrier. Model A has decent continuity, but Model B is more compositionally faithful and atmospheric overall despite the bird looking somewhat like a heron and the teal action not being evident in either entry. (Second judge pass, order swapped — scores are the average of both: Model B matches more of the prompt structure: a static observational marsh view with three numbered poles visible and a plausible darkening, muted atmosphere over the reedbed, though the featured bird reads more like a heron than a piebald marsh harrier and the teal action is unclear. Model A is cleaner visually, but it misses key prompt elements by showing only two poles, lacking the specified numbering set and cloud-driven tonal transition, and the bird again does not convincingly match a marsh harrier.)
Orbit around civet shrine
A one continuous 7-second shot in 16:9 of a masked palm civet stepping lightly around a tiny hillside fox shrine made of stacked black stones, nosing at a red persimmon and then flicking its tail as moths circle a smoking incense coil; execute a smooth clockwise orbit around the subject, maintaining the civet near center frame the entire time as the background reveals terraced tea rows, a rusted rain drum, and a painted wooden tag reading K-27, with cool blue pre-dawn light and a faint lantern glow from off-frame left brushing the stones; mood secretive and uncanny.
Model A matches the requested civet, stacked black-stone shrine, persimmon, incense coil, tea terraces, rusted drum, and K-27 tag, and the sampled frames suggest a coherent clockwise orbit with the animal kept near center. Model B is more visually atmospheric but departs heavily from the prompt by featuring foxes instead of a masked palm civet, a different shrine structure, and weaker subject-centric framing and action fidelity. (Second judge pass, order swapped — scores are the average of both: Model A matches the prompt far better: it shows a civet-like animal near center with a stacked black-stone shrine, persimmon, incense coil, tea terraces, rusted drum, and K-27 tag while sustaining a plausible clockwise orbit and coherent scene layout. Model B has stronger mood lighting in places, but it introduces an obvious fox-like second animal and a ring-shaped stone structure instead of the tiny shrine, which hurts adherence and consistency.)
Single continuous shot
One unbroken take gliding slowly through a candlelit cathedral from the entrance toward the altar, no cuts, jumps, or transitions, dust and warm light in the air, 16:9.
Model B better matches the requested single continuous glide through a candlelit cathedral, with stronger forward camera progression, richer warm haze and dust-lit atmosphere, and more cohesive cinematic framing. Model A is visually appealing and consistent, but the motion feels subtler and the aisle of floor candles reads slightly staged rather than a natural cathedral glide toward the altar. (Second judge pass, order swapped — scores are the average of both: Model B better matches the prompt with a centered, slow forward glide from the cathedral entrance toward the altar, maintaining warm candlelit atmosphere and visible dust-filled light beams with strong continuity across frames. Model A is visually appealing and consistent, but it feels less like a glide from the entrance and more like a staged interior aisle shot, with cooler lighting and a less convincing sense of the specified camera movement.)
Matchup powered by OpenRouter.