Head to head: Gemini Omni Flash 1.1 Image to Video vs LTX-2.3 22B
Gemini Omni Flash 1.1 Image to Video vs LTX-2.3 22B
By Ryan Merket · Published
This wasn’t competitive. Across 12 judged runs, Gemini Omni Flash 1.1 Image to Video beat LTX-2.3 22B on every single task, with a 35.6 to 17.0 aggregate and a 95% confidence verdict.
Gemini Omni Flash 1.1 Image to Video didn’t just edge out LTX-2.3 22B; it swept it. The aggregate score gap — **35.6 to 17.0** — is huge, and the task record is even harsher: **12 wins to 0, with no ties**. Statistically, this lands as a **clear win at 95% confidence**, which matches what the task-level evidence shows. The pattern was consistent: Gemini was the model that actually executed the prompt. In **physics realism**, it gave judges a readable ice-cube drop, visible splash, believable buoyant bobbing, and coherent fizz. LTX often looked attractive in still frames, but the motion logic broke down — weaker splash cues, muddier cube behavior, and less convincing temporal progression. That same story repeated in **crowd motion**, where Gemini delivered a real overhead Tokyo-style scramble with independently moving pedestrians and stable geometry, while LTX slipped into ghosting, warped figures, synchronized movement, or simply the wrong scene composition. The gap was just as obvious on prompt fidelity and identity consistency. In **archivist identity lock**, Gemini repeatedly preserved the subject while also hitting the specifics: narrow records vault, wheeled ladder, teal 1947-B file, cardigan-and-blouse styling, dusty archival mood, and the requested profile-to-frontal camera progression. LTX was serviceable at basic continuity, but it kept missing the actual assignment — wrong wardrobe, wrong props, wrong atmosphere, and a subject design that drifted away from the archivist described in the prompt. And in **night kiln follow**, Gemini was the only model that reliably produced the shot being asked for: a shoulder-height lateral follow of a glassblower carrying a glowing amber vase on a blowpipe from furnace 3 toward the marver table, with coherent workshop lighting and background action. LTX repeatedly lost the vase, substituted a flame-like object, obscured the subject, or collapsed into blurrier, less stable motion. That’s the core editorial takeaway here: Gemini wasn’t merely prettier or smoother — it was far more faithful, far more readable, and far more in control of motion over time. **Final call: Gemini Omni Flash 1.1 Image to Video is the decisive winner. LTX-2.3 22B never seriously threatens it on this test set.**
Physics realism
A close-up of a single ice cube dropped into a tall glass of soda: it falls, splashes, bobs, and settles with believable buoyancy and fizzing bubbles rising naturally, soft kitchen light, 16:9.
Model A better matches the prompt with a clear close-up of a single ice cube entering soda, visible splash droplets, buoyant bobbing near the surface, and natural fizzing bubbles with consistent glass geometry and lighting. Model B is visually pleasing but the motion progression is less believable and less clearly readable as a drop-splash-bob-settle sequence, with weaker buoyancy cues and less prompt-specific action. (Second judge pass, order swapped — scores are the average of both: Model A better matches the prompt with a clearly visible ice cube interacting with the soda, convincing buoyant rise and settling near the surface, and natural fizzing bubbles in a soft kitchen setting. Model B looks visually plausible but the sampled motion reads less clearly as a drop-and-bob event, with weaker splash evidence and less convincing temporal progression of the cube's buoyancy.)
Archivist identity lock
A single continuous 16:9 shot in smooth cinematic slow motion: a municipal archivist named Sabela Quist walks from left to right through a narrow records vault, pulling a wheeled ladder with one hand and reading a faded teal file marked 1947-B with the other, while the camera performs a gentle forward dolly and slight arc from three-quarter profile to frontal close medium, keeping her perfectly framed under flickering fluorescent tubes and dusty skylight beams; her appearance must remain exactly unchanged from first frame to last — same short copper curls pinned with a silver triangle clip, same faint gap in the left eyebrow, same moss-green cardigan over a mustard blouse, same ink smudge on the right cuff, same oval tortoiseshell glasses — with no morphing, flicker, or wardrobe drift, creating a hushed, meticulous mood.
Model A matches the archivist prompt closely, with the narrow records vault, wheeled ladder, teal file marked 1947-B, and a coherent camera move from profile toward frontal framing while preserving the character’s look across frames. Model B is more stylized and stable enough, but it misses key setting and appearance details, lacks the dusty vault mood, and the subject design diverges notably from the specified curls, wardrobe, and overall archivist identity. (Second judge pass, order swapped — scores are the average of both: Model A matches the prompt far more closely: it clearly depicts the narrow records vault, wheeled ladder, dusty skylight beams, teal file marked 1947-A, and a smooth progression from profile toward frontal framing while preserving the character’s look. Model B has decent framing continuity, but it misses key scene details and the specified wardrobe/setting, and feels less faithful overall despite acceptable temporal stability.)
Night kiln follow
In a single continuous 16:9 shot, track a glassblower named Mirek striding sideways through a midnight workshop while carrying a glowing amber vase on a long steel blowpipe from furnace 3 to the marver table; the camera moves in a smooth lateral follow at his shoulder height, subtly adjusting distance to keep his face, hands, and the vase centered and tack-sharp as sparks drift, assistants cross in the deep background, orange kiln light pulses against cool blue shadows, reflections slide across the concrete floor, and the mood is tense, focused, and cinematic.
Model A closely matches the prompt with a shoulder-height lateral follow of a glassblower carrying a glowing amber vase on a blowpipe from furnace 3 toward the marver table, with coherent workshop lighting, reflections, and background assistants. Model B misses key prompt elements—the vase is absent or unclear, the setting and camera feel less like the specified cinematic kiln follow, and the imagery is blurrier and less temporally stable. (Second judge pass, order swapped — scores are the average of both: Model A matches the prompt far better: it clearly shows a continuous lateral follow of a glassblower carrying a glowing amber vase on a blowpipe from furnace 3 toward the marver table, with strong lighting contrast, floor reflections, and background assistants. Model B has some sideways walking and kiln color, but the object on the pipe is unclear or missing, the scene is blurrier and less coherent, and the motion/continuity feel less controlled.)
Crowd motion
A busy Tokyo scramble crossing seen from above, dozens of pedestrians crossing in different directions, each moving independently without merging or warping into one another, overcast daylight, 16:9.
Model A closely matches the prompt with an overhead Tokyo-style scramble crossing, many pedestrians moving independently, and stable geometry across frames; motion appears coherent and people do not noticeably merge or warp. Model B shows severe ghosting, sparse and inconsistent pedestrian motion, and strong temporal artifacts that break both realism and the requested independent crowd movement. (Second judge pass, order swapped — scores are the average of both: Model A clearly matches the prompt with an overhead Tokyo-style scramble crossing, many pedestrians moving independently in multiple directions, and stable, realistic crowd motion. Model B shows severe ghosting, sparse and warped figures, and poor temporal coherence, so it fails the key requirement that pedestrians not merge or distort.)
Matchup powered by OpenRouter.