Head to head: Cosmos 3 Super vs Luma Photon Flash Reframe
Cosmos 3 Super vs Luma Photon Flash Reframe
By Ryan Merket · Published
This matchup pits broad prompt fidelity and compositional control against a narrower talent for keeping object attributes straight. Repeated trials expose a decisive difference in reliability across complex scenes, typography, anatomy, and reflections.
Cosmos 3 Super wins this decisively: **59.7 to 49.9**, with **95% statistical confidence**. It took **21 of 24 task trials**, leaving no ties and no plausible argument that the result came down to judging noise. Its advantage was breadth. Cosmos repeatedly delivered the requested 16:9 framing and preserved complicated scene requirements in the boxing gym, canal relay, goalkeeper, tunnel, reflections, typography, and hand-anatomy tests. It was not flawless—mirror geometry, minor role details, and some reflective elements slipped—but it consistently retained more of each prompt’s essential structure. Luma Photon Flash Reframe’s three wins all came from the same place: attribute binding. It correctly put the yellow duck on the green cube while Cosmos repeatedly attached it to the wrong object. That is a meaningful weakness for Cosmos, but Luma could not convert its precision there into broader competitiveness; square framing, omitted actions, cropped subjects, theatrical expressions, and altered text layouts repeatedly cost it elsewhere. **Final call: Cosmos 3 Super is the clear winner. Luma owns the attribute-binding upset, but Cosmos is the substantially more dependable model across the full benchmark.**
Attribute binding
A photorealistic tabletop scene: a matte green cube to the LEFT of a glossy red sphere, a small blue cylinder BEHIND both, and a single yellow rubber duck sitting ON TOP of the cube. Colors must stay on the correct objects. Soft studio light, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B correctly places the yellow duck on the green cube while preserving the requested object colors and spatial ordering, though its framing is not 16:9 and the cylinder is oversized. Model A has stronger widescreen composition and studio presentation, but incorrectly places the duck on the red sphere, violating the key attribute-binding instruction. (Second judge pass, order swapped — scores are the average of both: Model B correctly binds the yellow duck to the green cube and preserves the | Anthropic: Claude Fable 5.1: Model A places the duck on the red sphere…
Boxing gym reflections
Ultra-detailed nighttime interior, 16:9 — in a vintage boxing gym with floor-to-ceiling street-facing windows during a thunderstorm, a boxer wraps her hands beside a chrome heavy-bag stand while her trainer watches from behind; the scene must accurately handle reflections and transparency: the windows show layered reflections of the ring ropes and ceiling fans while also revealing a neon pharmacy sign, rain-streaked taxis, and pedestrians outside; a wall mirror at an angle captures the boxer from the side without impossible geometry, a puddle on the rubber floor mirrors the red EXIT sign and overhead strip lights, and the chrome stand produces distorted but coherent highlights; moody mixed lighting from cool lightning flashes, warm tungsten gym lamps, and magenta neon.


OpenAI: GPT-5.6 Sol Pro: Model A closely follows the requested 16:9 scene, including hand wrapping, a chrome stand, coherent angled mirror imagery, rainy street visibility, and a puddle reflecting the EXIT sign, though the trainer and ceiling fans are missing and the pharmacy text is malformed. Model B has attractive storm lighting and street detail but misses the aspect ratio, hand-wrapping action, chrome stand, trainer, ring-rope reflections, and key puddle and mirror requirements. (Second judge pass, order swapped — | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame,…
Canal relay scramble
Documentary photojournalism style, 16:9 — at a rain-slick urban canal race in Ghent at blue hour, four athletes and two support crew interact in one tightly composed scene: a female kayaker in a lime helmet reaches from her boat to grab a relay baton from a sprinting teammate on the stone quay, a third teammate kneels to steady the kayak with one hand while holding a coiled rescue rope, a fourth runner vaults a low barrier in the background, and a medic in a yellow poncho leans in beside a bicycle-mounted camera operator; wet cobblestones, canal water, dock lines, orange safety cones, and scattered event signage create depth, with balanced spacing, believable overlap, and crisp motion frozen under mixed sodium streetlights and cool evening sky.


OpenAI: GPT-5.6 Sol Pro: Model A closely matches the requested 16:9 relay scene, clearly showing the baton exchange, kneeling rope handler, approaching runner, medic, bicycle-mounted camera operator, wet quay, and mixed blue-hour lighting, though some athlete roles are conflated and the barrier vault is unclear. Model B is visually polished but uses the wrong aspect ratio and omits the central baton handoff, sprinting teammate, and background vaulting action. (Second judge pass, order swapped — scores are the average of | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame…
Defeat in the tunnel
Cinematic close-up portrait, documentary photojournalism look — inside a concrete stadium service tunnel after an upset loss, a young male middle-distance runner in a navy singlet sits on a folding crate, bib half-peeled, staring just past camera with unmistakable restrained devastation: watery eyes, clenched jaw, one cheek trembling, lips pressed as if trying not to cry, shoulders collapsed but not theatrical; his coach's hand enters frame resting lightly on his back, fluorescent overhead fixtures create harsh top light and deep under-eye shadows, and the shallow depth of field isolates the expression against blurred orange extension cords and scuffed cinderblock walls.


OpenAI: GPT-5.6 Sol Pro: Model A closely captures the requested restrained devastation, off-camera gaze, documentary framing, half-peeled bib, supportive hand, harsh tunnel lighting, and blurred orange cords. Model B is visually striking but depicts overt, direct-to-camera crying rather than restrained emotion, while its garbled bib and duplicated uniform logos reduce realism. (Second judge pass, order swapped — scores are the average of both: Model A better captures the restrained, just-past-camera devastation and docu | Anthropic: Claude Fable 5.1: Model A delivers the requested close-up wi…
Legible multi-line text
A minimalist event poster with three lines of crisp, correctly-spelled text stacked and centered: 'NIGHT MARKET' large on top, 'Fridays · 6–11pm' in the middle, 'Riverside Pier 4' at the bottom, on a deep navy background, clean sans-serif, subtle grain.


OpenAI: GPT-5.6 Sol Pro: Model A renders all requested text correctly in a clean sans-serif with strong hierarchy and compact centering, though the headline wraps across two visual lines. Model B uses a serif typeface, omits the middle dot, and breaks the three requested text lines into six, weakening both adherence and composition. (Second judge pass, order swapped — scores are the average of both: Model A renders all requested text correctly in a clean sans-serif with strong hierarchy, centered alignment, a deep navy | Anthropic: Claude Fable 5.1: Model A renders all three lines correctly s…
Goalkeeper glove check
Hyperreal editorial sports photograph — a women's handball goalkeeper standing full-body on the goal line during a timeout, clearly visible from head to toe, adjusting the wrist strap of her left glove with the fingers of her right hand while holding a resin-marked ball against her hip with the crook of her arm; natural anatomy, accurate finger count, realistic knuckles, tendons, and proportionate limbs are essential; she wears a teal jersey with the number 27, black knee pads, and one sock slightly slouched, with sweat on her forearms and a focused posture; shot from slightly below eye level with soft arena sidelights and a dim crowd behind her.


OpenAI: GPT-5.6 Sol Pro: Model A closely matches the requested full-body framing, teal number 27 jersey, knee pads, ball placement, and glove-adjustment action, though the hands are somewhat muddled and the court context is inaccurate. Model B is more photographically polished, but it is cropped below the knees, shows number 21, and omits the crucial two-handed wrist-strap adjustment. (Second judge pass, order swapped — scores are the average of both: Model A clearly shows the goalkeeper head to toe, correctly renders j | Anthropic: Claude Fable 5.1: Model A delivers the full head-to-toe fram…
Reflections & glass
A photorealistic close-up of a chrome teapot on a polished dark marble counter next to a glass of water; the window and a red apple on the counter must be correctly reflected in BOTH the chrome and the water's surface, dramatic side light, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model A delivers the requested 16:9 framing, stronger composition, and more convincing photorealism, though the apple is not visibly reflected in the water surface. Model B better suggests the apple in both reflective subjects, but its square format, ambiguous double waterline, and less natural reflections weaken the result. (Second judge pass, order swapped — scores are the average of both: Model A better satisfies the requested 16:9 format and clearly shows both the window and apple reflected | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame w…
Hands & anatomy
A natural candid photo of a person tying a friendship bracelet, both hands clearly visible with anatomically correct fingers and a plausible grip on the threads, soft daylight, shallow depth of field, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B better fulfills the anatomy-focused request with exactly two convincing hands and a plausible thread-working grip, though it misses the requested 16:9 framing. Model A has attractive lighting and composition but appears to show an extra third hand, significantly weakening prompt adherence. (Second judge pass, order swapped — scores are the average of both: Model B shows exactly two anatomically convincing hands with a plausible thread grip, natural light, and strong shallow focus, though | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame…
Matchup powered by OpenRouter.