Head to head: Wan vs OpenAI: GPT-5.4 Image 2

Wan vs OpenAI: GPT-5.4 Image 2

By · Published

RuntimeWire Head-to-Head: Head to head: Wan vs OpenAI: GPT-5.4 Image 2
RuntimeWire Head-to-Head matchup

Wan and OpenAI: GPT-5.4 Image 2 face off across eight image-generation tasks spanning documentary scenes, illustration, and precise spatial instructions. The comparison turns on a recurring trade-off: Wan’s more reliable widescreen framing versus OpenAI’s stronger scene detail and layout control.

OpenAI takes the matchup, winning six of eight tasks and leading 64.8 to Wan’s 54.8, with a tentative lean verdict. That edge comes from repeatedly making busy prompts feel more fully realized: its underpass scene sells the shared chess-table moment, its Art Nouveau tram stop has richer ironwork and ornament, and its bedroom layout puts the requested objects where they belong. Wan’s wins show where OpenAI’s advantage breaks down. In the missed-last-bus scene, Wan makes exhaustion and disappointment more legible and gets the 23:40 timetable into a faithful widescreen composition. Its perspective task is similarly stronger on the requested 16:9 aisle and convincing scale recession. Aspect ratio is OpenAI’s most consistent liability: several otherwise strong images arrive square when the prompt calls for widescreen. Both models also struggle with the plant-free reading nook, showing greenery outside despite the negation. And in the attribute-binding task, both get the colors and object relationships right, leaving framing—not basic instruction comprehension—as the meaningful distinction. **Final call: OpenAI wins clearly on this set, especially when a prompt demands rich scene detail or exact spatial arrangement. Wan is the better pick when widescreen fidelity and a sharply readable emotional beat matter most.**

Underpass chess crowd

A coherent multi-subject urban scene under the Cormorant Avenue underpass at late afternoon: five people interacting around a folding chess table—an elderly man in a flat cap reaching for a bishop, a teenage courier still wearing a teal bike helmet leaning in, a woman with a yellow grocery tote laughing while pointing at the board, a little boy on a scooter peeking between them, and a street saxophonist pausing nearby with his case open—balanced composition with layered depth, concrete pillars covered in faded stickers, passing buses glimpsed beyond, believable spacing, gestures, eye lines, and shared attention, documentary photojournalism feel, 16:9

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model A includes all five people and the requested details in a clear 16:9 underpass scene, with believable shared attention and buses visible beyond. Model B has a strong documentary mood and renders the avenue sign, but its square framing is less faithful to the requested format and the scooter is less clearly depicted. (Second judge pass, order swapped — scores are the average of both: Model B more vividly captures the shared chess-table interaction, including the boy with a scooter, the read | Anthropic: Claude Fable 5.1: Model A honors the 16:9 ratio and includes all…

Neon locksmith realism

Photorealistic documentary-style image of a midnight locksmith shop on Brackett Lane after light rain, the roll-up door half open, a woman in a charcoal wool coat examining a ring of newly cut keys beside a scratched aluminum workbench, believable mixed lighting from a cool LED strip inside and a magenta convenience-store sign across the street, physically accurate reflections on wet asphalt, brushed steel, tarnished brass, cracked varnished wood, damp denim, and human skin, natural shadow falloff, subtle steam from a street grate, 16:9

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model A more closely matches the requested 16:9 scene, with the half-open shop, clearly separated cool and magenta lighting, wet reflections, and visible steam. Model B has convincing locksmith details and a strong documentary mood, but misses the requested aspect ratio and its prominent signage is less cleanly rendered. (Second judge pass, order swapped — scores are the average of both: Model A better matches the requested 16:9 documentary composition, with the half-open roll-up door, readable | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame, corre…

Art Nouveau tram stop

A rainy twilight street scene in the named Art Nouveau style: a narrow tram stop on Rue Vespera in the fictional district of North Calder, with a green enamel kiosk labeled 'Line 7B', curling wrought-iron posters, iris and ginkgo motifs climbing the shelter glass, elongated elegant figures in plum and amber coats waiting beneath sinuous lamp standards, puddles reflecting peacock-blue and brass tones, every line fluid, ornamental, and unmistakably faithful to late-1890s Art Nouveau design, wide cinematic framing, 16:9

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model B captures the late-1890s Art Nouveau character more convincingly, with richly curling ironwork, botanical ornament, rainy reflections, and the requested plum-and-amber figures; it also includes the tram and additional location text. Model A has the requested wide framing and a clearly legible Line 7B kiosk, but its cleaner illustration style feels less distinctly Art Nouveau. (Second judge pass, order swapped — scores are the average of both: Model B more fully realizes the late-1890s Art | Anthropic: Claude Fable 5.1: Model B delivers a far more convincing late-189…

Negation

A cozy reading nook with an armchair, a stack of books, and a mug of tea by a window — with absolutely NO plants, NO lamps, and NO artwork or picture frames anywhere in the frame. Warm afternoon light, 16:9.

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model A matches the requested 16:9 framing and includes all the key nook elements, though visible potted plants outside the window violate the negation. Model B is especially cozy and polished, but also shows greenery through the window and misses the requested widescreen format. (Second judge pass, order swapped — scores are the average of both: Model B better preserves the plant-free reading nook, though distant foliage is visible outside its window; Model A has more conspicuous greenery by th | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame but c…

Missed last bus

Close, emotionally clear portrait of a transit worker sitting alone on a plexiglass bench at the deserted Elmridge night stop, convincingly conveying the specific emotion of exhausted disappointment after missing the last bus: shoulders dropped, jaw slack, eyes glossy but not crying, one hand still holding a folded timetable stamped '23:40', reflective orange safety vest half unzipped over a navy hoodie, sodium-vapor streetlight mixing with cold blue ad-panel glow across tired skin and creased fabric, shallow background of empty lanes and drizzle, documentary photojournalism realism, 16:9

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model A more clearly conveys exhausted disappointment through the slumped posture and weary expression, and it matches the requested 16:9 scene with a clearly visible 23:40 timetable. Model B is atmospheric and realistic, but its square framing and upward-looking expression are less aligned with the prompt. (Second judge pass, order swapped — scores are the average of both: Model B gives the more intimate, emotionally legible portrait and clearly shows the folded timetable stamped “23:40,” with | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame, a cle…

Perspective & scale

A photorealistic one-point-perspective shot down a long empty library aisle, shelves receding to a single vanishing point, books getting consistently smaller with distance and nothing warped or wrongly sized, warm overhead light, 16:9.

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model A more closely matches the requested 16:9 photorealistic aisle, with warm lighting, a clear single vanishing point, and convincing scale recession. Model B also has strong perspective and warm light, but its square framing misses the requested aspect ratio and its aisle feels more compressed and dramatic. (Second judge pass, order swapped — scores are the average of both: Model B creates a particularly strong, long one-point-perspective aisle, but its square framing misses the requested 16 | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame with…

Spatial layout

A clean isometric illustration of a bedroom: a bed against the LEFT wall, a round rug centered on the floor, a desk under the WINDOW on the back wall, and a floor lamp in the FRONT-RIGHT corner. Flat-vector style, consistent perspective.

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Model B places the bed, centered rug, desk directly beneath the window, and floor lamp clearly in their requested positions, with a cohesive isometric composition. Model A better matches the flat-vector style, but its rug is partly obscured and its desk is not clearly beneath the main window. (Second judge pass, order swapped — scores are the average of both: Model B places the bed, centered round rug, desk beneath the back-wall window, and floor lamp in the front-right corner as requested, with | Anthropic: Claude Fable 5.1: Model B places every element correctly—bed on t…

Attribute binding

A photorealistic tabletop scene: a matte green cube to the LEFT of a glossy red sphere, a small blue cylinder BEHIND both, and a single yellow rubber duck sitting ON TOP of the cube. Colors must stay on the correct objects. Soft studio light, 16:9.

Wan: Wan
OpenAI: GPT-5.4 Image 2: OpenAI: GPT-5.4 Image 2

OpenAI: GPT-6 Luna: Both models correctly bind the colors and place the duck on the cube, with the cylinder behind the other objects. Model A better matches the requested 16:9 framing, while Model B has a clean arrangement but is square. (Second judge pass, order swapped — scores are the average of both: Model A follows the requested 16:9 framing while clearly placing the duck on the green cube, the red sphere to its right, and the blue cylinder behind them. Model B is also polished and gets the object colors and p | Anthropic: Claude Fable 5.1: Model A honors the 16:9 framing and places colo…

Matchup powered by OpenRouter.