Head to head: AuraFlow vs Ideogram V4.5 Edit

AuraFlow vs Ideogram V4.5 Edit

By · Published

RuntimeWire Head-to-Head: Head to head: AuraFlow vs Ideogram V4.5 Edit
RuntimeWire Head-to-Head matchup

This matchup tests whether visual polish can survive exacting demands around layout, counting, text, color, anatomy, and exclusions. The result turns on consistent instruction-following rather than isolated standout images.

Ideogram V4.5 Edit wins decisively: **55.5 to 41.3**, with **17 task wins against AuraFlow’s 5**, plus two ties. The statistical verdict confirms this is a clear result at **limited confidence**, not a marginal lead inflated by a few favorable prompts. Ideogram’s advantage was broad and practical. It repeatedly nailed the seven-cup count and single-row arrangement, handled the bedroom layouts more accurately, reproduced the “ZONE 7C” parking-meter sticker more reliably, and assembled the color-bound kiosk objects with fewer substitutions or extras. It also came closer on the empty delivery lane, where AuraFlow repeatedly inserted prohibited cars and motorbikes. AuraFlow’s strongest counterpunch was the subway violinist: it won all three versions by preserving the requested mid-thigh framing, complete face, station context, and visible hands, while Ideogram cropped too tightly and produced malformed anatomy or instrument geometry. AuraFlow also picked up isolated wins on negation and restricted palette, but neither category was a convincing strength—both models frequently ignored exclusions, exceeded the specified colors, introduced gradients, and missed the requested 16:9 framing. The recurring aspect-ratio failures keep Ideogram from looking meticulous, and its human anatomy remains a liability. But AuraFlow’s counting errors—showing 10, 15, or 20 cups instead of seven—and repeated violations of core scene constraints are more damaging than Ideogram’s imperfections. **Final call: Ideogram V4.5 Edit is the clear winner. Its instruction-following is materially more dependable across spatial, textual, counting, object, and scene-compliance tasks.**

Spatial layout

A clean isometric illustration of a bedroom: a bed against the LEFT wall, a round rug centered on the floor, a desk under the WINDOW on the back wall, and a floor lamp in the FRONT-RIGHT corner. Flat-vector style, consistent perspective.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model B clearly places the bed along the left wall and the floor lamp in the front-right corner, with a more coherent bedroom composition, though the desk is beside rather than under the window and the rug is slightly off-center. Model A centers the rug and places the bed on the left, but the window, desk, and lamp placements substantially miss the requested layout. (Second judge pass, order swapped — scores are the average of both: Model B clearly places the bed along the left wall and the floo | Anthropic: Claude Fable 5.1: Model A places the bed against the left wa…

Negation

A cozy reading nook with an armchair, a stack of books, and a mug of tea by a window — with absolutely NO plants, NO lamps, and NO artwork or picture frames anywhere in the frame. Warm afternoon light, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model A better captures the warm afternoon atmosphere and includes the requested chair, books, tea, and window, but it flagrantly violates every negation with plants, lamps, and framed artwork. Model B is technically cleaner and well composed, yet contains even more prominent plants and artwork, also includes a lamp, and neither model follows the 16:9 requirement. (Second judge pass, order swapped — scores are the average of both: Both outputs severely violate the negation requirements and the r | Anthropic: Claude Fable 5.1: Both entries comprehensively fail the nega…

Restricted palette

A flat-vector illustration of a mountain campsite at dusk using ONLY four colors — cream, burnt orange, teal, and charcoal — with no gradients and no stray off-palette colors, balanced composition, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model B has the stronger, more balanced flat-vector composition and avoids the conspicuous gradients seen in Model A. Both models violate the exact four-color restriction by using multiple shades and miss the requested 16:9 aspect ratio, but Model B is closer stylistically and technically. (Second judge pass, order swapped — scores are the average of both: Model B has the stronger balanced campsite composition and cleaner flat-vector treatment, though it is square and uses several extra shades b | Anthropic: Claude Fable 5.1: Model A violates the restricted palette he…

Rainy parking-meter macro

Documentary photojournalism style macro close-up of a rain-speckled turquoise parking meter on Eldridge Cut, filling most of the frame, with a tiny peeling sticker reading "ZONE 7C" and a single copper coin half-inserted in the slot; shot at street level with a true macro lens effect, razor-sharp droplets and scratched paint, extreme shallow depth of field, the background dissolving into soft neon bokeh from a late-night noodle shop and a distant crosswalk signal, cinematic sodium-vapor and magenta reflections, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model B more convincingly delivers the macro framing, extreme shallow depth of field, neon reflections, and correctly rendered “ZONE 7C” sticker, while Model A’s sticker incorrectly reads “ZINE” and its centered framing feels less photographic. Neither model clearly shows the coin half-inserted in the slot or provides the requested 16:9 aspect ratio, though Model B comes closer overall. (Second judge pass, order swapped — scores are the average of both: Model B closely captures the macro framing | Anthropic: Claude Fable 5.1: Model B renders the exact 'ZONE 7C' sticke…

Color-bound kiosk still life

Top-down flat lay product-style photograph on the stainless-steel counter of a cramped midnight newsstand kiosk: a matte cobalt-blue ceramic espresso cup, a translucent amber glass ashtray, a brushed silver aluminum lighter, a moss-green waxed-canvas coin pouch, a crimson lacquered chopstick sleeve, and an ivory paper receipt curled at one corner; each object clearly separated with its own correct color and material, no color bleeding between items, crisp overhead fluorescent lighting with subtle urban grime and realistic shadows, highly detailed, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model B includes the requested six distinct objects with substantially more accurate colors and materials, stronger separation, and a cleaner product-style arrangement, though the amber ashtray resembles a tumbler and the image is not 16:9. Model A is technically polished but adds several extraneous items, uses a clear rather than amber ashtray, and misrepresents the pouch and chopstick sleeve. (Second judge pass, order swapped — scores are the average of both: Model B includes all requested obj | Anthropic: Claude Fable 5.1: Model B renders all six specified objects…

Exact counting

A clean overhead flat lay on white marble of EXACTLY seven identical espresso cups arranged in a single straight row, each clearly separate and individually countable, even soft lighting, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model A shows fifteen cups in a three-row grid, fundamentally failing the exact-count and single-row requirements. Model B clearly presents exactly seven separate cups in one straight row on marble, though the output is square rather than 16:9 and the cups vary slightly in appearance. (Second judge pass, order swapped — scores are the average of both: Model B correctly shows exactly seven clearly separated espresso cups in one straight row on white marble, though the image is square rather than | Anthropic: Claude Fable 5.1: Model A shows fifteen cups in a 3x5 grid, b…

Subway violinist hands

Documentary photojournalism image of a young woman busking on the downtown platform at Bracken Square station, shown from mid-thigh up, standing naturally in a charcoal coat and rust scarf while playing a weathered violin; both hands fully visible and anatomically correct, left hand accurately fingering the strings and right hand holding the bow with five natural fingers, realistic wrists, shoulders, and facial proportions, commuters blurred in the distance, cool platform fluorescents mixed with warm train-door light, decisive candid composition, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model A better matches the requested mid-thigh documentary composition, fully visible subject, station context, and complete hands, though the bow grip remains unnatural and the image is not 16:9. Model B has appealing depth and lighting, but its tight crop cuts off much of the face and body, while the violin geometry and fingering hand are visibly malformed. (Second judge pass, order swapped — scores are the average of both: Model A better matches the requested mid-thigh framing, fully shown fa | Anthropic: Claude Fable 5.1: Model A shows the full face and setting bu…

Empty delivery lane

Early-morning urban street scene in documentary photojournalism style: a narrow delivery lane behind the fictional Marrow & Finch printworks, wet cobblestones, stacked bread crates, steam rising from a basement vent, faded murals, orange safety cones, overhead laundry lines, and a lone cyclist walking a bicycle across the frame — but absolutely no cars, no vans, no buses, and no motorbikes anywhere in the scene; soft overcast light, deep perspective down the alley, 16:9.

AuraFlow: AuraFlow
Ideogram V4.5 Edit: Ideogram V4.5 Edit

OpenAI: GPT-5.6 Sol Pro: Model B closely fulfills the scene requirements with no motor vehicles, a person walking a bicycle, wet cobblestones, crates, steam, murals, cones, laundry lines, and strong alley perspective. Model A is visually polished but prominently includes several prohibited vehicles, shows the cyclist riding, renders the printworks sign poorly, and neither model follows the requested 16:9 aspect ratio. (Second judge pass, order swapped — scores are the average of both: Model B closely follows the request | Anthropic: Claude Fable 5.1: Model A violates the core negative constra…

Matchup powered by OpenRouter.