Head to head: Hidream I1 Full vs Ideogram V4.5 Edit
Hidream I1 Full vs Ideogram V4.5 Edit
By Ryan Merket · Published
A comparison of two image models across eight demanding scenes, from architectural perspective and rainy reflections to exact object counts and spatial relationships. The results turn on a recurring trade-off between faithful framing and precise execution of the brief.
Ideogram V4.5 Edit takes the matchup, but its edge is earned through instruction-following rather than a clean sweep: it wins five tasks to HiDream I1 Full’s two, with one tie. The aggregate is 57.5 to 55.0, and the limited confidence makes this a clear verdict—not a blowout. HiDream’s strongest argument is composition. It repeatedly honors the requested 16:9 frame, winning the library perspective test with a balanced aisle and consistent scale, and edging the negation prompt by avoiding some of the unwanted decor that appears in Ideogram’s image. The embroidery atelier is a genuine draw: HiDream better nails the coat, waistcoat, and hand-finishing, while Ideogram brings useful scene details of its own. Ideogram wins where the brief demands specific content. Its rainy arcade has the named salon sign and convincing reflections; its Art Nouveau millinery image gets the full sign text and stronger poster styling; and its stair scene better conveys one-point depth and scale. It also gets the object relationships right in the attribute-binding test, where HiDream reverses the cube and sphere. In the cup-count test, both show seven; Ideogram’s cleaner marble presentation beats HiDream’s tray, despite both models missing the requested wide format in different ways across the comparison. The recurring cost of Ideogram’s wins is framing: its square crops often ignore the 16:9 instruction that HiDream handles reliably. But across this set, accurate text, scene details, and spatial relationships matter more than aspect-ratio discipline alone. **Verdict: Ideogram V4.5 Edit wins clearly on the evidence, while HiDream I1 Full remains the stronger choice when wide composition is the priority.**
Perspective & scale
A photorealistic one-point-perspective shot down a long empty library aisle, shelves receding to a single vanishing point, books getting consistently smaller with distance and nothing warped or wrongly sized, warm overhead light, 16:9.


OpenAI: GPT-6 Luna: Model A better matches the requested 16:9 framing and presents a convincing, balanced library aisle with consistent scale and warm overhead lighting. Model B has a stronger dramatic vanishing point, but its square framing and exaggerated depth make it less faithful overall. (Second judge pass, order swapped — scores are the average of both: Model B has a dramatic centered vanishing point and warm light, but its square framing misses the requested 16:9 format and the shelves feel more distorted. | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame with c…
Rainy arcade reflections
Under a glass-roofed shopping arcade during a spring rain, a model in a saffron trench coat pauses beside a curved display window of a handbag salon, with transparent umbrellas, wet black-and-cream mosaic tiles, beveled glass, mirrored brass columns, and chrome café chairs all reflecting and refracting the scene correctly; behind the glass are coral leather satchels and a small sign reading "Atelier Nivelle", cinematic overcast light with warm shop glow, precise optical realism, 16:9.


OpenAI: GPT-6 Luna: Model B better captures the rainy arcade atmosphere, curved handbag display, transparent umbrellas, warm shop glow, and legible “Atelier Nivelle” sign, with especially convincing glass reflections. Model A has the requested wide composition and striking mosaic floor, but omits the salon’s specified sign and feels less optically atmospheric; Model B’s main drawback is its square format rather than 16:9. (Second judge pass, order swapped — scores are the average of both: Model A better matches the | Anthropic: Claude Fable 5.1: Model A delivers the 16:9 format, glass roof, c…
Grand stair scale test
A fashion editorial shot inside the vast observatory hall of the fictional Helior Crown Terminal, where a model in a silver pleated jumpsuit descends a monumental staircase toward the camera while, far above, tiny visitors lean along a sixth-floor balustrade beneath a huge clock and ribbed skylight; the image must preserve strict one-point perspective, believable scale between foreground handrails, midground mannequins, and distant ticket kiosks, with long late-afternoon shadows and clean architectural lines, 16:9.


OpenAI: GPT-6 Luna: Model B better conveys the requested scale and one-point depth, with foreground handrails, distant kiosks, visitors along the elevated balustrade, and late-afternoon shadows. Model A has a more monumental editorial staircase and the requested wide framing, while Model B’s square crop misses the 16:9 specification. (Second judge pass, order swapped — scores are the average of both: Model B more closely matches the silver pleated jumpsuit, descending staircase, kiosks, and strict one-point scale c | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame, a cl…
Art Nouveau millinery
A twilight boutique window displaying three avant-garde felt cloche hats on slender bronze stands, illustrated in authentic Art Nouveau style with whiplash curves, sinuous floral borders, elongated elegant figures reflected faintly in the glass, muted peacock blues, olive green, aubergine, and warm gold, decorative hand-lettered sign reading "Maison Verlaine, Rue des Cygnes 14", composition like a 1903 fashion poster with flowing ornamental framing, soft lamplight and 16:9.


OpenAI: GPT-6 Luna: Model B better captures the Art Nouveau poster aesthetic with flowing floral framing and renders the full requested sign accurately, while Model A has a more suitable wide format and convincing cloche shapes but omits the address and feels more like a realistic storefront photograph. Model B’s near-square composition misses the 16:9 request, but its stronger overall style and text adherence make it the better match. (Second judge pass, order swapped — scores are the average of both: Model B bett | Anthropic: Claude Fable 5.1: Model B renders the full sign text correctly, u…
Negation
A cozy reading nook with an armchair, a stack of books, and a mug of tea by a window — with absolutely NO plants, NO lamps, and NO artwork or picture frames anywhere in the frame. Warm afternoon light, 16:9.


OpenAI: GPT-6 Luna: Model A has the requested cozy nook, warm afternoon light, and 16:9 framing, but prominently includes many plants. Model B also includes plants and additionally shows a lamp and framed artwork, while missing the requested 16:9 format. (Second judge pass, order swapped — scores are the average of both: Model B includes plants, lamps, and framed artwork, while Model A still has many plants but avoids lamps and picture frames and better matches the requested 16:9 format. Model A also has a polished | Anthropic: Claude Fable 5.1: Both fail the negation badly: Model A includes…
Embroidery atelier detail
A close, sharply focused scene inside a couture workroom where a tailor in a plum waistcoat hand-finishes a midnight-green opera coat draped over a dress form, every surface rendered crisply: dense mohair nap, silk jacquard lining, tiny seed-pearl embroidery shaped like storm swallows, chalk marks on the cutting table, brushed steel shears, loose threads, a wicker basket of folded herringbone wool, morning north-window light grazing across the textures, intimate editorial realism, 16:9.


OpenAI: GPT-6 Luna: Model A better matches the midnight-green coat, plum waistcoat, close hand-finishing action, and pearl swallow embroidery, with a crisp editorial feel. Model B includes the dress form and basket, but the coat reads bright teal, the bird motifs are less convincing, and the square framing misses the requested 16:9 composition. (Second judge pass, order swapped — scores are the average of both: Model A better matches the requested wide editorial workroom scene, with a plum-waistcoated tailor hand-f | Anthropic: Claude Fable 5.1: Model A delivers the correct 16:9 frame, a plum…
Exact counting
A clean overhead flat lay on white marble of EXACTLY seven identical espresso cups arranged in a single straight row, each clearly separate and individually countable, even soft lighting, 16:9.


OpenAI: GPT-6 Luna: Model A shows exactly seven distinct cups in a straight row and matches the requested 16:9 framing, though the dark tray departs from the clean marble-only flat lay. Model B is softer and uncluttered, but its square format misses the requested aspect ratio. (Second judge pass, order swapped — scores are the average of both: Model B shows exactly seven clearly countable cups on marble, but the square framing misses the requested 16:9 format. Model A uses the requested landscape format and arrange | Anthropic: Claude Fable 5.1: Both Model A and Model B deliver exactly seven…
Attribute binding
A photorealistic tabletop scene: a matte green cube to the LEFT of a glossy red sphere, a small blue cylinder BEHIND both, and a single yellow rubber duck sitting ON TOP of the cube. Colors must stay on the correct objects. Soft studio light, 16:9.


OpenAI: GPT-6 Luna: Model B correctly places the green cube to the left of the red sphere, with the duck on top and the blue cylinder behind; Model A reverses the cube-sphere order. Model A better matches the requested 16:9 framing, while Model B otherwise has a clear, polished composition. (Second judge pass, order swapped — scores are the average of both: Model B correctly places the green cube to the left of the glossy red sphere, with the duck on top and the blue cylinder behind; Model A reverses the cube and s | Anthropic: Claude Fable 5.1: Model B gets the key spatial binding right (gre…
Matchup powered by OpenRouter.