Head to head: Longcat Image vs Ideogram V4.5 Edit
Longcat Image vs Ideogram V4.5 Edit
By Ryan Merket · Published
We tested two image models on eight prompts spanning anatomy, reflective surfaces, precise object relationships, typography, composition, style, and palette control. The comparison rewards not just attractive images, but how faithfully each model follows the brief.
Longcat Image takes the matchup with a 58.4 to 51.9 aggregate score, winning six tasks to Ideogram V4.5 Edit’s one, with one tie. The limited confidence verdict makes this a decisive result—not a photo finish—and the task breakdown shows why: Longcat was more reliable at honoring the specific constraints that defined each prompt. Its clearest advantages came in spatial layout, named art style, and multi-line text. Longcat put the bedroom desk beneath the window, evoked ukiyo-e with muted indigo and cream, and preserved the poster’s requested wording and capitalization. It also edged the reflections and negation prompts, where the requested widescreen framing and fewer unwanted scene elements mattered more than polish alone. Neither model was flawless: the cozy nook images both included prohibited decor, and the tabletop object-relations task was a tie. Ideogram’s win was meaningful: its bracelet image showed both hands with a more plausible grip and clearer anatomy. It also produced especially legible reflections in the chrome teapot, but the square crop missed the requested 16:9 format. Longcat’s win on restricted palette likewise wasn’t a clean execution of the brief—its scene strayed into extra colors—yet Ideogram’s square composition and broader range of orange and yellow shades strayed further. **Verdict: Longcat Image wins decisively on prompt fidelity across a varied set of image tasks. Ideogram V4.5 Edit has the stronger hand-and-anatomy result, but that isolated strength doesn’t overcome Longcat’s broader consistency.**
Hands & anatomy
A natural candid photo of a person tying a friendship bracelet, both hands clearly visible with anatomically correct fingers and a plausible grip on the threads, soft daylight, shallow depth of field, 16:9.


OpenAI: GPT-6 Luna: Model B shows both hands clearly with a more natural, plausible grip on the colorful threads and convincing soft daylight. Model A matches the requested widescreen framing, but its fingers and grip around the bracelet look less anatomically clear. (Second judge pass, order swapped — scores are the average of both: Model B more clearly shows one person’s two hands holding the bracelet threads, with convincing anatomy and a natural shallow-focus look. Model A has a stronger wide composition, but i | Anthropic: Claude Fable 5.1: Model B delivers a single person with both hand…
Reflections & glass
A photorealistic close-up of a chrome teapot on a polished dark marble counter next to a glass of water; the window and a red apple on the counter must be correctly reflected in BOTH the chrome and the water's surface, dramatic side light, 16:9.


OpenAI: GPT-6 Luna: Model A better matches the requested close-up, dramatic 16:9 framing, with the chrome teapot, glass, apple, and polished dark counter all clearly present. Model B shows especially legible window and apple reflections in the chrome, but its square framing departs from the requested format and the reflections on the water surface are not clearly conveyed in either image. (Second judge pass, order swapped — scores are the average of both: Model A better matches the requested 16:9 composition and pl | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 frame, dram…
Negation
A cozy reading nook with an armchair, a stack of books, and a mug of tea by a window — with absolutely NO plants, NO lamps, and NO artwork or picture frames anywhere in the frame. Warm afternoon light, 16:9.


OpenAI: GPT-6 Luna: Model A better captures the requested cozy nook, with a steaming mug of tea, books, an armchair, and warm afternoon light, though it includes flowers and a visible picture frame. Model B also has the requested nook elements, but prominently includes plants, a lamp, and framed artwork, and is not widescreen. (Second judge pass, order swapped — scores are the average of both: Model B depicts the requested nook but prominently includes many plants, a lamp, and framed artwork, violating all three ex | Anthropic: Claude Fable 5.1: Model A delivers the armchair, book stacks, ste…
Attribute binding
A photorealistic tabletop scene: a matte green cube to the LEFT of a glossy red sphere, a small blue cylinder BEHIND both, and a single yellow rubber duck sitting ON TOP of the cube. Colors must stay on the correct objects. Soft studio light, 16:9.


OpenAI: GPT-6 Luna: Model A follows the requested 16:9 tabletop setup with the green cube left of the glossy red sphere, the blue cylinder behind them, and the yellow duck on top. Model B also gets the object colors and placement largely right, but its square framing misses the requested aspect ratio and its lighting is darker. (Second judge pass, order swapped — scores are the average of both: Model B clearly shows the requested objects and relationships, but its square framing misses the requested 16:9 format and | Anthropic: Claude Fable 5.1: Model A delivers the requested 16:9 framing, co…
Legible multi-line text
A minimalist event poster with three lines of crisp, correctly-spelled text stacked and centered: 'NIGHT MARKET' large on top, 'Fridays · 6–11pm' in the middle, 'Riverside Pier 4' at the bottom, on a deep navy background, clean sans-serif, subtle grain.


OpenAI: GPT-6 Luna: Model A better matches the deep navy, subtle-grain poster style and preserves the requested mixed-case supporting text, though both models break the title across two lines. Model B is crisp and legible but uses all caps for the supporting lines and a flatter charcoal background. (Second judge pass, order swapped — scores are the average of both: Model A preserves the requested wording and capitalization more closely, with crisp centered type and a subtle grainy navy background. Model B is also c | Anthropic: Claude Fable 5.1: Model A renders all three lines exactly as spec…
Spatial layout
A clean isometric illustration of a bedroom: a bed against the LEFT wall, a round rug centered on the floor, a desk under the WINDOW on the back wall, and a floor lamp in the FRONT-RIGHT corner. Flat-vector style, consistent perspective.


OpenAI: GPT-6 Luna: Model A places the bed, centered round rug, desk beneath the back-wall window, and floor lamp in the requested locations. Model B is polished and well composed, but its desk sits beside rather than beneath the window, and its oversized rug is partly covered by the bed. (Second judge pass, order swapped — scores are the average of both: Model A more faithfully places the desk beneath the back-wall window while also keeping the bed on the left and the lamp in the front-right corner. Model B is lar | Anthropic: Claude Fable 5.1: Model A correctly places the desk directly unde…
Named art style
A ukiyo-e woodblock print of a fishing boat riding a large cresting wave at dawn, faithful to the flat color planes, bold outlines, and stylized foam of the tradition, muted indigo and cream palette.


OpenAI: GPT-6 Luna: Model A more faithfully evokes ukiyo-e woodblock printing through its muted indigo-and-cream palette, flat color planes, and stylized foam, while still showing a dawn sky and fishing boat. Model B is bold and polished, but its black-heavy contrast and vivid pink and yellow boats depart from the requested palette and tradition. (Second judge pass, order swapped — scores are the average of both: Model A more convincingly evokes a traditional ukiyo-e print, with stylized foam, a dramatic cresting w | Anthropic: Claude Fable 5.1: Model A nails the muted indigo-and-cream palett…
Restricted palette
A flat-vector illustration of a mountain campsite at dusk using ONLY four colors — cream, burnt orange, teal, and charcoal — with no gradients and no stray off-palette colors, balanced composition, 16:9.


OpenAI: GPT-6 Luna: Model A better follows the requested 16:9 campsite scene and is closer to a flat-vector treatment, though it includes off-palette greens and gray tones. Model B is more elaborate, but its square format and numerous extra orange and yellow shades depart further from the strict four-color palette. (Second judge pass, order swapped — scores are the average of both: Model B uses the requested cream, burnt-orange, teal, and charcoal palette more consistently and presents a detailed, balanced dusk cam | Anthropic: Claude Fable 5.1: Model A delivers the required 16:9 frame and a…
Matchup powered by OpenRouter.