Head to head: Ideogram Transparent vs Ideogram V4.0 Text to Image
Ideogram Transparent vs Ideogram V4.0 Text to Image
By Ryan Merket · Published
A razor-thin contest tests prompt obedience, typography, anatomy, counting, and photographic realism. The results expose complementary strengths—and several shared blind spots.
The aggregate scores—55.5 for Ideogram Transparent and 55.7 for Ideogram V4.0 Text to Image—are functionally identical. With only limited confidence that either model is genuinely better, the 0.2-point gap is noise, not a meaningful lead. Transparent showed its best work in typography, cinematic noodle-lane imagery, and macro food photography. Its cleaner text treatment and subtler grain often beat V4.0’s heavier texture, while several gooseberry-tart and street-scene outputs had stronger focal hierarchy, steam, reflections, and atmosphere. It also claimed multiple tea-pouring and exact-counting rounds—but neither category produced a consistent sweep. V4.0 was stronger when an image depended on readable physical action. It more often delivered convincing bracelet grips, active bakery preparation, orderly cup arrangements, and coherent kettle-to-cup movement. Yet those advantages were intermittent: cup counts and tea arcs swung between models across runs, showing that both remain sensitive to generation variance. The shared failures matter more than the fractional score difference. Both repeatedly ignored 16:9 framing, violated explicit negations by adding plants, lamps, and wall art, and put vehicles into supposedly car-free lanes. **Final call: TOO CLOSE TO CALL. Ideogram Transparent and Ideogram V4.0 Text to Image are effectively tied; choose Transparent for cleaner graphic styling and atmospheric polish, or V4.0 for somewhat stronger action staging—not because either is better overall.**
Tea arc leap
A high-speed editorial food photograph of a tattooed tea master in a saffron apron snapping a copper kettle to pour a ribboning arc of scarlet hibiscus tea into six tiny celadon cups, droplets suspended midair, steam tearing sideways, linen sleeve whipping with the motion, crowded night-market stall hinted as colorful streaks behind him, dramatic frozen action with convincing kinetic energy, crisp subject against motion-blurred background, cinematic tungsten and neon spill, 16:9


OpenAI: GPT-5.6 Sol Pro: Model B delivers a clearer kettle-to-cup pouring action, stronger kinetic arc, convincing market streaks, and a more cohesive editorial composition, though it shows seven cups rather than six and lacks visible steam. Model A has vivid color and good sharpness but includes too many cups, an awkward kettle pose, and a liquid arc that appears disconnected from the spout; both also miss the requested 16:9 format. (Second judge pass, order swapped — scores are the average of both: Model B delivers th | Anthropic: Claude Fable 5.1: Model A's scarlet arc materializes from th…
Negation
A cozy reading nook with an armchair, a stack of books, and a mug of tea by a window — with absolutely NO plants, NO lamps, and NO artwork or picture frames anywhere in the frame. Warm afternoon light, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model A and Model B both severely violate the negation requirements by prominently including multiple plants, a lamp, and a wall frame, while also using a square rather than 16:9 format. Model A wins narrowly for its stronger warm-afternoon lighting, more balanced composition, and slightly higher visual polish. (Second judge pass, order swapped — scores are the average of both: Model B and Model A both severely violate the negation requirements by prominently including multiple plants, a lamp, a | Anthropic: Claude Fable 5.1: Both Model A and Model B completely fail t…
Legible multi-line text
A minimalist event poster with three lines of crisp, correctly-spelled text stacked and centered: 'NIGHT MARKET' large on top, 'Fridays · 6–11pm' in the middle, 'Riverside Pier 4' at the bottom, on a deep navy background, clean sans-serif, subtle grain.


OpenAI: GPT-5.6 Sol Pro: Model A is cleaner, better balanced, and closer to the requested subtle grain, with a reasonable dot separator; Model B substitutes an asterisk and uses a heavier, more distracting texture. Both models incorrectly split “NIGHT MARKET” across two lines and alter the requested capitalization and time styling. (Second judge pass, order swapped — scores are the average of both: Model A has cleaner spacing, subtler grain, and more refined hierarchy, while its middle dot more closely matches the reque | Anthropic: Claude Fable 5.1: Model A renders all three lines correctly…
Bakery prep realism
A photorealistic documentary-style scene inside a small dawn bakery: a baker with flour-dusted hands scoring three oval rye loaves on a scarred walnut worktable beside a brushed steel bench scraper, a folded indigo linen, and a cracked ceramic bowl of wet dough, pale winter sunlight entering through a fogged window casting soft directional shadows, believable skin texture, flour on fabric, sheen on metal, pores in the wood, and translucent dough edges, composed naturally in 16:9


OpenAI: GPT-5.6 Sol Pro: Model A more clearly depicts the active scoring process and includes the cracked dough bowl, indigo linen, textured walnut, and convincing material detail, though it shows only two loaves and is not 16:9. Model B correctly includes three loaves and has strong lighting and composition, but the baker appears to press already-scored loaves rather than score them, and the bowl is not visibly cracked. (Second judge pass, order swapped — scores are the average of both: Model A more convincingly depict | Anthropic: Claude Fable 5.1: Model A captures the scoring action, a cra…
Hands & anatomy
A natural candid photo of a person tying a friendship bracelet, both hands clearly visible with anatomically correct fingers and a plausible grip on the threads, soft daylight, shallow depth of field, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B presents a clearer, more colorful bracelet-making action with a convincing grip, strong focus, and more appealing depth of field, while Model A is slightly harsher in lighting and less visually refined. Both show plausible hands but miss the requested 16:9 aspect ratio. (Second judge pass, order swapped — scores are the average of both: Model B shows a clearer bracelet-making action, a more convincing grip, and stronger color and depth, while Model A is anatomically plausible but flatter | Anthropic: Claude Fable 5.1: Model B presents cleaner, more anatomicall…
Car-free noodle lane
A rainy evening street-food scene in the old harbor district of Port Cendre, photorealistic cinematic style, a noodle vendor under a striped awning ladling broth for two customers at a standing counter while puddles mirror red signage and steam drifts through the alley, wet cobblestones, bicycles leaning against plaster walls, glowing shop windows and umbrellas, but absolutely no cars, no vans, no buses, and no motorcycles anywhere in the scene, 16:9


OpenAI: GPT-5.6 Sol Pro: Model A is more cinematic and polished, with stronger steam, reflections, lighting, and alley depth, though Model A includes a van and car and does not clearly show two customers. Model B depicts the vendor serving two customers more accurately, but Model B also contains two prohibited vans, has weaker realism, and both models miss the requested 16:9 format. (Second judge pass, order swapped — scores are the average of both: Model A has stronger cinematic composition, realism, lighting, and stea | Anthropic: Claude Fable 5.1: Both fail the core 'car-free' constraint—M…
Exact counting
A clean overhead flat lay on white marble of EXACTLY seven identical espresso cups arranged in a single straight row, each clearly separate and individually countable, even soft lighting, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B clearly presents exactly seven matching, separate espresso cups in a cleaner and more consistent straight row, while Model A includes visibly different cup colors and designs. Both use a square format rather than the requested 16:9 aspect ratio. (Second judge pass, order swapped — scores are the average of both: Model B clearly presents exactly seven separate, matching cups in a precise straight row with softer, cleaner lighting. Model A also counts correctly, but several cups differ in | Anthropic: Claude Fable 5.1: Both Model A and Model B show exactly seven…
Dew on gooseberry tart
A macro close-up studio product photograph of a single gooseberry tartlet on a matte slate crumb board, shot with a true macro lens so sugar crystals, dew beads, and hairline cracks in the glazed fruit are tack-sharp while the fluted crust falls into creamy blur, low three-quarter angle, shallow depth of field, cool morning window light from the left with a silver bounce fill, luxurious fine-art food styling


OpenAI: GPT-5.6 Sol Pro: Model B better captures the requested tack-sharp dew, delicate fruit surface detail, cool side lighting, and shallow macro depth while maintaining an appetizing, natural presentation. Model A is polished and well composed, but its uniformly arranged droplets and very glossy fruit appear slightly synthetic. (Second judge pass, order swapped — scores are the average of both: Model B more convincingly captures cool left-side morning light, abundant fine dew, slate texture, shallow macro depth, and | Anthropic: Claude Fable 5.1: Model A delivers a moodier, more luxurious…
Matchup powered by OpenRouter.