Head to head: CogView vs Cosmos 3 Super
CogView vs Cosmos 3 Super
By Ryan Merket · Published
This matchup tests prompt fidelity across wildlife, typography, anatomy, constrained illustration, lighting, and geometric perspective. The scorecard looks decisive at first glance, but the sample read tells a more complicated story.
Cosmos 3 Super posts the higher aggregate score, 65.1 to CogView’s 50.1, and it is the more consistent prompt-follower in the individual comparisons. It better captured the pika’s extended shale-to-shale leap, respected the four-color flat-vector constraint, produced more plausible bracelet-tying hands, and built the cleaner one-point-perspective library aisle. That advantage also appeared in the denser narrative scenes. Cosmos more convincingly assembled the mangrove setting, fishing cat, candlelit mood, salt-flat boardwalk, stilts, researcher, bioluminescent algae, kelp, terns, and horseshoe crabs. Neither model was exact: Cosmos missed requested object counts, duplicated figures, added unwanted moons or animals, and sometimes left compositions sparse. CogView’s recurring weakness was substitution: rabbit or squirrel anatomy for a pika, lighthouse-like buildings for observation towers, ordinary crabs for horseshoe crabs, and gradients where a restricted palette demanded flat color. Its clearest counterpunch came in one typography variant, where it preserved the requested three-line hierarchy, wording, navy field, and grain while Cosmos broke the title and introduced white margins. The conflicting text results show how brittle both systems remain across generations. The crucial result is statistical, not cosmetic: despite the 15-point aggregate gap, the evaluation assigns only limited confidence that either model is genuinely better. That makes the apparent Cosmos lead descriptive rather than conclusive. **Final call: too close to call — CogView and Cosmos 3 Super are effectively tied in this test.**
Saltflat Observation Towers
A hyper-detailed cinematic matte painting of three weathered birdwatching towers on the pale lavender salt flats of Laguna Brumosa, viewed from a low angle along a narrow boardwalk whose planks converge to a single distant vanishing point; a giant rust-red wind pump stands far in the background, two black-necked stilts pick through shallow reflective water in the foreground, and a tiny field researcher in a teal rain cape leans on the nearest railing, with all relative sizes perfectly believable and the horizon straight, no warped geometry, crisp atmospheric depth, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B more convincingly depicts weathered observation structures, two black-necked stilts, a leaning teal-caped researcher, reflective flats, and believable boardwalk perspective, though it includes only two towers rather than three. Model A includes three towers and strong convergence, but the buildings resemble lighthouses, the researcher is oversized and not leaning, and the distant wind pump is not giant. (Second judge pass, order swapped — scores are the average of both: Model B more conv | Anthropic: Claude Fable 5.1: Model B delivers a coherent scene with thr…
Pika Leap Through Shale
An ultra-sharp wildlife editorial illustration of a silver-eared pika launching mid-air between two jagged slabs of blue-gray shale on Mount Verdan Coil, tiny pebbles scattering behind it, whiskers swept back, hind legs fully extended, a spray of alpine snow dust and blurred moss fragments emphasizing speed, the camera frozen at the peak of the jump with dynamic diagonal composition and palpable momentum, naturalistic anatomy, dramatic motion energy, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B closely matches the silver-eared pika, fully extended leap, jagged blue-gray shale, scattering pebbles, and dynamic diagonal momentum requested. Model A is polished but resembles a rabbit, uses a static frontal composition, and lacks the extended anatomy and shale-to-shale trajectory central to the prompt. (Second judge pass, order swapped — scores are the average of both: Model B closely captures the silver-eared pika in a dynamic lateral leap between layered blue-gray shale, with scatt | Anthropic: Claude Fable 5.1: Model B depicts a convincing pika (compact…
Restricted palette
A flat-vector illustration of a mountain campsite at dusk using ONLY four colors — cream, burnt orange, teal, and charcoal — with no gradients and no stray off-palette colors, balanced composition, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B follows the four-color, flat-vector restriction far more closely, with clean color separation and no obvious gradients, though its large empty margins weaken the composition. Model A is more detailed and visually dynamic, but extensive gradients, shading, and stray off-palette hues directly violate the core prompt. (Second judge pass, order swapped — scores are the average of both: Model B closely follows the strict four-color, flat-vector requirement with a balanced 16:9 campsite scene, | Anthropic: Claude Fable 5.1: Model B strictly uses the four specified f…
Candlelit Mangrove Ritual
A luminous storybook-style painting of a fishing cat standing knee-deep in a dark mangrove pool at twilight, illuminated only by seven floating beeswax candles trapped among arching roots, warm amber reflections trembling across the water while the surrounding foliage falls into cool indigo shadow, fireflies faint in the distance, the mood hushed and reverent, with precise candlelit illumination shaping the cat’s face and wet fur, intimate composition, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B better captures a true mangrove environment, a fishing-cat-like subject, knee-deep water, and the hushed candlelit reverence, though it depicts only four candles. Model A appears to include seven candles and has strong reflections, but the animal resembles a domestic tabby and the bright cyan illumination weakens the prompt’s candle-only lighting. (Second judge pass, order swapped — scores are the average of both: Model B more convincingly depicts a stocky fishing cat, dark mangrove sett | Anthropic: Claude Fable 5.1: Model B delivers a convincing fishing cat…
Bioluminescent Tide Gathering
A richly composed semi-realistic digital painting of a moonless coastal cove where five royal terns, three horseshoe crabs, a juvenile sea lion, and a lone marine botanist in orange waders interact around a ribbon of glowing blue bioluminescent algae on wet black sand; one tern hovers with a small fish, two argue over drift kelp, the sea lion noses a shell beside the botanist’s sample crate, and the horseshoe crabs form a subtle foreground arc, all arranged in one balanced coherent scene under dim neon-blue surf glow and soft overcast night light, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B better captures the glowing algae ribbon, wet black sand, kelp interaction, fish-carrying tern, and foreground arc of recognizable horseshoe crabs in a coherent semi-realistic scene. Model A is polished but substitutes ordinary crabs, omits key interactions, and both models incorrectly add a moon and duplicate the lone botanist, while Model B also has too few terns and an extra sea lion. (Second judge pass, order swapped — scores are the average of both: Model B better captures the richl | Anthropic: Claude Fable 5.1: Model B captures far more of the specified…
Legible multi-line text
A minimalist event poster with three lines of crisp, correctly-spelled text stacked and centered: 'NIGHT MARKET' large on top, 'Fridays · 6–11pm' in the middle, 'Riverside Pier 4' at the bottom, on a deep navy background, clean sans-serif, subtle grain.


OpenAI: GPT-5.6 Sol Pro: Model B precisely delivers the requested three centered lines with correct spelling, crisp sans-serif typography, strong hierarchy, and a subtle textured navy background. Model A splits the title across two lines, misspells “Riverside,” adds an unrequested divider, and has distracting white side margins. (Second judge pass, order swapped — scores are the average of both: Model B closely matches the requested three-line hierarchy with crisp, centered text and correct wording, though the middle se | Anthropic: Claude Fable 5.1: Model A misspells 'Riversride', renders '1…
Hands & anatomy
A natural candid photo of a person tying a friendship bracelet, both hands clearly visible with anatomically correct fingers and a plausible grip on the threads, soft daylight, shallow depth of field, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B more clearly depicts multicolored friendship-bracelet threads being tied with a plausible two-handed grip, while Model A resembles handling a thick wrapped cord rather than making a bracelet. Both are well composed, but Model B adheres more closely to the requested activity despite slightly smooth, synthetic-looking skin. (Second judge pass, order swapped — scores are the average of both: Model B clearly depicts both largely plausible hands tying multiple bracelet threads, with attractiv | Anthropic: Claude Fable 5.1: Model A delivers more photorealistic, anat…
Perspective & scale
A photorealistic one-point-perspective shot down a long empty library aisle, shelves receding to a single vanishing point, books getting consistently smaller with distance and nothing warped or wrongly sized, warm overhead light, 16:9.


OpenAI: GPT-5.6 Sol Pro: Model B delivers a longer, more precise one-point perspective with a clearly unified vanishing point and highly consistent diminution of shelves, books, and lights. Model A is attractive and convincing, but its blown-out endpoint and slightly softer geometry make the depth and scale less exact. (Second judge pass, order swapped — scores are the average of both: Model B more precisely delivers a long, symmetrical one-point-perspective aisle with highly consistent scaling, a clear vanishing point, | Anthropic: Claude Fable 5.1: Model B delivers a longer, crisper one-poi…
Matchup powered by OpenRouter.