Mona-lisa-1 surfaces on Arena as a possible GPT Image successor
The anonymous test has prompted speculation about an OpenAI successor, though the codename does not establish who built it.
By Ryan Merket · Published
Why it matters
Arena tests often provide the first public look at frontier models. If mona-lisa-1 is an OpenAI system, it suggests the current image-generation leader is already testing its successor.

A codenamed image generator called "mona-lisa-1" appeared in Arena testing on August 9th, prompting speculation that OpenAI is evaluating a new GPT Image model in public.
AI visual creator Riccardo Wolf (@WolfRiccardo) published raw outputs from the model in a two-post thread. Wolf said the images showed a noticeable improvement in realism, particularly in reducing the glossy, synthetic skin and surfaces that often make generated photographs look artificial.
The developer behind mona-lisa-1 remains unverified. Arena frequently places prerelease systems into anonymous comparisons under aliases, allowing model providers to collect human preference data before deciding which versions to ship. Arena says it works directly with commercial and open-source developers to test prerelease models, and that providers may submit multiple experimental variants.
That process makes the appearance of a codename meaningful without making its rumored ownership a fact. A model running on Arena can be a release candidate, a narrow experiment or a checkpoint that never becomes a public product. The mona-lisa-1 name itself provides no reliable attribution.
Testing realism through ordinary details
Wolf tested mona-lisa-1 with a prompt for a present-day New York City street. The instructions called for a parked car with a realistic brand logo, a specific license plate, authentic storefronts, signs, sidewalks and traffic details. Those requirements probe several persistent weaknesses in image generators at once: accurate text, coherent objects, recognizable commercial marks and the dense geometry of a real urban scene.
Wolf shared outputs at high and medium settings without edits. The images show the type of crowded, everyday composition that can expose failures hidden by more forgiving prompts for portraits, concept art or cinematic scenery. Storefront lettering, vehicle details and repeated architectural elements give viewers more opportunities to spot malformed text and inconsistent geometry.
A handful of images cannot establish whether mona-lisa-1 consistently outperforms released systems. Wolf's post contained no controlled comparison using the same seed and settings across competing models, and it did not include enough generations to measure failure rates. His tests provide evidence that the model is active on Arena and can produce convincing individual outputs. They do not establish its overall ranking, speed, editing stability or reliability across prompts.
Those distinctions matter because OpenAI already holds the top position on Arena's published text-to-image table. Arena's July 31st leaderboard ranked gpt-image-2 medium first with a score of 1,381 plus or minus five, based on 66,665 votes. Mona-lisa-1 does not yet appear on that published table under its codename.
Why the OpenAI theory is plausible, but unconfirmed
OpenAI released ChatGPT Images 2.0 on April 21st. OpenAI described improved instruction following, text rendering, world knowledge and realism as central gains over earlier image systems. The model is available across ChatGPT plans, while a separate thinking mode can research and plan an image before generating it.
Those capabilities overlap with the areas Wolf's test was designed to examine. That overlap supports the speculation around mona-lisa-1, although it does not prove lineage. Rival image labs are pursuing the same problems and can use arbitrary codenames during testing.
A stronger attribution route would come from provenance data. OpenAI says images created through ChatGPT, Codex and its API contain C2PA metadata and invisible SynthID watermarks. Its public verification tool can inspect an uploaded image for those signals. A positive result can associate an image with OpenAI's tools, though it would identify the provider rather than disclose a product name or confirm that mona-lisa-1 will ship.
The timing puts pressure on any prospective release. OpenAI's gpt-image-2 already leads Arena, while Reve, Meta, Google, ByteDance and Microsoft occupy much of the published top tier. The practical bar for a successor is therefore higher than producing a sharper demonstration image. It would need to preserve detail through repeated edits, follow crowded prompts reliably and improve generation speed or cost enough to justify replacing an incumbent that users already prefer.
For now, mona-lisa-1 is best understood as an active Arena experiment with promising public samples. Its provenance, production readiness and relationship to OpenAI's released image models remain unresolved by the codename alone.