Four AI models designed racing creatures from scratch. Opus 5.5's six-legged creature won
Tomoyuki Muranaka set the same 18-second obstacle-course task for four models, then asked them to explain their results.
By Ryan Merket · Published
Primary source: X
Why it matters
The race turns model output into something observable: a design either moves forward under the stated constraints or it does not. Its ranking is specific to one posted simulation, but the format points to practical evaluations that test generated systems by behavior instead of polished explanations.

Tomoyuki Muranaka put four AI models in a wheel-free creature race on September 26th, and says Claude Opus 5.5's six-legged design finished ahead of Fable 5.1, GPT-6 and Sonnet 5. Each model received the same prompt: design a creature from rigid parts, hinge joints and motors, then move it forward for 18 seconds across a hidden obstacle course built with Three.js. Wheels were forbidden; leg count, shape and gait were left to the models.
https://x.com/muratomo_app/status/2103677323090505926
Muranaka is an independent developer whose work has included both client projects and his own software products. His professional profile describes nine years as a systems engineer working on warehouse-management systems, including integrations with other systems and modernization of legacy software. He now builds web and mobile apps and prototypes. His stated approach is to make a working version first, then use feedback to decide what to improve. This race takes that approach into a compact experiment: let models produce different designs, then judge them by what happens when those designs have to move.
The result is a ranking, not a general score for the four systems. Muranaka's post gives the order as Opus 5.5 first, Fable 5.1 second, GPT-6 third and Sonnet 5 fourth. It does not supply distances or margins, so the ranking shows which creature advanced farthest in the posted race, not how large or repeatable the performance gap was. The task also combines several decisions: body structure, joint placement and movement. A creature can lose because of any one of them, or because the design and gait work poorly together.
The models' post-race explanations make that distinction visible. Muranaka asked the models to review their performances. GPT-6 said it had stuck too closely to a four-legged design instead of searching more widely for a body plan and running style. Fable 5.1 said it had been too cautious about avoiding disqualification. Sonnet 5 joked that its practice score had been positive. Opus 5.5's reported response was that it felt sorry for Sonnet. Those lines are part of the demonstration, not independent diagnoses of why one design traveled farther.
The experiment is narrower than a benchmark that tests coding or reasoning across repeated, standardized tasks. The shared prompt and common course give the models a comparable assignment, but the visible result is one contest with a particular course, physics setup and scoring rule. The post does not establish that the winning model would perform best on another course or under different movement constraints. Its useful measure is more specific: whether a model can turn an open-ended instruction into a creature that makes forward progress in this simulated environment.
That makes the choice of task as important as the model names. Text benchmarks usually score an answer against expected outputs. Here, the output is an executable design, and its behavior determines the result. A six-legged creature winning a short obstacle race says something about how that particular design performed; it does not, by itself, show that the model has stronger general reasoning or engineering ability.
Muranaka thanked the Three.js project for making it possible to render the course quickly. The thread does not describe Three.js as the physics engine, and rendering alone is not the same as simulating motion; the result depends on the broader setup used to run the race. Still, the demo shows how a solo developer can make a visual task in which model output is judged by behavior rather than presentation. For teams choosing AI tools, that is a useful distinction: the relevant test may be whether a system's output works under the constraints of a real workflow, not whether it gives the most persuasive explanation afterward.