Atomic.chat says Qwen 3.7-max beat Opus 4.7 and GPT-5.5 in agentic Tetris-bot test
In a 10-iteration code-and-rewrite challenge, the team let each model build and train a Tetris bot, then compared the final agents.
By Ryan Merket
· Published
Why it matters
Operator-style, agentic evals are closer to how teams will actually use frontier models. If Qwen is edging rivals in code-and-iterate loops, stack choices and costs could shift quickly.
Reader comments
Conversation for this story loads after sign-in.