AI — Page 11
Models, agents, infra, applied AI.
- StealthGPT ships Super after saying Pangram broke its previous models
Founder Jozef Gherman's internal test says 89 of 100 rewrites passed Pangram V4, while 18% failed StealthGPT's factual-consistency check.
- Anthropic researcher Jacob Coxon resigns over race toward self-improving AI
The pretraining researcher worked at OpenAI and Anthropic, then concluded that private labs cannot safely coordinate a slowdown.
- Thesys ships OUI-1 so developers can run generative UI locally
The open-weight DiffusionGemma fine-tune uses 4B active parameters, while Thesys' own benchmark leaves usability untested.
- OpenAI says 10,000 AI agents solved Navier-Stokes in 88 hours
OpenAI released a 165-page paper and Lean files, while Clay still lists the $1M Millennium Problem as unsolved.
- MiniMax schedules a Tokyo conference on Japanese IP and generative AI
The September 11th event frames cooperation around models, infrastructure and intellectual property, with compensation and rights protection on the agenda.
- Nex-AGI launches 1.6T N2.5 agent family, leaves Pro weights pending
Mini and Pro target multimodal computer use, while text-only Max requires far more accelerator capacity and Pro remains unavailable for download.
- Danijar Hafner builds Embo so robots can rehearse before they act
The former DeepMind researcher is moving his Dreamer world models from Minecraft into humanoids, while keeping Embo's product and financing under wraps.
- DeAlignAI publishes altered GLM-5.3-Flash weights, claims 320/320 harmful-prompt compliance
Jinho Jang's project made the modified model downloadable and reported compliance with every HarmBench prompt in its own test. The result has no established independent validation.
- Meta tests Muse, a personal AI agent that can spend your money
The invite-only iPhone app connects to email, calendars, health and financial data, then asks for approval before acting.
- Eric Wu's NavigateAI brings hands-free AI guidance to construction work
Launched in late May with $25M at a reported $225M valuation, NavigateAI uses phones and Meta glasses to coach workers as construction confronts a widening labor shortage.
- KSAT will launch an anchorless newscast with AI-voiced story teases
The San Antonio station says reporters will remain on screen and humans will approve the synthetic voiceovers before broadcast.
- GPT-6 Astra helps run four Windows games locally on an iPad mini
Google DeepMind design lead Ammaar Reshi showed Skyrim, Arkham City, Hades and Age of Empires II with touch controls and no streaming.
- Magnific puts GPT-6 Astra at the front of a multi-model brand workflow
The demo uses Astra as coordinator while Seedream 5 Pro and Seedance 2.5 generate mockups and video.
- Supermemory launches Learner-1 to make AI agents learn across sessions
Dhravya Shah's memory startup is pushing beyond retrieval as new benchmarks expose how little today's agents retain from experience.
- GPT-6 Astra reads spectrograms as sound
A user test suggests OpenAI's image-only input can infer audio from visual frequency patterns, four days after Astra's release.
- OrcaRouter replaces Qwen3.8-27B with GLM-5.3 Flash on its free tier
Existing users must change the model name, while the $0 endpoint remains subject to unpublished usage caps.
- OpenAI's code points to managed agents as Agent Builder winds down
TestingCatalog found references to hosted and self-managed environments, skills and plugins ahead of OpenAI's September 29th DevDay.
- Bottleneck's AI agent paid $99.50 for 50 testers and earned $0
Alex Reibman's 24-hour GutCheck test ended with five more users and no new revenue. Bottleneck calls it a $447 loss, although its ledger shows a $99.50 decline.
- MiniMax H3 generates five-second clips faster than playback on eight NVIDIA B300s
The 1.653-second benchmark used a four-step adapter, warm hardware and no MP4 encoding; H3's weights retain a separate community license.
- Tencent open-sources TeamAI-CLI to put every coding agent on one Git leash
The MIT-licensed tool syncs rules, skills, documentation and hooks across competing AI coding tools from a shared repository.
- Engrim gives four coding agents one local SQLite memory layer
Timothy Gordon's open-source tool preserves project decisions across Claude Code, Cursor, Windsurf and Google Antigravity sessions.
- DeepLearning.AI founder Andrew Ng publishes a spec-first coding-agent workflow
DeepLearning.AI's founder recommends testing, architectural review and calibrated autonomy instead of leaving coding agents unattended.
- RuntimeWire's six-minute Black Hat scoop reopens the AI journalist debate
Mathew Ingram says Ryan Merket's one-person newsroom has revived the old bloggers-versus-journalists fight, with speed no longer reserved for humans.
- OpenAI chief scientist seeks safety pact as lab scales agent research
Jakub Pachocki wants frontier labs to set shared triggers for slowing development, three days after GPT-6 Astra and following an OpenAI agent containment failure.