AI — Page 23
Models, agents, infra, applied AI.
- Stephen Cresswell ships Yadda 3 after a one-day Claude-assisted rebuild
The longtime JavaScript maintainer says Claude handled most of the work, while an existing test suite kept the agent from redefining correctness.
- Sankalp uses Codex loop to place 12th in GPU Mode's B200 QR contest
Sankalp placed 12th among 183 entrants after using Codex, profiling tools and more than 1,500 submissions to optimize a B200 QR kernel.
- Andon Labs carries out first firing recommended by its AI manager
Luna needed a human prompt to recover its attendance policy, and people still reviewed and executed the termination.
- Xia Chen ships ThoughtDAG v0.3.13 to expose LLM context
The open-source desktop app lets users branch, prune, merge and preview the exact conversation history sent to a model.
- California DMV permits Aurora and Kodiak to test autonomous heavy trucks
Kodiak has begun supervised runs near Mountain View, while Aurora and Kodiak must each complete 500,000 autonomous miles before driverless testing and another 500,000 before deployment.
- Community attributes animated SVG to Qwen3.8-27B without reproducible run details
A pelican-on-a-bicycle animation attributed to Alibaba's open-weight model lacks the prompt, raw SVG and execution details needed for independent reproduction.
- GitHub adds xAI's Grok 4.6 to Copilot for long-running coding agents
xAI's Grok 4.6 is entering GitHub's IDE, command-line and enterprise workflows two days after launch, under usage-based billing.
- Anthropic stole ChatGPT's aura. Then it gave it back.
Claude Code made Anthropic Silicon Valley's chosen winner. Its models and business remain formidable. Self-inflicted errors and a contested warning about collapsing morale have punctured the sense that the company could do no wrong.
- Ukraine finds Nvidia Jetson module in Russia's S-71M cruise missile
Ukraine says the robotics computer may support onboard vision, raising questions about resale routes that move commercial silicon into Russian weapons.
- MiniMax will present H3 at ModCon in San Francisco on August 18th
Morgan Suo will discuss MiniMax's multimodal video model at Modular's San Francisco conference for AI infrastructure developers.
- Grok 4.6 tops BBEH Mini benchmark, scoring 0.67 on 460 tasks
In a 460-task BBEH Mini evaluation, SpaceXAI’s Grok 4.6 led the field with a 0.67 score at $0.0034 per task. Anthropic’s Claude Opus 4.8 followed at 0.62, while OpenAI’s GPT-5.6 Sol Pro posted 0.59.
- Meta releases local Muse Glimmer as Zuckerberg argues for AI 'for everyone'
Meta released the 30B open-weight model alongside Zuckerberg's case for widely distributed personal AI, while keeping Muse Spark 1.2 behind its products and API for now.
- Andor Health adds remote monitoring to Epic Toolbox for CMS ACCESS
Raj Toleti is packaging devices, AI triage and clinical labor inside Epic as Medicare tests outcome-linked payments through its new ACCESS model.
- Alibaba schedules Qwen3.8-27B open-weight release for August 14th
Alibaba's Qwen team set August 14th for a 27B dense vision-language model, but the supplied materials do not confirm when downloadable weights will appear.
- Pony.ai and Uber plan more than 2,000 robotaxis across four European cities
Pony.ai will supply the autonomy stack, Uber will integrate rides into its network, and local partners may own and operate fleets in four cities that remain unnamed.
- Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79
In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.79 at an estimated $0.0056 per task. OpenAI’s GPT-5.6 Sol followed at 0.77, while several GPT-5.6 variants and Anthropic’s Claude Opus 4.8 clustered close behind.
- Apple trains China-specific AI model with Alibaba as Beijing clears rollout
The model would sit alongside Qwen and Baidu technology in a China-only stack expected to reach Apple devices in the coming months.
- NOPE founder explains how AI text watermarks survive copying and fade under rewrites
James Padolsey's guide explains how statistical text marks survive copying, while Declaude tests full rewrites against open watermark schemes.
- Alibaba's Qwen Cloud recap lays out an agent-first developer platform
Alibaba Cloud's June 30th Qwen Live recap describes an agent-first platform as Eddie Wu pushes paid model access, agent tooling and cloud adoption.
- SpaceXAI's Grokathon crowns a binary decompiler built in 12 hours
Nova beat social simulation tool Signal and brain-sensing speech prototype ThinkVoice in the 2026 competition.
- Assembly launches an AI app builder for client-facing software
The New York software maker is targeting service firms that need authentication, CRM data, permissions and payments built into generated apps.
- Google replaces Gemini 3.6 Flash after three weeks, cuts prices through year-end
Tulsee Doshi's team improved coding and workflow scores, though Google's full evaluation table shows gains were uneven.
- OpenAI launches ChatGPT Computer History to remember work across Mac apps
The opt-in feature turns screen activity into reusable context for Pro, Business, and Enterprise users, with a timeline and privacy controls.
- OpenAI previews GPT-5.6 Sol at up to 14x standard speed
Cerebras hardware drives the API tier to 750 output tokens per second, with access restricted to selected customers while capacity grows.