AI — Page 21
Models, agents, infra, applied AI.
- Bedrock-RL makes Minecraft repeatable enough to train vision-language agents
The open-source framework keeps worlds and rewards fixed while researchers swap models and trainers, aiming to make Minecraft VLM experiments reproducible.
- TestMu AI launches Agent Assurance to test AI agents by their effects
The former LambdaTest is extending its quality-assurance stack to agents that write files, call tools and access APIs.
- A four-hour Claude Code session built a Mac bridge for HP's Laser 1008a
Kuber Mehta's fix routes CUPS through a Linux VM and HP's SPL3 codec, then writes print jobs directly over USB.
- OpenAI grew 18% in Q2. Anthropic grew roughly 140%
OpenAI's sales reached $6.7B as losses deepened, while Claude Code helped Anthropic pass its older rival for the first time.
- Anthropic says 354 Claude protein designs bound in lab tests
Anthropic released the prompts and data behind 1,440 designs, though none were tested for biological function or structurally solved.
- Z.ai launches GLM-5.3 API for coding and defensive security
The text-only model keeps GLM-5.2 pricing, adds a 1M-token context window, and arrives before its delayed public weights.
- Fable 5 responses surface an internal 'Kettle' route, researcher says
A decoded thinking signature points to a backend change first seen August 16th, though Anthropic says the public model ID remains pinned.
- Firefox taps Exa to power cited AI search in Smart Window
The deal puts Exa retrieval inside Firefox's desktop AI workspace and its voice-based Quick Answers feature on iOS.
- Tomas Koutsky explains how Meta fits Muse Glimmer into 24 GB
Meta says its 30B local agent fits within 24 GB or 32 GB alongside a 131K context, vision encoder and speculative drafter. Byteline co-founder Tomas Koutsky explains the memory arithmetic.
- Anthropic opens Claude Cowork to every paid plan, beta label intact
The agent runs on web and mobile, while Anthropic's help pages retain rollout caveats and keep desktop as the full experience.
- Converge AI connects four products through one shared business context layer
Oliver Zhang's suite links knowledge work, software, creative work and games through shared context as Converge stages its US rollout; its growth product is next.
- FORT Robotics plans a SPAC merger at a reported $557M enterprise value
Samuel Reeves' FORT Robotics has not disclosed revenue or detailed merger economics, while Newbury Street II faces a November 4th, 2026 deal deadline.
- Bartowski releases an imatrix dataset for the ugly end of model quantization
The public corpus targets chat templates and tool use; without imatrix, Qwen3.6-35B-A3B fell about 28 BFCL points at Q2_K.
- Webml-community runs three Qwen3.5 models locally in browsers with WebGPU
Webml-community's demo runs Alibaba's 0.8B, 2B and 4B multimodal models locally through Transformers.js and WebGPU, with no cloud API required.
- Ashutosh Shrivastava's May Gemini Omni Flash demo exhausted a five-hour limit
Ashutosh Shrivastava's May 25th demo showed Gemini Omni Flash generating a personal avatar video, while a single prompt reportedly exhausted a five-hour usage allowance.
- Researchers evolved AI "mind viruses." The antivirus was one paragraph
The preprint showed natural-language payloads spreading through coding agents and files, while describing the current risk as limited.
- MiniMax's Vercel workout invite has no date or visible terms
MiniMax's open-weights workout joke with Vercel's v0 offers an invite, but the supplied posts disclose no date, venue, eligibility rules or formal partnership.
- Replit adds black-box pen tests and sends Agent to patch the holes
Level 3 scans test a sandboxed copy from inside and outside the codebase, then hand confirmed findings to Replit Agent for reviewable fixes.
- Devin migrated Cognition's site and recorded its own verification runs
Jared Palmer says the Astro-to-Next.js project used recorded tests to reproduce subtle failures, though Cognition published no scope or performance data.
- Alvys puts a build-your-own AI agent layer inside its freight TMS
Nick Darman's Foundry starts with 20-plus templates, human approval gates and a waitlist, extending the $40M Series B bet into freight operations.
- Google tests phone photos as a screen for insulin resistance
PhotoScan estimated body-fat distribution in a clinical study, though its closest DXA comparison rests on a 132-person validation cohort.
- J-Space says its text harness pushed DeepSeek V4 Pro past Fable 5
The reported gains are substantial, but the viral claim outruns single-run tests that mix vendor benchmark methods.
- OpenAI awards $1M to 14 projects studying AI's economic and social effects
The six-month grants span jobs, benefits, data-center energy, biosecurity and democratic oversight, with results due in 2027.
- Sam Hogan launches Lumbridge, a RuneScape-style test world for AI agents
The open-source experiment lets Claude, Codex, Hermes and other coding agents control persistent characters through a TypeScript SDK.