AI — Page 22
Models, agents, infra, applied AI.
- Nous Research turns Hermes profiles into reusable specialist Bots
The desktop update gives each named Bot its own model, memory and skills, with cross-Bot communication inside the same app.
- Khosla and Emanuel argue autonomous AI could outperform doctors using AI
The JAMA Perspective draws on simulated studies and challenges medicine's prevailing physician-led deployment model.
- Amazon buys and destroys printed books for AI training data
A tracker hidden in a roughly 1,000-book order led 404 Media to a Las Vegas site where workers cut bindings and scan pages.
- 3M expert told ChatGPT to build zero-fault case before $61M explosion verdict
Discovery produced 350 pages of prompts, including a request to show 3M was "0% at fault," before jurors assigned 30% of the blame to 3M.
- MiniMax pitches H3 as a video model for consistent game-sprite motion
MiniMax is pitching its general-purpose multimodal H3 model for game-sprite animation, though the supplied materials include no independent comparative benchmark.
- CarSignal launches a $200-a-month AI operating system for auto repair shops
The YC-backed startup says more than 40 shops use its software for booking, diagnostics, estimates, parts and payments.
- World Liberty links its tokens to Chinese AI models under US scrutiny
WorldClaw takes USD1 payments and offers 43 models tied to Chinese companies flagged by the Trump administration, Reuters found.
- Uber backs Zipline to target 1 million daily drone deliveries by 2029
The Uber Eats integration will begin in Houston and Dallas before expanding across dozens of U.S. cities.
- Speko routes voice AI across providers using language-specific benchmarks
Speko has raised $1.1M and received a $500,000 YC commitment, for roughly $1.6M total.
- Bedrock Robotics deploys operatorless excavators at three commercial sites
Bedrock's former Waymo and robotics engineers must prove autonomy can handle changing jobsites and lower contractors' cost per yard.
- NODA AI claims $100M award, but federal challenge spans multiple winners
Philong Duong's startup says it received $100 million, but the federal program offers that amount across multiple autonomous-systems awards.
- Intersignal tests signed context handoff between different local AI models
David Seaman's self-funded lab let one Mac verify context from another, then rebuild it inside a different embedding space.
- Alibaba releases Qwen3.8-Max weights without clear license terms
Alibaba paired an Apache-licensed 27B model with downloadable flagship parameters on August 17th, but Qwen3.8-Max's license and official specifications remain unresolved.
- WildClawBench ranks Alibaba's Qwen3.8-27B 15th on agent tasks
The 27B open-weight model scored 48.0% across 60 WildClawBench agent tasks, trailing larger hosted systems while beating its Qwen3.6 predecessor.
- San Mateo County weighs on-site human supervision for commercial humanoid robots
The framework could require on-site human supervisors, kill switches, job-loss assessments and fees for fire response.
- MathCode converts plain-language problems into Lean 4 theorems and attempts formal proofs
Math-AI, the open research community associated with Princeton researcher Yifan Zhang, built a local terminal assistant that stores and reuses successfully proved theorems.
- Pika's April disclosure details one-GPU architecture for lower-cost real-time video
Founders Demi Guo and Chenlin Meng connect Pika's creator pricing to a one-GPU streaming architecture while keeping model costs private.
- Unslop offers a $19 editor for restoring voice to AI-polished drafts
Founder Spartak built the English-focused tool after AI grammar fixes made his own posts sound generic; paid rewrites run through Claude.
- Anthropic nears $7B Decart deal after founders favor it over Nvidia
Decart's founders and Sequoia reportedly chose Anthropic despite a higher Nvidia offer; most of the consideration would be Anthropic stock.
- OpenAI lets GPT-5.6 Sol delegate grunt work to cheaper Luna agents
The new routing can keep Sol in charge while assigning bounded work to faster, lower-cost Luna subagents.
- Duolingo cuts AI video-call cost below 1 cent with open models
The cost decline is helping Luis von Ahn move conversation practice into Super and reconsider Duolingo's premium Max tier.
- OpenAI benchmarks a 16x load-time gain for giant ChatGPT and Codex threads
A 741-turn, 231 MB thread loaded in 1.66 seconds in a test, with lower memory growth and 98% fewer requests.
- Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.78
In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.78. OpenAI’s GPT-5.6 Luna Pro followed at 0.77 with the lowest reported cost among the top three, at $0.0005 per task.
- Netflix tests GenRec on 10% of traffic, reports 0.006% relative lift
Netflix engineers Ying Li, Arjun Rao and Shradha Sehgal tested an LLM-backed ranker using about 40 times fewer Phase 2 labels than the production model, reporting a 0.006% relative lift in an undisclosed core metric.