AI — Page 19
Models, agents, infra, applied AI.
- Instant's team joins OpenAI as its cloud service heads for shutdown
The database developer tool will support existing cloud apps until August 31st, 2027, then preserve backups for another year.
- Munder Difflin says one agent redrew 119 blog heroes for $0
Founder Chaitanya Giri turned Reddit criticism into a Claude Code workflow built from SVG parts, 16 composition archetypes and one media manifest.
- FreeToken reports 22-25 tok/s for a 284B DeepSeek model on one RTX 5090
Shuo Yang, Xiaoze Fan and collaborators report the decode rate using CPU compute, a 32 GB RTX 5090 and 192 GiB of system memory.
- MCP roadmap prioritizes agent identity, server events and transport unification
Co-creator David Soria Parra is steering the protocol toward long-running, delegated work across enterprise systems.
- Munder Difflin fixes agent cost, memory and messaging failures in v0.4.5
Chaitanya Giri's local-first harness wraps 12 agent CLIs. Its v0.4.5 release notes document repairs to cost reporting, Apple Silicon semantic memory and communication between workers.
- Dan Luu reports a 7% ripgrep speedup after minutes of prompting
Dan Luu argues coding agents have made once-specialized optimization cheap enough to customize software for each workload.
- Outer Biosciences has raised about $23M to build a skin-assay licensing business
Outer Biosciences, Michael Polansky's 19-person company, uses machine learning and human skin tissue that remains viable for four weeks to find compounds for beauty and pharmaceutical companies.
- Cloudflare syncs AI bot rules with robots.txt, closing its own policy gap
Product manager Jin-Hee Lee's feature is headed for an upcoming rollout, publishing Search, Agent and Training choices while preserving existing Disallow rules.
- SpaceXAI puts Grok 4.6 in Google Cloud's Vertex AI Model Garden
Google's August 21st release notes list the model as a Preview offering, while SpaceXAI quotes $0.30 cached input and a 500,000-token context window.
- MiniMax hints at M3 on SambaNova, provides no benchmark result
An undated MiniMax post mentions M3 and SambaNova hardware, while the companies' disclosed demonstration identifies M2.7 and provides no M3 setup or result.
- OpenAI cuts GPT-5.6 Sol output pricing by one-third for three months
The flagship model now costs $4 per million input tokens and $20 per million output tokens, while paid-plan usage limits stay put.
- Grok declared a Chinese influence campaign real. Ten seconds after a ChatGPT audit, it backed down.
After showing a 120-source counter, Grok treated a disputed $23.6B data-center tally as evidence of a foreign-directed US campaign. ChatGPT found that the underlying sources did not establish the money, instructions or control required for that conclusion.
- Shanghai AI Lab published an August 13th paper on its 397B scientific-agent model
Intern-S2-Preview combines visual pretraining, tool use and long-horizon reinforcement learning in an open 397B-parameter scientific model.
- NVIDIA's AVO scores 100% on ARC-AGI-3's public set
Powered by Claude Opus 5, the coding-agent system cleared 183 levels in 6,624 actions, though NVIDIA did not test the private sets.
- Z.ai offers up to 5 trillion GLM-5.3 tokens to recruit ZCode users
The two-day promotion gives 50,000 new users 100M tokens each, extending a coding-agent push that Z.ai says already reached 1M users.
- MyDataWork puts a $7,500 analytics context plan on AWS Marketplace
Founder Gib Bassett is betting that easier cloud procurement can move his metadata workspace from analyst trials into shared team deployments.
- Argentic bets AI scrapers will pay 10 sats for a proxy hop
The pseudonymous builder said the zero-cost prototype had no paying agents after seven weeks, a candid test of whether bots will pay instead of route around.
- DeepSeek's experimental vision model spans three formats, caps images at 384 tokens
Liang Wenfeng's separately named `deepseek-v4-flash-vision-exp` supports Chat Completions, Messages and Responses requests, with each image billed at no more than 384 V4-Flash tokens.
- SGLang publishes one-GPU Qwen3.8-27B recipes, claims 206.1 tokens per second
Ying Sheng and Banghua Zhu's inference project added NVFP4 and DFlash2 recipes, with project-reported throughput of 206.1 tokens per second on one RTX 5090.
- Anthropic's $1.5B Ode builds an AI consultancy around private equity portfolios
Chris Taylor's 100-engineer operation gives Claude a deployment arm and its financial backers a ready-made customer pipeline.
- Pew finds AI fingerprints on 35% of newer webpages it could date
The 490,000-page study found a 10% rate overall, with commercial sites far ahead of government and education domains.
- Apple Music will label AI-made tracks, artwork and videos later this year
Providers must flag material AI use across recordings, compositions, artwork and videos, while Apple's public delivery spec still calls the metadata optional.
- Locus launches one balance for 600 APIs, because agents collect subscriptions too
Cole Dermott's YC-backed startup is selling platforms a metering layer spanning models, search, scraping and data providers.
- Liquid AI trains 4-bit LFM2.5 checkpoints to retain roughly 97% of BF16 performance
The MIT spinout used quantization-aware distillation across four LFM2.5 models, then tested them on a phone, laptop, mini PC and Raspberry Pi.