AI — Page 3
Models, agents, infra, applied AI.
- Nous Research adds Perplexity Fast Search to Hermes Agent, says it is free
The September 24th post promises access across Nous Portal tiers, while the Portal's pricing page describes hosted tools as a paid-tier benefit.
- Isoquant launches GLM-5.3-Flash inference cloud with prices below the model maker's list rates
Aniket's September 24th post quoted $0.07 per million input tokens, while Isoquant's website now shows different rates and performance claims based on its own benchmark.
- OpenAI's reported $500 Pro Max plan may run on Cerebras infrastructure
TestingCatalog reports that the proposed ChatGPT tier would add 'Fastest Work and Codex' to Pro, with higher usage limits also possible. OpenAI has not listed it on its pricing page.
- micro1 launches a PII model that keeps synthetic identities connected
Founder Ali Ansari says flow-transform 1.0 scored 96% F1 on PrivacyBench; micro1's new end-to-end benchmark is synthetic and limited to tabular data.
- Anthropic signs $11.6B Akamai CPU deal, with an equity warrant attached
Anthropic is locking in seven years of CPU capacity; Akamai says the deal could expand by another $9B and estimates about $5.5B in related capital spending.
- Google will test its TPUs in orbit on SpaceX's Transporter-18
Sundar Pichai says a Planet-built prototype will fly next week, beginning an orbital test of Google's plan to put AI computing in space.
- Instinct users report slowdowns as capacity warnings surface
One user described delayed email alerts. The Information reported that Instinct had sometimes warned users it was at full capacity and responses might be slower; that warning does not establish how many users experienced delays.
- Liquid AI adds a 280M draft model to speed up its 3B vision model
Ramin Hasani's team says the experimental add-on cuts decoding time by up to 3.13x on an M5 Max, while leaving image processing and prompt prefill unchanged.
- JetBrains expands Air into a control system for rival coding agents
JetBrains CEO Kirill Skrygan is pushing the IDE maker into team workflows, cost controls and governance without demanding loyalty to Junie.
- Meta adds live video chat to Muse, giving its AI agent a face and voice
Alexandr Wang demonstrated the feature on September 24th, a day after Meta outlined the real-time avatar technology behind it.
- Meta gives Muse a live avatar, with subsecond video responses
Alexandr Wang's team pairs Muse Realtime Voice with a video model Meta says runs at about 870 milliseconds and watermarks generated output.
- HeyBrain opens Brain to all, with two data connectors listed as live
The company pitches a shared workspace for people and AI agents; its product site lists Google Drive and Notion as the sources currently available.
- Wang says Meta's Muse Charm keychain will ship in December
Alexandr Wang said the device is shipping in December, in time for the holidays. Demo screens show gesture timings and an avatar interface that Meta has not confirmed as shipping specifications.
- Meta acquired WaveForms for its voice research; Conneau's team joined MSL
The deal was reported in August 2025, after WaveForms raised $40 million from a16z; Alexis Conneau says some of the team's work is set for Meta Connect.
- Meta says Muse will make money by taking a cut of transactions
At Connect, Mark Zuckerberg also described new Mac control for the agent, extending a product that already works across files and native apps.
- Meta says Muse will come to all its AI glasses, with a spoken wake word
Alexandr Wang said users can wake Muse by saying its name. Meta has not given a release date or explained how glasses will approve consequential actions.
- Inferact's TPU megakernel hits 709 tokens per second on Kimi K3
Inferact reports 709 tokens per second on 16 TPU v7 chips, versus 452 on 16 GB200 GPUs, in a low-concurrency speculative-decoding test. Co-founder Woosuk Kwon, a vLLM co-creator, helped build the open-source kernel.
- Jev matched 72.5% of Kepler labels after a prompt rewrite
Pavel Rabtsevich's linked test classified 8,054 historical signals, but its stronger result came after the first run informed a redesigned prompt.
- Rabbit asks OS3 to compare itself, a day after the agent launch
The company’s prompt arrived after OS3 opened beyond invite-only beta, pitching one assistant across models, computers and its r1 device.
- Robocurve says Opus 5.5 matches Astra on robots for 21% less
Across 360 real-robot trials, Opus 5.5 averaged 36% progress at an estimated 90 cents per run; Astra completed more tasks outright.
- HeyPCB launches Circuit World with 8,800 hardware designs to build on
Co-founder Inky Ganbold is pitching a browser workspace where engineers can fork open boards, edit them with AI and export files for fabrication.
- OpenAI brings ChatGPT Voice into Work, with access to connected apps
The September 23rd update pairs voice with GPT-6 models and workplace tools, extending ChatGPT Work to web and mobile.
- ACTx486 turns a podcast or video into a responsive AI conversation
Karina Nguyen and Jakub Zegzulka built the research demo around synthetic versions of Elon Musk and Joe Rogan, with generated dialogue labeled on screen.
- Anthropic says Claude made its apps 3.1x faster across 13 measurements
Anthropic says a two-week sprint improved 13 performance measurements across claude.ai and the Claude desktop app, with engineers setting targets and steering the work.