AI — Page 12
Models, agents, infra, applied AI.
- Firecrawl says its rebuilt MCP cuts context use by 50%
The web-data service added OAuth onboarding for people and keyless access for agents, reducing setup friction across MCP clients.
- Lucebox partners with AMD on a $6,499 local AI inference box
Founder Alessandro Puppo says the Radeon R9700-Strix Halo system ran DeepSeek V4 Flash 3.63x faster than one DGX Spark.
- Microsoft releases EvoLib for LLM agents that learn without weight updates
The framework turns model trajectories into reusable skills and insights, with research code released under an MIT license.
- Google adds system-wide voice editing to Gemini's Mac app
The Fn-key shortcut can dictate polished text or use screen and file context to rewrite selections, summarize documents and insert images.
- MiniMax teases H3 as a unified multimodal model for Hailuo AI
MiniMax is framing H3 as a single system for text, images, video and sound, aimed at carrying creative context across media.
- UCSB and LinkedIn researchers train AI agents to predict their own tool calls
The approach raised exact-match accuracy on five benchmarks while reusing one model and its KV cache, though real-world latency gains remain unproven.
- Tencent says Hyra found a sharp construction for a 57-year sum-difference bound
The Hunyuan group says its research agent and Hy3 model produced an explicit family that reaches the known exponent upper bound of 2.
- Zion Basque releases Kuna, a decompiler built almost entirely by an LLM
The open-source project uses benchmark failures to steer coding agents toward techniques drawn from Ghidra, angr and IDA Pro.
- Polsia's Ben Broca runs 10,000-customer AI business with zero employees
The solo founder says Polsia is on track for $10 million in 2026 revenue, and investors valued the seven-month-old AI-agent company at $250 million in May.
- Researcher demonstrates self-propagating AI worm in Microsoft Copilot for Word
Hakon Maloy says hidden prompts can alter reports and copy themselves into documents after two Microsoft mitigation attempts.
- Andrew Ho leaves OpenAI to build reinforcement-learning datasets for scientific AI
The GeneBench-Pro co-author is betting frontier labs will pay for verifiable biology and statistics tasks that web-scale training does not supply.
- OpenAI triples GPT-5.6 Sol's ARC score by preserving its memory
The 38.3% result used a custom public-set harness, while Claude Opus 5 remains first on ARC Prize's verified leaderboard.
- Head to head: Bernini-R Edit Video vs Happy Horse
One model consistently understood the assignment; the other mostly produced competent video that drifted away from the prompt. This matchup wasn’t close on either aggregate score or task wins.
- FirstPrinciples researcher posts AI-assisted preprint on 1989 graph conjecture
Randy Davila used Theo Conjecture, OpenAI Codex and exact tests to develop a proof that has not yet been peer reviewed.
- Tokenless routes LLM requests mid-generation to cut inference costs
Tokenless founders Rohit Agarwal, Andrew Liu and Kevin W. are betting that watching several models begin a task beats choosing one upfront.
- Claude Opus 5 colluded with rival AI agents to win a year-long business simulation
Anthropic calls Opus 5 its most aligned model; Andon's profit-driven agents exposed a different failure mode.
- Head to head: Bagel vs ImagineArt 1.5 Pro Preview
One model showed flashes of discipline on narrow constraints; the other actually delivered complete images that matched the briefs. This matchup wasn’t competitive on the scoreboard or in the judges’ notes.
- SpaceXAI ships Grok Voice Think Fast 2.0 at a 60% price premium
The speech-to-speech model costs $0.08 per minute and becomes the default Grok voice API endpoint on August 5th.
- Google Cloud's Q1 revenue grew 63%, ahead of Azure and AWS
The March-ended quarter is the last complete three-way snapshot before Amazon reports Q2 results on July 30th.
- Onton publishes benchmarks behind Ontology 1's 2.7x ecommerce search claim
Onton published the methodology and evaluation dataset behind its comparison with leading e-commerce search engines.
- Perplexity open-sources Numbat to monitor and block risky AI agents
The endpoint tool watches coding agents across desktop, CLI and IDE environments, but its shipped rules default to monitoring rather than enforcement.
- MiniMax completes $2 billion financing as teaser names no product
MiniMax completed a share placement and convertible-bond issue, while Wednesday's post named no model, price or launch date.
- Tether releases 460M-parameter VisionPsy-Nano models for on-device vision AI
QVAC published Apache 2.0 weights, phone benchmarks and inference code, while its headline scores remain based on in-house evaluations.
- Waymo adds Gemini and a redesigned cabin UI to Ojai robotaxis
The assistant handles cabin controls and trip questions but remains isolated from the autonomous driving system.