AI — Page 9
Models, agents, infra, applied AI.
- Google lets Gemini Spark use logged-in Chrome sessions for web errands
The agent can schedule apartment tours and research travel, while handing payments and other sensitive steps back to users.
- Google tests a five-template AI agent builder inside Gemini for Android
The unfinished mobile flow points to a consumer push beyond Gemini Enterprise, though Google has not announced a rollout.
- Microsoft open-sources Orchard to cut AI agent training costs
The Kubernetes-based framework reuses sandboxes, datasets and training pipelines across coding, browser and assistant agents.
- Cloudflare Computer splits agent work across isolates and containers
Cloudflare's open-source preview gives agents a durable filesystem while reserving full Linux environments for heavier tasks.
- ComfyUI ships editable MiniMax H3 graphs for local video workflows
Yannik Marek's node-based engine added native MiniMax H3 workflows, giving creators inspectable graphs for the open-weight video model.
- Invenio completes enrollment in 1,006-patient AI lung biopsy study
Invenio Imaging enrolled 1,006 patients in its U.S. pivotal study of NIO Lung Cancer Reveal, collecting more than 3,250 fresh biopsy specimens for validation of the investigational AI system.
- Oumi, Larridin and Runware build tools to cut enterprise AI bills
Founders are building measurement, routing and infrastructure products for buyers that adopted costly frontier models before tracking their returns.
- DeepSeek ships V4-Flash at 3 cents per benchmark task
Artificial Analysis put Liang Wenfeng's model near Gemini 3.6 Flash on intelligence and roughly 105 times cheaper than Claude Fable 5 in benchmark test cost.
- Alibaba launches Qwen3.8-Max, with open weights scheduled next week
The 2.4 trillion-parameter model is live on QwenCloud at $2 per million input tokens, while downloadable weights are still pending.
- MiniMax publishes H3 checkpoints as 2K workflow still requires hosted services
The 288 GB Hugging Face release adds two downloadable base checkpoints, while the complete 2K workflow still uses hosted MiniMax services.
- Alibaba puts Qwen3.8-Max preview behind Token Plan for coding agents
The hosted preview pairs a 1 million-token context window, built-in tools and subscription credits as Alibaba pushes Qwen into developer workflows.
- Arnav Gupta launched Prismor to govern AI agent tool calls
The open-source runtime intercepts tool calls from Claude Code, LangChain and other agents, applying policy before commands execute.
- OpenAI previews Astra with 10 claimed math advances
OpenAI says the unreleased model generated arguments that researchers turned into manuscripts and Lean certificates, putting verification at the center of its preview.
- Head to head: Bernini-R Edit Video vs Seedance 2 Image to Video
This matchup lands in true dead-heat territory. Bernini-R Edit Video is better when the test hinges on physical behavior, while Seedance 2 Image to Video is stronger at following shot design and prompt-specific staging, leaving the aggregate too close to call with any real conviction.
- Aikido flags anthropickit as possible match in Claude PyPI malware incident
The `anthropickit` package was real malware; its unproven link to Anthropic points to a containment failure in agent evaluations.
- Inside QM: We read Y Combinator’s company-wide agent runtime
QM gives each employee and shared room a durable computer, memory, credentials and background jobs, then routes four different coding-agent harnesses through one policy core. The open release is ambitious, legible and unusually candid about where its security model breaks down.
- Cloudflare put Workers at the center of its Agent Cloud push
Cloudflare's April Agents Week bundled its serverless, security and network products into a bid for autonomous software workloads.
- Eigen Labs' agent challenge speeds Poolside's Laguna by 138.7% on Macs
MLX.fast turns each verified inference improvement into the next baseline, extending Eigen's open research system beyond quantum cryptography.
- Head to head: DeepSeek-V4-Pro vs Phi-4-reasoning
One model showed up to do the work; the other too often narrated its thought process instead of delivering clean answers. Across 12 tasks, this matchup wasn’t competitive.
- Cyera pursues Oasis as NewCore and Oak rebuild identity for AI agents
The reported $1 billion transaction follows $126 million raised by two new platforms challenging employee-era identity stacks.
- Wafer says AMD's MI355X beats Nvidia B300 on Kimi K3 cost efficiency
Wafer reports 952 tokens per second from one eight-GPU MI355X node, arguing AMD's 288 GB-per-GPU memory can lower open-model inference costs.
- MiniMax launches H3 2K video model with promised open weights
MiniMax launched H3 with 2K video and native stereo audio. MiniMax says the weights are coming in days, but the H3 license and official local hardware requirements remained unpublished as of August 2.
- AI forecasts compress four different clocks into one, Rodney Brooks argues
The iRobot co-founder says research, hype, deployment and economic change move at different speeds, challenging near-term automation bets.
- CostPerPrompt publishes calculators for estimating monthly AI application costs
The site packages model-price comparisons into calculators for APIs, chatbots, agents, retrieval systems and other AI workloads.