AI — Page 6
Models, agents, infra, applied AI.
- StepFun launches a 600B agent model at $1 per million input tokens
Jiang Daxin's flagship pairs a 1M-token window with aggressive pricing, while its own tests still trail top rivals on several coding benchmarks.
- Alibaba previews Qwen-Image-2.1 for character-consistent storyboards
The model accepts multiple reference images and turns a three-view character sheet into six scenes, while its open release remains listed as upcoming.
- Open-weight models reach 78.4% of Vercel AI Gateway token volume
Guillermo Rauch's daily snapshot shows open models dominating volume as Vercel moves closer to the center of the model market.
- Artificial Analysis starts counting when coding agents refuse the job
Micah Hill-Smith and George Cameron's index now separates blocked attempts from agents that recover through a fallback.
- NASA and IBM released an open lunar AI model trained on two million data bundles
Rachel Slank aligned 11 data types and labeled 49,000 craters; the GitHub release includes downstream code but omits the pretraining code.
- Alibaba open-sources RADAR to detect 146 findings in abdominal CT scans
A Science study reports 0.913 mean AUC and faster reads with AI assistance, while [clinical clearance and patient-outcome evidence remain unestablished](https://ichgcp.net/clinical-trials-registry/NCT07040358).
- OpenAI seeks a $1.2T-plus valuation while projecting $278B cash burn
Its five-year plan pairs $840B in revenue with $856B in compute spending, leaving investors to finance the gap.
- Cua open-sources a 706,048-parameter model for filling forms
CUA-S1-FORMS makes one-pass choices among predefined actions, then hands the plan to Cua Driver for execution.
- SpaceXAI launches a $0.20-an-hour transcriber that tops streaming accuracy
Grok Voice Transcribe 2.0 reached 2.7% final word error at 0.49 seconds, while its batch result placed fifth in Artificial Analysis testing.
- Meta opens Muse connectors so developers can plug services into its agent
The September 18th release turns Muse into a distribution layer where users can invoke outside services through natural-language requests.
- Meta tests a Mail tab that could give Muse its own inbox
The unreleased interface centralizes agent email, but does not show whether Muse gets an address or reads connected accounts.
- Sony and UMG sue Suno again, saying licensed v6 inherited old infringement
The labels say Mikey Shulman's licensed reset carried user feedback and model learnings from earlier systems into Suno's v6 lineup.
- SpaceXAI ships Grok Voice Transcribe 2.0 as Loom transcribes every video
SpaceXAI kept pricing at $0.10 per batch hour and $0.20 for streaming while claiming a 2x accuracy gain.
- Precis AI recruits policy wonks for Willard's citation-first beta
David Fuscus is extending his domain-specific AI play from PR into congressional work, where fluent errors carry statutory consequences.
- TypeSafe raises $40M for Jev, its AI model built to skip chat
TypeSafe AI's Jev made 47 Pong decisions in 12 seconds in an Ably demo; Gemini, Claude and GPT made two or three, while still choosing correctly in most runs.
- Dnotitia gets its vector-search ASIC back from fab before Q4 tests
Dnotitia reported a 5.77x FPGA gain, but silicon performance, power and cost remain untested ahead of planned evaluations.
- Anthropic and OpenAI seek smaller data centers while gigawatt sites wait
The labs are exploring 20-30 MW deployments as inference demand outruns the timetable for gigawatt campuses.
- LangChain adds TypeSafe's Jev to the agent control loop
The integration gives developers typed answers and confidence scores for routing, escalation and tool-call checks inside AI agent workflows.
- Vals ranks Tencent's Hy4 Preview fourth among open-weight models
The model finished 21st overall, but its $1.28 test cost and task-specific wins make aggregate rank only half the buying decision.
- Nunchux speeds video diffusion with value smoothing and FP8 softmax casting
Nunchux AI's Muyang Li and collaborators report up to 1.70x faster end-to-end generation without retraining, with MiniMax H3 among four models tested.
- Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools
The text-output model accepts text, images, audio and video across a 1M-token context window, with function calling and web search for work beyond analysis.
- Pruna AI promises five-second video clips in two seconds
The MiniMax H3-based API generates 5- to 15-second clips with audio, charging $0.02 to $0.075 per output second.
- GPT-6 Astra built a Mac aquarium wallpaper, battery optimizations pending
AI engineer Chase Lean's prototype turns the cursor into fish bait, though performance still depends on the machine running it.
- Thesys open-sources AppLess, its no-app phone demo, with OUI-1
The React Native experiment now uses a self-hosted diffusion model, though every booking, payment and order remains simulated.