AI — Page 20
Models, agents, infra, applied AI.
- GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score
In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models.
- Daniel Vaughn publishes Huzzah, an AI editor built around persistent pseudocode
The experimental editor treats persistent `.hz` pseudocode as a human-authored specification and regenerates source code from changes to it.
- Meta says Muse Spark 1.2 gains 12.2 points when given tools
Alexandr Wang's MSL reports that its model trails its predecessor without tools and pulls ahead when paired with Meta's runtime.
- Monk says its AI accounts receivable platform now manages $2B
George Kurdin is pairing a customer payment portal with Stripe distribution as Monk pushes finance teams to consolidate invoice-to-cash work.
- OpenAI expands Computer History to Mac users across Europe
OpenAI's opt-in timeline records selected activity from apps and websites, while admins must enable it for business accounts.
- Metric releases Armenian speech benchmark; closed systems take first eight spots
Hrant Davtyan's lab tested nearly 30 systems on 20.7 hours of Armenian audio across five speech datasets, with closed systems leading the combined ranking.
- Mistral launches Agentic Search so enterprise AI can keep looking
Arthur Mensch's Mistral says its retrieval loop lifted FinanceBench accuracy from 26.7% to 86% in a company-run test.
- TrueFoundry open-sources TrueForge to put its gateway beneath more AI agents
TrueFoundry reports lower costs than Claude Managed Agents, though its largest claimed saving also comes from switching models.
- Replit launches paid Free Mode with OpenAI's cheapest GPT-5.6 model
Core and Pro subscribers get credit-free everyday Agent tasks, with five-hour limits and up to 30 hours of monthly chat on Core.
- OpenAI pitches Codex for tax prep after a 7,000-return pilot
The 2025 tax-year pilot saved accountants 31% of preparation time, according to Current, while keeping practitioners responsible for review.
- Cloudflare engineer proposes web apps whose users generate the missing features
Jeremy Morrell argues LLM-written extensions and sandboxed runtimes could serve workflows too specific for a conventional product roadmap.
- Court stalls Google's $10M Spirit data deal over worker privacy
The sale covers 100 million emails and 500 million Teams items; a union says stripping names leaves confidential employment content exposed.
- SpaceX pursued Cognition after closing its $60B Cursor acquisition
Scott Wu's $26B coding-agent company stayed independent, while the two sides kept discussing a compute deal that echoes SpaceX's Cursor playbook.
- Inco AI releases DFlash 2, reports 21% longer accepted token sequences
Inco AI adds a path selector and local convolution to DFlash 2, reporting 21% longer accepted drafts with 1.3% added cycle latency.
- NVIDIA research veteran Sanja Fidler launches Veeda AI for robot world models
The former NVIDIA AI research vice president is building Veeda with Zan Gojcic and Huan Ling around simulated environments for embodied agents.
- OpenAI previews cross-session safety checks designed to preserve zero data retention
Private Safety Processing will trace risk patterns across related API interactions while giving OpenAI only limited safety signals.
- Unsloth says Dynamic 3.0 quants beat rivals by 10% at the same size
Daniel and Michael Han's post-training quants target local AI users, with the headline gains measured through Unsloth's own tests.
- Ornith AI ships open models that write their own training curriculum
The three-model family extends self-scaffolding into task generation; its benchmark comparisons come from Ornith AI's own evaluation runs.
- xAI puts Grok 4.6 on Amazon Bedrock one week after launch
xAI's flagship model brings a 500,000-token context window to AWS at $2 per million input tokens and $6 per million output tokens.
- Fei-Fei Li says World Labs wants AI to understand space beyond chatbots
The World Labs co-founder is betting on persistent 3D environments for creators, designers and robotics researchers.
- Vivodyne opens human-tissue data center to train drug-discovery AI
Andrei Georgescu and Dan Huh's robotic lab says each HIVE can test 10,000 tissues, generating human biological data before clinical trials.
- Flock Safety's AI searches can start with behavior, not a suspect
WIRED reports that OS Investigate lets police search vehicle movements and connected records by behavior, extending Garrett Langley's camera network beyond plate lookups.
- Modular open-sources Mojo three weeks after Qualcomm acquisition
The Apache 2.0 release includes Mojo's compiler and tooling as Qualcomm-owned Modular argues its AI platform can serve rival chips.
- Palomar opens a Lean proof registry for the AI math pileup
The project registers fixed GitHub snapshots, checks Lean proofs mechanically, and uses an LLM to compare formal statements with informal descriptions.