AI — Page 16
Models, agents, infra, applied AI.
- Ref founder Matt Dailey publishes a seven-step guide to avoiding AI design slop
Matt Dailey's method front-loads constraints and review before generated UI reaches Ref's production code.
- Intel gives coding agents 20 ways to stop guessing at Arc GPU flags
Kushal Mittal and nine Intel contributors packaged setup, sizing, benchmarking, profiling and CUDA migration into one Apache-2.0 repository.
- IBM brings its US Open AI demo back to Madison Square Park
The free September 10th-13th event pairs match screenings with a serve simulator, AI-generated ping-pong recaps and a photobooth beside IBM's Manhattan office.
- Thinking Machines' ReViSQL-K2.6 scores 91.37% on cleaned SQL benchmark
ReViSQL-K2.6 reached 88.55% with greedy decoding and 91.37% with 16-sample self-consistency on the corrected Arcwise-Plat-SQL benchmark.
- Anthropic wins order lifting federal ban on its technology
Judge Rita F. Lin ordered the administration to lift its ban on Anthropic technology after finding the exclusion inadequately justified, Bloomberg reported.
- Experiential Labs open-sources a router that learns from agent traffic
The Y Combinator Summer 2026 startup is betting that production traces can push routine work onto cheaper models customers own.
- Marvell hits record $2.739 billion revenue as AI-driven data center demand soars
The company raised its full-year outlook as custom silicon and optical interconnects outpace traditional segments of the business.
- Google DeepMind pilots double-blind AI evaluations without sharing prompts or model weights
Google DeepMind says its pilot keeps external evaluation prompts and model weights private, though its first-of-its-kind claim remains independently unverified.
- HeyGen details TAVR, a video-reference system for steadier AI avatars
First disclosed on March 3rd, the production research system uses up to 48 reference frames to preserve a person's identity across new scenes.
- Z.ai will release GLM-5.3 weights on August 28th after safety delay
The Tsinghua-born AI lab held back the coding model's checkpoints for two weeks after its cyber capabilities grew faster than expected.
- xAI puts Grok 4.6 on Microsoft Foundry after AWS and Google Cloud
The public preview gives Azure customers access to Grok 4.6 with a 500,000-token context window, extending xAI's multi-cloud distribution push.
- Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark
In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 (Fast) led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost.
- SpaceXAI tests voice calls for Grok Bot's persistent AI workers
A test interface shows users speaking directly with agents that already browse websites, operate software and run jobs in the cloud.
- Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94
In a 12-task Editorial Craft evaluation, Z.ai’s GLM 5.3 Flash ranked first with a 0.94 score and an estimated cost of $0.0002 per task. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard.
- Imageat serves unidentified uncensored Qwen derivative, leaving customers to test it
The third-party provider offers metered access without identifying the checkpoint, revision, modification method or endpoint evaluation behind its service.
- NVIDIA hits $96.2 billion revenue as Data Center demand surges 117%
The AI chip giant crushed Street expectations on top line and provided a massive $108 billion outlook for the next quarter.
- Latitude opens Voyage, betting AI roleplay needs rules after all
Nick Walton's open beta separates world state from AI narration, with free public multiplayer and paid private servers across web and mobile.
- Linux Foundation accepts OPAQUE's TRACE standard for AI runtime evidence
Imran Siddique's cryptographic receipt format has backing from AMD, Intel and Microsoft, though its governance and core verification work remain unfinished.
- x1 launches an AI iPhone builder that asks questions before writing code
Manil Lakabi's YC-backed tool plans, designs and packages React Native apps for testing and App Store review.
- Perplexity launches its Computer agent locally on Nvidia's $4,699 desktop
Portable Computer keeps models and files on a DGX Spark, while cloud escalation still requires approval and credits.
- Google ships Gemini 3.5 Transcribe for real-time speech apps
The public-preview model handles 85+ languages, custom jargon and up to three speakers across live and recorded audio.
- Anthropic lets outsiders study Claude usage without showing them the chats
Stanford, Oxford and METR analyzed three separate samples of about 250,000 conversations through Anthropic's privacy filter.
- Google announces Gemini 3.5 Transcribe, leaves its API identity unclear
Google says its latest speech-to-text model handles structured strings, filler removal, formatting and custom vocabulary, while pricing and the access path remain undisclosed.
- Hugging Face explains why DeepSeek wants lookup tables inside LLMs
Blackroot's new guide translates a January architecture that trades some MoE experts for hashed N-gram embeddings.