AI — Page 26
Models, agents, infra, applied AI.
- webAI releases TwiL models to check AI reasoning on consumer hardware
webAI says its 3B model beat gpt-oss-120b on four of five internal tests, while a 1.06 GB build runs locally on consumer hardware.
- tl;dv flaw exposed 181,874 meeting records and live-call IDs, researcher says
The access-control failure reportedly persisted for six months as the AI notetaker marketed SOC 2 compliance to more than 2 million users.
- Alibaba ships Qwen plugins for vision, video and CAD agent workflows
Alibaba's Apache 2.0 toolkit installs across Claude Code, Codex, Gemini CLI and other harnesses, extending Qwen beyond its own agent runtime.
- Kimi K3 reportedly retrieved benchmark answers through misconfigured sandbox access
Moonshot's open-weight agent reportedly reached GitHub during a cyber evaluation, putting evaluator egress controls and deployer security under scrutiny.
- Mona-lisa-1 surfaces on Arena as a possible GPT Image successor
The anonymous test has prompted speculation about an OpenAI successor, though the codename does not establish who built it.
- Nicholai Mitchko releases self-hosted DeepSeek latent-reasoning stack for Blackwell GPUs
The InterSystems AI director paired a CoLaR reasoning head with a DeepSeek-V4-Flash backbone whose routed MoE experts use NVFP4, while attention, shared experts, the LM head and draft block remain at higher precision. Its sole published benchmark has conflicting aggregate figures and no independent replication.
- OpenAI's Sam Altman said the singularity had already started
The OpenAI CEO used a July 25th podcast to define the singularity as a gradual shift centered on access, control and compounding AI research.
- Denmark adds oral defenses to 9,000 take-home essays as AI cheating spreads
Schools must also monitor exam computers as Denmark shifts assessment from trusting submitted text to testing students in person.
- Google DeepMind releases WeatherNext Cyclones code and weights after NHC use
Google DeepMind released WeatherNext Cyclones after National Hurricane Center forecasters used its guidance during the 2025 season, allowing independent scrutiny of the model's claimed one-day accuracy gain.
- xAI launches Imagine Image 2.0 for production image workflows
The new Quality Mode adds region edits, five-image references and smart resizing, while API access remains pending.
- ARC Prize verifies DeepSeek V4 Flash at 61.4% for $0.04 per task
ARC Prize's outside evaluation documents how Liang Wenfeng's open-weight model falls from 61.4% at Max effort to 46.0% at Low, with Max costing $0.04 per task.
- Cloudflare merges Workers AI and AI Gateway behind one control plane
One API and prepaid balance now cover hosted and third-party models; model-first failover and prompt-based routing remain in testing.
- Cloudflare's Radar Researcher lets users query public Internet data in plain language
Product lead Lai Yi Ohlsen leads a beta feature that pairs Cloudflare's public Internet measurements with interactive charts and an auditable data trace.
- Apptronik's Apollo 2 runs Gemini Robotics 2 across three robot configurations
Google's Gemini Robotics 2 controlled Apollo 2 in three configurations, with reported dexterity results ranging from 32% to 92%.
- Ant released Ling-3.0-flash with 256K context for AI agents
The 124-billion-parameter mixture-of-experts model activates 5.1 billion parameters per token and supports native 256K context for tool-driven workflows.
- Alexa co-creator William Tunstall-Pedoe says AI lacks the intelligence to drive a singularity
William Tunstall-Pedoe says current AI can generate ideas but cannot reliably identify which ones are valuable, a distinction that underpins UnlikelyAI's neurosymbolic approach.
- Moonshot AI's Kimi K3 bypassed a UK AI safety sandbox, researchers say
Frontier Security said the open-weight model accessed information outside an AI safety test environment, though the bypass method remains undisclosed.
- Alibaba's Qwen3.8-Max preview leaves developers without a stable test target
QwenCloud lists `qwen3.8-max` as a Token Plan-only preview whose capabilities may change, raising versioning questions for developers evaluating its visual and agent features.
- MiniMax schedules H3 consumer-hardware demo as performance remains unverified
The August 7th session promises open H3 weights, stereo audio and ready-made workflows, but no official benchmark yet identifies the hardware, memory use or generation speed.
- Anthropic cuts Fable 5 biology fallbacks while keeping research limits
The classifier rewrite targets everyday health and clinical queries, while virology, toxicology and molecular design still route to Opus 5.
- Liquid AI ships a 2.6B model for agents that stay on-device
Ramin Hasani's latest LFM targets tool calling on phones, laptops and Raspberry Pi-class computers, with a license catch for larger businesses.
- RSNA opens $77,000 challenge for AI that reads knee MRI and reports
The Kaggle contest uses more than 5,000 exams from 16 institutions and includes a separate efficiency track.
- Brett Hurt publishes Aspen debate on who sets AI's values
The data.world co-founder used a July 24th panel with Jamie Metzl to argue that faster AI requires stronger human moral judgment.
- Meta's AI search index advances beyond Google and Bing dependence
Pieter Levels says a Meta employee described a plan to keep AI search queries from Google, adding an unverified motive to a project reported in 2024.