Ant 发布了 Ling-3.0-flash,具有 256K 上下文,用于 AI 代理。

这款拥有1240亿参数的混合专家模型在每个令牌上激活51亿参数,并为工具驱动的工作流程原生支持256K上下文。

By · Published

Primary source: Aligned News - AI Intelligence

Why it matters

Ling-3.0-flash gives developers another low-active-compute option for long-context agents and tool calls. Its value will depend on whether Ant's sparse architecture reduces total workflow cost after retries, context processing and failed actions are counted.

Illustration of layered digital processes funneling into a jewel-like core labeled Ling-3.0-flash, representing Ant's 124B model with native 256K context for AI agents.

On July 24, 2026, Ant Group released Ling-3.0-flash, giving its Ling model family a sparse mixture-of-experts system built for long-running agent workflows, tool calls and high-throughput inference, according to the 官方发布页面. Ant 的文档 列出总参数量为1240亿,总计每个 token 激活51亿个参数,原生 256,000-token 上下文窗口,并可选择将上下文扩展到 1,000,000 token。

The model can switch between reasoning and non-reasoning modes, allowing an application to reserve additional computation for harder prompts. Ant 将 Flash 定位于工具调用和生产级 AI 代理工作负载.

面向代理工作负载的稀疏计算

Ling-3.0-flash uses a mixture-of-experts architecture, which routes each token through a selected portion of the model rather than activating all 124 billion parameters. Its 5.1 billion active parameters represent the immediate computational workload disclosed for each token, though total serving cost also depends on memory requirements, context length, output volume, hardware and the efficiency of the provider's implementation.

Ant claims Flash can reach inference speeds of up to 1,000 tokens per second and time-to-first-token below 100 milliseconds, according to its 模型文档. Those figures are vendor claims and were not independently established in the reviewed research. Performance inside an agent also depends on tool latency, prompt caching, retries and whether the model can complete a workflow without escalating the task to a larger system.

The 256,000-token native context gives developers room to pass large code repositories, document collections or long interaction histories into a single session. Ant says the window can be extended to 1 million tokens, although a larger context can raise memory use and processing cost even when the model activates a relatively small share of its parameters.

Flash is available through the Ant Ling API, OpenRouter and ZenMux, according to Ant's 发布页面 and 开发者资料. The documentation also lists integrations with Claude Code, OpenClaw, Kilo Code and Hermes Agent.

Those channels give developers several ways to test the model without operating inference hardware. Ant has not disclosed Ling-specific customer counts, paid API volume, revenue or model-download figures in the materials reviewed for this story, leaving limited evidence about adoption beyond the published integrations and distribution channels.

Ling 作为 Ant 的内部研究项目

Ling is an internal Ant Group research and product initiative. Ant's Ling model documentation describes Ling as an independently developed, open-source general-purpose model series. The research is attributed to a collective Ling Team. A recent Ling and Ring technical report lists Ang Li and Ben Liu first among hundreds of authors, but the public materials do not identify a single project founder or establish that either researcher leads the program.

Ant Group grew out of Alipay, Alibaba's payment service established in 2004, and traces its corporate formation to 2014. The company is headquartered in Hangzhou, China. Ling is funded through Ant's corporate research and product operations; no Ling-specific investors, valuation or capital raised have been disclosed.

Ant's published roadmap places Ling 1.0 in March 2025, Ling 2.0 in October 2025, Ling 2.5 in February 2026 and Ling 2.6 in April 2026, followed by Ling 3.0 in July 2026. Flash therefore extends a model program Ant has revised several times in less than 18 months.

竞争正转向以激活计算为主

Ant is competing with model developers that increasingly market sparse architectures and smaller serving footprints. Alibaba's Qwen3-30B-A3B has 30 billion total parameters and activates 3 billion, while also supporting thinking and non-thinking modes. Qwen launched that generation in April 2025 as part of an open-weight lineup spanning several model sizes.

Other companies are approaching lower-cost inference from different directions. Liquid AI offers LFM2 models ranging from 350 million to 2.6 billion parameters for on-device and on-premise use with a non-transformer architecture. IBM's Granite 4.0 includes small and hybrid mixture-of-experts variants aimed at constrained hardware. Fastino, which develops small task-specific enterprise models, 在 2025 年 5 月筹得由 Khosla Ventures 领投的 1,750 万美元.

Ling-3.0-flash's commercial case will depend on completed work per dollar. Low active-parameter counts can reduce the cost of individual calls, but agent systems often incur hidden expense through retries, failed tool calls and repeated context processing. Ant has supplied the architecture, access routes and vendor performance targets. Production usage data will determine whether Flash can carry long agent sequences at the speed and cost its specifications imply.

Reader comments

Conversation for this story loads after sign-in.