ARC Prize verifies DeepSeek V4 Flash at 61.4% for $0.04 per task

ARC Prize's outside evaluation documents how Liang Wenfeng's open-weight model falls from 61.4% at Max effort to 46.0% at Low, with Max costing $0.04 per task.

By · Published

Why it matters

ARC Prize's outside evaluation gives developers a checkpoint-specific measurement of DeepSeek V4 Flash's ARC-AGI-2 performance, reasoning-effort curve and benchmark cost.

A lone, incredibly efficient AI system's verifiable triumph over a monumental, abstract reasoning challenge (Oil painting, meticulously crafted to evoke the precise brushwork, strong light, and melancholic, human-scaled compositions of Edwa

ARC Prize verified DeepSeek V4 Flash 0731 at 61.4% on ARC-AGI-2's Semi-Private evaluation at maximum reasoning effort, with a reported cost of $0.04 per task. The result is tied to the July 31 checkpoint and includes three effort levels: 61.4% at Max, 56.0% at High and 46.0% at Low.

ARC Prize evaluated the named checkpoint through its verification program and reported both task performance and cost. That outside evaluation is the focus here. The checkpoint comes from DeepSeek, the Hangzhou AI company founded by Liang Wenfeng.

The ARC Prize results page marks the result as verified and identifies Max, High and Low reasoning variants. The page also reports 89.0% on ARC-AGI-1 Semi-Private at maximum effort, at $0.02 per task.

ARC Prize created its verified program because results can vary with prompting, dataset curation and testing methods. Its methodology treats resource efficiency as part of the measurement, which is why the result page reports cost alongside task completion.

The $0.04 figure applies to each ARC-AGI-2 task at maximum effort under this evaluation. It is not a general price for coding-agent jobs, research runs or other workloads.

Three effort levels expose the ARC-AGI-2 cost trade

ARC Prize tested Max, High and Low reasoning variants on the Semi-Private benchmark. V4 Flash scored 61.4% at Max, 56.0% at High and 46.0% at Low.

That 15.4 percentage-point spread shows how strongly the ARC-AGI-2 result depends on inference effort. A model's top score offers limited guidance for production planning unless operators can also see how much computation produced it. DeepSeek exposes reasoning effort as a parameter, allowing developers to adjust deliberation according to the latency and spending limits of a workload.

ARC Prize's $0.04 figure remains specific to the maximum-effort ARC-AGI-2 evaluation. Repeated tool calls, long context windows and large outputs can raise the cost of an application well beyond a single benchmark task.

DeepSeek's current API pricing lists V4 Flash at $0.14 per million uncached input tokens, $0.0028 per million cached input tokens and $0.28 per million output tokens. Those token rates explain how DeepSeek bills ordinary API use; they do not establish a universal per-task price.

Liang's efficiency discipline came from quantitative trading

Liang founded DeepSeek in 2023 after building High-Flyer, a quantitative hedge fund that applied machine learning to automated trading. He earned a bachelor's degree in electronic information engineering and a master's degree in information and communication engineering at Zhejiang University, according to Fortune.

That background helps explain DeepSeek's attention to the relationship between model capability and computation. Quantitative trading rewards improvements that survive a cost calculation. DeepSeek applies that discipline through model architecture, API pricing and open releases.

V4 Flash contains 284 billion total parameters, while activating about 13 billion for each token, according to DeepSeek's technical paper. Its mixture-of-experts design routes each token through a subset of the model instead of using every parameter for every request.

The V4 paper lists Liang among more than 300 authors and describes compressed sparse attention, heavily compressed attention, manifold-constrained hyper-connections and the Muon optimizer. The architecture supports a one-million-token context window, while DeepSeek's API documentation lists a maximum output of 384,000 tokens.

DeepSeek released the broader V4 family on April 24, according to its transparency center. ARC Prize dates its result for the V4 Flash 0731 checkpoint to July 31 and identifies Max, High and Low variants. Naming the checkpoint ties the verified scores to a specific model version.

Open weights widen the checkpoint's distribution

DeepSeek published the V4 Flash 0731 weights on Hugging Face under an MIT license. The model card provides deployment paths for Transformers, vLLM, SGLang and Docker, along with instructions for local inference.

The release gives DeepSeek two routes to adoption. Developers can use the hosted API, while infrastructure teams can operate the checkpoint themselves and modify the surrounding inference stack. Downloadable weights and permissive licensing allow the model to reach users running it outside DeepSeek's servers.

ARC-AGI has limits. It evaluates adaptation to unfamiliar visual reasoning tasks rather than the full range of reliability, security, coding and tool-use behavior required in production agents. ARC Prize lists no ARC-AGI-3 score for V4 Flash 0731. That benchmark moves toward interactive environments in which an agent must gather information and act over time.

The verified result answers a narrow question under documented conditions: DeepSeek V4 Flash 0731 solved 61.4% of ARC-AGI-2 Semi-Private tasks at maximum effort, at a reported $0.04 per task. The lower scores at High and Low effort show how much that result changes when the model spends less computation on reasoning.

Reader comments

Conversation for this story loads after sign-in.