Vals ranks Tencent's Hy4 Preview fourth among open-weight models
The model finished 21st overall, but its $1.28 test cost and task-specific wins make aggregate rank only half the buying decision.
By RuntimeWire Staff · Published
Primary source: Vals AI
Why it matters
Hy4 Preview shows why model selection is moving from one leaderboard score to cost per completed workflow. That shift gives cheaper open-weight models room to win specific jobs without claiming the overall crown.

Rayan Krishnan and Langston Nashold built Vals AI to grade models on professional work. Their latest result gives Tencent's Hy4 Preview a useful distinction: fourth place among open-weight models, with a bill low enough to make its uneven scores worth examining.
Vals said in a September 17th thread that Hy4 Preview scored 55.40% on the Vals Index, ranking No. 21 of 58 models overall. At $1.28 per test, it finished behind DeepSeek V4.1 Flash, Moonshot AI's Kimi K3 and Z.ai's GLM 5.3 among models with downloadable weights.
Krishnan and Nashold met as Stanford computer science students on a pre-orientation backpacking trip. Krishnan later worked as a software engineer at Palantir and was preparing for Ph.D. applications, while Nashold had worked across Hudson River Trading, Facebook, NVIDIA, Snap and Salesforce. ChatGPT's arrival clarified their shared premise: model capability would advance faster than dependable methods for measuring what those models could accomplish inside a workplace.
That premise has become increasingly valuable as labs release models with self-selected benchmarks, settings and competitors. Hy4 Preview offers a clean example. Tencent's own August 28th release presented the model as a productivity flagship for coding, office work and scientific research. Vals' independent evaluation produced a narrower conclusion: Hy4 Preview can be economically attractive on certain agent tasks, though it remains far from the strongest model across the full test suite.
A price result, with limits
The strongest number came from Code Migration, which tests whether a model can reimplement a working program in another language. Hy4 Preview scored 47.43%, ranking seventh among 61 models and first among the open-weight entries tested by Vals. Each test cost $3.41, compared with $24.91 for GLM 5.3.
Hy4 Preview also scored 68.61% on Public Benefits Bench, placing fourth among 36 models behind three Anthropic systems. The benchmark measures answers to questions about Supplemental Nutrition Assistance Program benefits, a task where an incorrect response carries a different cost from a weak chatbot summary.
The reasoning results were competitive without reaching the top of the table. Hy4 Preview scored 75% on ProofBench v1.1 and 59.33% on the International Olympiad in Informatics benchmark. It placed 10th of 61 on Legal Research Bench and 13th of 62 on Harvey's Legal Agent Benchmark.
The weaker scores matter for anyone considering Hy4 Preview as a general agent engine. It ranked 42nd of 66 on Terminal-Bench 2.1, 37th of 93 on medical coding and 35th of 95 on medical scribing. Vals also measured medium-high latency for the model's size and price, with its model page showing 47 minutes and 51 seconds at low concurrency across the Vals Index evaluation.
Cost per test is a more useful measure than token pricing alone because agent workflows can vary sharply in length, tool use and failed attempts. It is still tied to the evaluator's harness and task mix. Vals accessed Hy4 Preview through OpenRouter, using a temperature of 0.9, Top P of 1 and a maximum output length of 64,000 tokens. Different providers, prompts and agent scaffolds can produce different economics.
Tencent built Hy4 around work products
Tencent released Hy4 Preview on August 28th with 770 billion total parameters and 49 billion activated for each token. The mixture-of-experts model supports a 1-million-token context window, while its Apache 2.0 weights and deployment instructions cover vLLM and SGLang.
Tencent priced API access at $0.834 per million input tokens and $2.501 per million output tokens. The Hy team says it trained the model with data developed alongside Tencent specialists in software engineering, game development, finance and security, then co-designed it with products including CodeBuddy and WorkBuddy.
That product connection is central to Tencent's model strategy. RuntimeWire reported in August that Tencent brought Hy3 to international users through WorkBuddy, Miora and TokenHub, combining downloadable weights with distribution through Tencent's own applications and cloud services. Hy4 Preview extends the same play: make the weights available to developers while using Tencent's productivity products to collect feedback and turn model improvements into features.
Tencent is also unusually direct about the preview's rough edges. The Hy4 repository says the model can spend longer than necessary reasoning through complex tasks and tends to over-verify its work. Those limitations line up with Vals' latency findings and explain why the cheapest capable model can still become expensive when completion time matters.
Tencent says Hy4 Preview participated in optimizing parts of its own training and inference process, including experiments that improved end-to-end throughput by 31.8% from Tencent's baseline. That figure comes from Tencent's internal work rather than an independently reproduced test, so it should be read separately from the Vals results.
Vals is selling the scorecard
For Krishnan and Nashold, each new model release is also a demonstration of Vals' role between model labs and customers. Vals uses private datasets developed with domain experts to reduce the risk that benchmark questions enter training data. Vals also grades accuracy alongside cost, latency and failure patterns, giving buyers a view that public exam-style scores rarely provide.
Private test sets protect benchmark integrity while limiting outside reproduction of the headline rankings. Vals publishes its methodology and has open-sourced parts of its evaluation infrastructure, but independent researchers cannot inspect every underlying task that produced Hy4 Preview's 55.40% index score. The ranking is evidence from Vals' controlled harness, rather than a universal ordering of model quality.
The market has rewarded the founders' attempt to become AI's independent scorekeeper. On August 13th, Vals announced a $40 million Series A at a $400 million valuation, led by Andreessen Horowitz with 8VC, Pear VC, Bloomberg Beta, HRT Ventures and NextLadder Ventures participating. Vals said revenue in 2026 had reached eight times its total for 2025, while its customer base doubled and its staff tripled over six months. Those operating figures are self-reported.
Hy4 Preview makes the founders' case in practical terms. A model ranked 21st overall can still be the rational choice when it performs well on the exact workflow being purchased and completes that work for a fraction of a rival's cost. It can also be a poor fit for terminal operations or medical administration. Krishnan and Nashold are betting that the widening gap between those outcomes will turn evaluation from a launch-day marketing exercise into permanent infrastructure.