AutoTrust AI releases JEV-27B, adding decisions without rewriting Qwen

Co-founder Daniel Tang's AutoTrust AI added a compact decision block to a frozen Qwen3.8-27B model, keeping typed decisions and ordinary generation on one set of weights.

By · Published

Primary source: PR Newswire

Why it matters

JEV-27B makes a practical infrastructure bet: keep agent decisions local and fast by adding a specialized decision path to a general model. Its benchmark edge over a hosted alternative is slim and company-measured, so deployment-specific testing matters more than the headline average.

A close-up view of a sophisticated server rack system with cool blue indicator lights, focusing on a central module with a subtle warm glow.

AutoTrust AI released JEV-27B on September 28th, a self-hosted model that adds fast, structured decision-making to an otherwise unchanged 27B language model. For co-founder and CEO Daniel Tang, the bet is that AI agents need a reliable way to make frequent choices - classify a message, route a ticket, select a tool - without sending every small decision to a hosted service or asking a general-purpose model to generate an answer from scratch.

Tang brings an academic and engineering background to that problem. He completed a PhD at the University of Luxembourg in 2025 and was first author of CodeAgent, a multi-agent framework for autonomous code review accepted at EMNLP 2024. He also created ScienceGuru, AutoTrust AI's scientific-work platform. AutoTrust AI describes itself as an applied AI lab working on agent systems and efficient deployable models.

JEV-27B's design reflects a practical constraint in agent infrastructure: a model used to make decisions often has to coexist with one that can reason and write. AutoTrust says JEV-27B handles both through one set of weights. Its ordinary generation path comes from a frozen Alibaba Qwen3.8-27B backbone; a detachable decision block handles yes-or-no, multiple-choice and 0-to-5 rating questions in a single forward pass, returning a probability for each option. The model is released under the Apache-2.0 license and its weights and files are on Hugging Face.

A small addition to a large model

AutoTrust says the trained decision block contains 108.9 million parameters, around 0.4% of the full model. AutoTrust reports that training took about 9.2 hours on one NVIDIA B200, also describing the run as approximately 9.2 B200-hours. With the decision block switched off, all 164 HumanEval completions matched the base model byte for byte, according to AutoTrust. That test shows the added capability can be turned off without changing the tested coding outputs; it does not establish identical behavior for every prompt or deployment.

The model is intended for cases where a system repeatedly needs an answer in a known format. A support agent could classify a request or decide whether to escalate it; a coding workflow could check whether a change violates a stated rule. Probability outputs also let operators choose thresholds for handing uncertain cases to a slower model or a person. AutoTrust suggests this routing pattern, while noting that it has not benchmarked the combined workflow.

The hardware claim needs context. AutoTrust's headline deployment uses a single B200, and the model repository lists roughly 54.7 GB of files. The setup requires substantial accelerator hardware; it does not establish that the model runs cheaply on an ordinary server. In a technical post published September 27th, AutoTrust also described a full-precision deployment on one H100 80 GB, with lower reported capacity. Local deployment shifts inference onto customer infrastructure, leaving hardware and serving costs with the customer. That is relevant to buyers comparing an API bill with the cost of operating a GPU fleet.

Benchmark lead, with a narrow margin

AutoTrust reports an equal-weight average of 84.07% across six public decision benchmark groups. It also ran hosted TypeSafe Jev 1.13 on the same groups and reports an average of 83.85%, a difference of 0.22 percentage points. JEV-27B scored higher on four groups and lower on two. AutoTrust ran both sides of that comparison, so the small lead is AutoTrust-produced evidence, not independent confirmation that JEV-27B is broadly superior.

The strongest external check in the release is an independent human-labeled benchmark of decisions with up to 16 answer options. There, JEV-27B reached 96% of TypeSafe Jev 1.13's accuracy. It remains behind the hosted model. AutoTrust trained JEV-27B on probability distributions produced by Jev 1.13: the 25,376-question test measuring how closely the models' distributions match uses the teacher's own outputs as labels. That test demonstrates how faithfully the student reproduces its teacher; it does not independently establish that either model is right.

AutoTrust says the model inherits known weaknesses of its teacher, including multi-hop reasoning, arithmetic and adversarial inputs. It also cautions that benchmark comparisons are mostly its own runs and that performance varies by task. A 137-millisecond median decision latency and about 130 decisions per second on one B200 are likewise AutoTrust measurements. AutoTrust's comparison with hosted API latency is not like-for-like: its local result excludes network time, and the hardware and serving conditions differ.

The deployment bet

AutoTrust is following its earlier JEV-9B with a larger model built around the same split between typed decisions and ordinary generation. The 27B version makes a particular case to organizations that want decisions to stay inside their own infrastructure, especially where every request to an outside API carries cost, privacy or control concerns. Tang's release statement frames AutoTrust AI's value as a decision path that runs alongside a reasoning model on customer-owned infrastructure.

The open weights make the system inspectable and deployable, while the benchmark record still calls for care. Operators will need to test it against their own inputs, choose confidence thresholds, and keep human review for consequential decisions. JEV-27B's contribution is a compact specialized capability attached to a general model. Its promise depends on whether teams can make that arrangement useful at the cost of running a full 27B backbone.

Reader comments

Conversation for this story loads after sign-in.