Cantina releases open-weight security model trained on 50 vulnerability cases
Co-founded by Harikrishnan Mulackal, Cantina built Apex Flash-1 with Yeta Labs and says it solved 40 of 60 held-out security tasks.
By Ryan Merket · Published
Primary source: X
Why it matters
Cantina is turning its vulnerability research into training data and a downloadable model, coupling security work with a reusable AI asset. Its initial evaluation is company-run and limited to 60 tasks, while the abliterated release makes the dual-use tradeoff unusually explicit.

Harikrishnan Mulackal's Cantina released Apex Flash-1, an open-weights model post-trained for security research using vulnerability cases drawn from the company's work. The model, built with Yeta Labs, is available to download, along with a separate version modified to produce fewer refusals. Cantina described the release in an October 1st post on X.
Mulackal founded Cantina after building Spearbit, a network of security researchers that grew from digital-asset audits into competitions and bug-bounty programs. The company now sells autonomous security software as well as researcher-led work. Its leadership page lists Mulackal as CEO and co-founder and Mike Leffer as co-founder and president; Leffer's background includes more than a decade in defense and security operations and service in the U.S. Army. The model release turns that research operation into a second asset: training examples for software that can investigate vulnerabilities.
Cantina says Apex Flash-1 is a fine-tune of GLM-5.3-Flash, trained with full-parameter fine-tuning and group relative policy optimization, or GRPO. For the initial training run, it created 150 tasks from 50 vulnerability cases. Each case was turned into three versions with different levels of information: one gives the model source code and guidance, another provides source code with limited direction, and a third gives limited guidance and access to a running target without its source code.

That structure is intended to test whether an agent can do more than identify a suspicious code pattern. In one example Cantina describes, the model had access to one Forgejo workflow but needed to retrieve a protected artifact from another. It had to trace a mismatch between what a signed download URL authorized and what the application later fetched, then demonstrate the effect against a running system. Cantina says its verifier checked whether the model recovered the protected artifact, and that apparent successes were reviewed for shortcuts.
The model card on Hugging Face reports results from 60 tasks built from 20 held-out vulnerability cases. Cantina says Apex Flash-1 passed 40 tasks, for a 66.7% pass-at-one score. The base GLM-5.3-Flash passed 36, or 60%; Claude Opus 5 High passed 43, or 71.7%. Cantina estimates the run cost at $2.38 for Apex Flash-1, $4.56 for the base model and $74.68 for Opus 5 High, using provider pricing.

Those figures come from Cantina's own evaluation, not an independent benchmark. The test is also small: each model ran the same set once, and the result measures performance on Cantina's own constructed environments. The comparisons give buyers an initial look at cost and task performance, but they do not establish how the model performs across different codebases, security teams or agent setups.
Cantina positions Apex Flash-1 as a worker model that a larger agent directs toward a focused task. Its model listing identifies the checkpoint as 321 billion parameters and provides instructions for running it locally with common inference software. The license is listed as MIT. The separately released Apex Flash-1 Abliterated is described as an experimental derivative with modified refusal behavior; Cantina's reported evaluation applies to the standard model, not that version.
The open release also puts Cantina's safety position into practice. The company says local deployment gives defenders more control over data, tools and model behavior, and argues that attackers can already adapt open models while legitimate teams face restrictions on proprietary services. The abliterated variant makes that dual-use tension concrete: fewer refusals may help authorized researchers fit a model to their workflow, while also weakening a barrier against harmful requests. Cantina frames its intended users as security researchers working in authorized environments; the downloadable weights are not limited to that audience.
The timing follows a broader shift in Cantina's business. In July, the company announced $8 million in new funding led by Framework Ventures, bringing total funding to $16.5 million, and launched Clarion, its security operations platform. Cantina's account of the shift says years of bug-finding work exposed a bottleneck after discovery: teams still have to prioritize, fix and verify reported issues. Apex Flash-1 extends the discovery side of that strategy while using the vulnerabilities Cantina finds as a basis for model training. The next test is whether the approach transfers beyond 50 cases and a 60-task internal evaluation.