DeAlignAI publishes altered GLM-5.3-Flash weights, claims 320/320 harmful-prompt compliance

Jinho Jang's project made the modified model downloadable and reported compliance with every HarmBench prompt in its own test. The result has no established independent validation.

By · Published

Primary source: jyn.dev

Why it matters

Open weights let developers inspect and deploy capable models locally while giving downstream users the same ability to modify safeguards. DeAlignAI's downloadable artifact demonstrates that redistribution path, although its self-reported benchmark does not establish real-world offensive capability.

A glowing, intricate digital model has a protective layer precisely stripped away by a laser, revealing its vulnerable core.

Independent AI researcher Jinho Jang has published altered weights for GLM-5.3-Flash, an open-weight model from Beijing-based Z.ai, that his DeAlignAI project describes as having its guardrails removed. The downloadable artifact shows what a permissive model license allows, even if DeAlignAI's performance claims remain unverified by outside researchers.

Z.ai released GLM-5.3-Flash on August 26th under an MIT license. Jang's independent research project, DeAlignAI, subsequently published an altered FP8 version with downloadable model files.

Z.ai's open-weight release lets downstream users alter refusal behavior outside the controls of its hosted service. DeAlignAI says its altered FP8 variant complied with all 320 prompts in its HarmBench-320 run, but the supplied research does not independently validate that result.

What DeAlignAI actually tested

HarmBench-320 measures how models respond to harmful requests involving areas such as cybercrime, disinformation and biological weapons. DeAlignAI used the benchmark to measure whether its altered model complied rather than refused. A recent analysis on jyn.dev cited the result while examining the security implications of downloadable, modifiable model weights.

DeAlignAI's README reports 320 compliant responses from 320 HarmBench-320 prompts. The result comes from DeAlignAI's own evaluation and has no established independent replication or external methodological validation.

DeAlignAI's model card calls CRACK its method for changing refusal behavior directly in model weights and claims the edit was kept conservative to preserve quality. DeAlignAI says its broader research includes more than 200 controlled experiments and over 40 findings across nine models.

The supplied materials do not document a controlled comparison between Z.ai's original model and DeAlignAI's FP8 variant under identical inference settings. They also do not establish how much of the original model's capability survived the modification.

The available evidence supports a limited conclusion: DeAlignAI produced downloadable altered weights and reported full compliance in its HarmBench run. It does not establish that the altered model can autonomously compromise production infrastructure, develop reliable exploits or sustain access to a target. HarmBench measures behavioral compliance with prompts rather than operational success against real systems.

For the larger GLM-5.3 model, Z.ai reports scores of 84.5% on CyberGym and 54.4% on ExploitBench. Those company-reported results cannot be assigned to Flash. The GLM-5.3-Flash model card does not publish direct CyberGym or ExploitBench results for the smaller model.

What the MIT release permits

The model's open distribution gives downstream users the ability to change its refusal behavior outside Z.ai's hosted controls. Developers can inspect the model, run it on their own infrastructure, alter its behavior and integrate it into products without paying Z.ai for every generated token. Z.ai's published announcement presents that distribution as part of an effort to make coding and agentic workloads cheaper to operate.

GLM-5.3-Flash has 320 billion total parameters and activates 18 billion for each token, according to Z.ai. The company describes it as the first natively multimodal model in the GLM-5 series. Its hybrid sparse and linear-attention architecture supports a claimed one-million-token context window, and Z.ai says pretraining used a 30 trillion-token multimodal corpus.

Z.ai reports a discounted cost of $0.045 per task on Artificial Analysis Intelligence Index v4.1.1, where Flash scored 57. Z.ai says it costs one-tenth as much as GLM-5.2. On Z.ai's internal Code Bench, Flash scored 29.0 at maximum effort, compared with 29.5 for Anthropic's Claude Opus 4.8. Z.ai also reports 63.4 on DeepSWE v1.1 and 48.8 on AutomationBench, compared with 46.2 and 26.2 for GLM-5.2. These are vendor-reported benchmarks rather than independent evaluations.

The smaller active footprint is relevant to distribution because inference cost affects how widely a model can be operated outside its developer's infrastructure. It does not make the 320 billion-parameter model a routine laptop download, but it lowers the operating burden relative to Z.ai's larger models.

The company behind the weights

Z.ai was founded in 2019 as a spinout from Tsinghua University's Knowledge Engineering Group. Tang Jie and Zhang Peng are Z.ai's publicly identified founders.

Tang, a Tsinghua computer-science professor and Z.ai's chief scientist, created AMiner, an academic-search and researcher-network platform. Zhang worked on AMiner and related knowledge-graph projects before the research was commercialized. Z.ai, formerly known internationally as Zhipu AI, built its GLM model family from that academic base.

Dealroom reports at least $1.4 billion in private funding before Z.ai's public offering. Reported backers include Alibaba, Tencent, Meituan, Xiaomi, Ant Group, Qiming Venture Partners, Legend Capital, HongShan and Prosperity7 Ventures.

Z.ai markets GLM-5.3 for software engineering, coding and agent tasks through coding plans and compatible agent tools. Flash's pricing and benchmark claims support that sales pitch. Its permissive license also means downloaded copies no longer depend on Z.ai's hosted behavioral controls.

What the artifact establishes

DeAlignAI's release is evidence that a third party can redistribute modified GLM-5.3-Flash weights and claim substantially different refusal behavior. The downloadable files and DeAlignAI's attribution are verifiable. The reported 320/320 benchmark result remains a self-evaluation, and the supplied research establishes neither an independent replication nor real-world offensive performance.

Security teams evaluating open models should assume provider-defined refusals can be modified after the weights leave the provider's infrastructure. They should also resist treating a behavioral compliance score as proof of autonomous exploitation capability. The useful next step is controlled testing of the original and altered weights under matched inference settings, with operational security claims measured against real systems rather than prompt compliance alone.

Reader comments

Conversation for this story loads after sign-in.