Build1 publisher3 min readPublished
DeAlignAI's downloadable FP8 build is the license working exactly as written, while its self-reported 320-of-320 HarmBench run remains unchecked by any outside researcher and measures compliance, not capability.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The license is the whole mechanism. Z.ai's MIT terms let downstream users inspect the model, run it on their own infrastructure, change its behavior and ship it in products without paying Z.ai for generated tokens [13]. Refusal controls on a hosted endpoint govern only the traffic that reaches that endpoint, not a checkpoint sitting on someone else's GPU.
DeAlignAI's model card names its technique CRACK and describes it as changing refusal behavior directly in the weights, with the edit kept conservative to preserve quality [8]. That phrasing treats refusal as localized enough in parameter space to edit without retraining. The supplied research does not document a controlled comparison between Z.ai's original model and the FP8 variant under identical inference settings, and does not establish how much of the original capability survived [10]. "Conservative" is therefore a claim about a diff that only DeAlignAI has measured.
The 320-of-320 figure is a compliance count from DeAlignAI's own README, on a benchmark that scores how a model responds to harmful requests across areas such as cybercrime, disinformation and biological weapons [4][6]. There is no established independent replication or external methodological validation [5]. For that count to carry the meaning a security reader will attach to it, the grader's compliance criterion would need to match your threat model, the compliant answers would need to be correct as well as compliant, and the edited model would need to have retained the capability to act on them. None of that is shown here; HarmBench measures behavioral compliance with prompts rather than operational success against real systems [7][11].
The cyber numbers that would speak to capability belong to a different model. Z.ai reports 84.5% on CyberGym and 54.4% on ExploitBench for the larger GLM-5.3, and the GLM-5.3-Flash model card publishes neither result for the smaller model [12].
What does transfer is the shape of the artifact. Flash carries 320 billion total parameters and activates 18 billion per token [15], which is 5.6% of the weights live on any given forward pass [21]. Sparse activation sets what hardware can run the thing, and inference cost is what decides how widely a model gets operated outside its developer's infrastructure [20]. Z.ai's own scoreboard supports the cheapness story on its own harness: a discounted $0.045 per task on Artificial Analysis Intelligence Index v4.1.1 at a score of 57, and one-tenth the cost of GLM-5.2 [17]. On Z.ai's internal Code Bench, Flash scored 29.0 at maximum effort against 29.5 for Claude Opus 4.8 [18], and Z.ai reports 63.4 on DeepSWE v1.1 and 48.8 on AutomationBench versus 46.2 and 26.2 for GLM-5.2 [19], the last of those an 86% relative gain [22]. A vendor's internal bench, scored by the vendor, comparing the vendor's model to a competitor's, is a claim about the vendor's workload.
Runtimewire, citing an analysis on jyn.dev, frames the release the same way the evidence does: the download demonstrates the redistribution path, and the self-reported benchmark does not establish real-world offensive capability [23][24]. Jinho Jang's project also reports more than 200 controlled experiments and over 40 findings across nine models [9], which is a body of work no one outside it has audited either. The load-bearing fact needs no benchmark at all. A modified checkpoint is sitting behind a download link, under terms that permit precisely that [3][1].
Ranked by verification strength, evidence, and original report placement.
Z.ai, based in Beijing, released GLM-5.3-Flash on August 26th under an MIT license.
Independent AI researcher Jinho Jang has published altered weights for GLM-5.3-Flash, which his DeAlignAI project describes as having its guardrails removed.
DeAlignAI published an altered FP8 version of GLM-5.3-Flash with downloadable model files.
DeAlignAI's README reports 320 compliant responses from 320 HarmBench-320 prompts for its altered FP8 variant.
The 320-of-320 result comes from DeAlignAI's own evaluation and has no established independent replication or external methodological validation.
HarmBench-320 measures how models respond to harmful requests involving areas such as cybercrime, disinformation and biological weapons; DeAlignAI used it to measure whether its altered model complied rather than refused.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Artifacts checkable, numbers self-graded
Two different grades of fact sit side by side here. That the license is MIT, that DeAlignAI's FP8 files exist under Jang's name, that Flash's model card carries no CyberGym or ExploitBench result: all of that can be opened and read by anyone. Every figure that would tell you whether the edited model is dangerous, or still good, was produced by the party with a stake in it. Runtimewire flags this itself, which is why the score sits mid-range rather than low.
Redistribution proven once, uptake unknown
One researcher took the weights, edited refusal behaviour and republished them. That single derivative is the entire observed use of the license. No download counts, no named deployments, no product built on Flash or on the altered FP8 build appears anywhere in this reporting, so what has been demonstrated is that the redistribution path functions, not that anyone is walking it.
The label outruns the measurement
'Guardrails removed' is the loudest phrase in this story and the least tested. A perfect score on HarmBench-320 means the model answered 320 prompts; it says nothing about whether the answers would work against anything, and Runtimewire draws that line in its own copy rather than leaving it to a reader. The residual gap comes from DeAlignAI's framing and from the missing capability comparison, not from the outlet overselling what it has.
Both scorekeepers are also contestants
DeAlignAI's standing as a research project rests on the 320-of-320 number, and DeAlignAI ran the test. Z.ai's coding-plan business rests on Flash edging Claude Opus 4.8 and costing a tenth of GLM-5.2, and Z.ai ran those tests too, one of them on a bench it built. Nobody with a reason to find a smaller number has looked at either set, and Z.ai is not quoted responding to a third party republishing its weights with the refusals taken out.
Firm floor, single thread above it
The floor is solid: an MIT release on a known date, a third-party FP8 rebuild with files attached, a self-graded compliance run, and an outlet careful about where its knowledge stops. Above that floor everything runs through one publisher reading one personal blog, with Z.ai silent and Jang's run unrepeated, so conclusions about the altered model's behaviour or reach should be held loosely.
build
Post-training alone took GLM-5.3 from 4.6 to 28.3 on Terminal-Bench 3.08 publishers
build
Open weights caught up on finding bugs. They did not catch up on using them.1 publisher
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 publisher
Publishers with included, body-backed reporting in this cluster.
1 article · September 7, 2026