Build1 distinct publisher3 min readUpdated
Jinho Jang's Qwen3.8-27B-CRACK-GGUF packages abliterated multimodal weights with seven quantizations and a vision projector. The control point is now your inference hosts, not a vendor contract.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
An independent researcher, Jinho Jang, has published what his project dealign.ai describes as a safety-modified, 27-billion-parameter derivative of Qwen3.8-27B, packaged for local image, video and text inference through llama.cpp [1]. The repository, Qwen3.8-27B-CRACK-GGUF, ships seven quantizations, a vision projector and llama.cpp commands [2][6], which turns refusal-stripped weights of this class into a provisioning question for whoever runs your inference hosts rather than something you settle in a vendor contract.
Hugging Face metadata reproduced by dealign.ai dates the repository to August 12, 2026 [3]. The model card calls the model "abliterated," a weight-modification process intended to remove refusal behavior, and says the result completes requests a stock version declines while keeping coding, reasoning and multimodal capability [4][5]. Those retention claims and the repository's benchmark results come from dealign.ai's own testing [5]. The specifications are project-reported too, with no independent validation: a dense 27B vision-language hybrid, 64 layers, a stated 262K-token context window, and a Multi-Token-Prediction head in every GGUF [9], with 48 GatedDeltaNet linear-attention layers, 16 full-attention layers and a hidden size of 5,120 [10].
The distribution detail matters more than the architecture. The file table lists seven language-model quantizations from a 10.5 GB IQ2_M to a 29.0 GB Q8_0, plus a separate 0.9 GB F16 vision projector that pairs with any of them [11][13]; dealign.ai recommends the 17.0 GB Q4_K_M as its size-performance balance [12]. Full multimodal at the recommended size is therefore a 17.9 GB download [23], and 11.4 GB at the smallest [24]. The usage instructions cover llama.cpp's terminal client and its OpenAI-compatible local server [14]. That last part is the operational consequence: once running, the thing speaks the same wire protocol your internal gateway already routes to, and it leaves no trace in a hosted provider's logs.
The headline benchmark measures willingness, not quality. On HarmBench-240 the card reports 98.8% compliance for the Q8_0, Q6_K_L, Q6_K and Q4_K_M builds, and 97.5% for IQ2_M [16], a spread of 1.3 points [25]. A high score there measures how well the modification elicits answers the stock system was configured to reject; it is not a conventional safety or quality score [17]. Reported post-modification MMLU runs from 76.0% for IQ2_M to 83.4% for IQ4_XS [19], a 7.4-point band [26], and those figures do not establish performance on real applications, long video or extended agent tasks [19]. No independent evaluation of the HarmBench or MMLU results was located [20]. The card labels the release a research artifact with reduced guardrails, limits stated intended use to research and authorized red-teaming, and places responsibility for lawful use on the operator [18].
This is not an isolated artifact. Heretic automates abliteration while optimizing for fewer refusals and lower KL divergence, OBLITERATUS packages refusal removal and ablation studies into an open-source toolkit, and Orion-zhen's Abliteration implements refusal-direction removal plus norm-preserving variants [21]. Jang, who describes himself as an ML research and systems engineer in Irvine, California, has also published MLX Studio and the JANGQ variable-bit quantization project [7]. Dealign.ai reports more than 200 controlled experiments, over 40 findings and work across nine models from 0.8 billion to 397 billion parameters, all self-reported [15], and discloses no funding, investors, pricing, revenue or customer count on its public pages [8].
Two things worth tracking: whether anyone outside the project reproduces the compliance and MMLU numbers [20], and Jang's stated research line on how safety behavior is distributed across layers and how quantization weakens, preserves or restores it [22]. If quantization itself moves refusal behavior, then the artifact your gateway serves and the artifact you evaluated are not necessarily the same model.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Jinho Jang, the independent researcher behind dealign.ai, published a model his project describes as a safety-modified, 27-billion-parameter derivative of Qwen3.8-27B that packages image, video and text processing for local use through llama.cpp.
The repository identifies the release as Qwen3.8-27B-CRACK-GGUF.
Hugging Face metadata reproduced by dealign.ai lists the repository as created on August 12, 2026.
The repository describes the model as "abliterated," a weight-modification process intended to remove refusal behavior.
The model card says the resulting system will complete requests that a stock version declines while retaining coding, reasoning and multimodal capabilities; those retention claims and the repository's benchmark results come from dealign.ai's own testing.
The release combines refusal-removal research with a practical local distribution package: seven quantizations, a vision projector and llama.cpp commands, prepared for local inference without requiring a hosted API.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Artifact well documented, performance and safety claims entirely self-reported
The existence, packaging and file inventory of the release are concretely documented — repository name, creation date, quantization sizes, projector, runtime commands — but every behavioral claim (refusal removal, capability retention, HarmBench compliance, MMLU) originates from dealign.ai's own testing, the article states no independent evaluation was located, and only one publisher covers the story.
Published and downloadable, no usage evidence
Adoption evidence stops at availability: the repository is created and the quantizations plus projector are downloadable with documented llama.cpp usage paths. The supplied source reports no download counts, forks, deployments, dependent projects or third-party usage, so nothing beyond publication can be measured.
Mildly overstated by the artifact's own numbers, partly offset by cautious framing
The project's headline figures — 98.8% HarmBench compliance and retained coding, reasoning and multimodal capability — are unverified self-measurements, which pushes the claim side ahead of the evidence side. The gap stays modest because the coverage itself repeatedly attributes those numbers to dealign.ai, reframes a high HarmBench score as refusal-removal success rather than quality, and notes the MMLU band does not establish real-application performance.
Author-published benchmarks for an author-published artifact
The only numbers in the story are produced by the same party that built and distributes the model, and they serve as a demonstration of dealign.ai's safety-modification research program, including its self-reported experiment counts. Offsetting factors are that no funding, pricing, revenue or customers are disclosed, so no direct commercial upside is visible, and the card voluntarily narrows intended use and shifts liability to the operator.
Confident on packaging, weak on behavior
Confidence is high that the artifact exists as described and is downloadable in the stated configurations, since those details are specific and internally consistent. It is low on how the model actually behaves after modification, because a single publisher relays a single project's unreplicated measurements and no adoption or third-party assessment is available.
build
Unsloth's 10% quant claim is really about which machines can run a 27B model1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
leadership
You Procured Qwen. Your Edge Boxes Are Running Somebody Else's File.1 distinct publisher
build
Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 16, 2026