Skip to content

Build1 publisher3 min readPublished

A refusal-stripped 27B model now ships as a 17.9 GB llama.cpp pull

Jinho Jang's Qwen3.8-27B-CRACK-GGUF packages abliterated multimodal weights with seven quantizations and a vision projector. The control point is now your inference hosts, not a vendor contract.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Jinho Jang, the independent researcher behind dealign.ai, published a model his project describes as a safety-modified, 27-billion-parameter derivative of Qwen3.8-27B that packages image, video and text processing for local use through llama.cpp.
  • The repository identifies the release as Qwen3.8-27B-CRACK-GGUF.
  • Hugging Face metadata reproduced by dealign.ai lists the repository as created on August 12, 2026.
  • The repository describes the model as "abliterated," a weight-modification process intended to remove refusal behavior.
  • The model card says the resulting system will complete requests that a stock version declines while retaining coding, reasoning and multimodal capabilities; those retention claims and the repository's benchmark results come from dealign.ai's own testing.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An independent researcher, Jinho Jang, has published what his project dealign.ai describes as a safety-modified, 27-billion-parameter derivative of Qwen3.8-27B, packaged for local image, video and text inference through llama.cpp [1]. The repository, Qwen3.8-27B-CRACK-GGUF, ships seven quantizations, a vision projector and llama.cpp commands [2][6], which turns refusal-stripped weights of this class into a provisioning question for whoever runs your inference hosts rather than something you settle in a vendor contract.

Hugging Face metadata reproduced by dealign.ai dates the repository to August 12, 2026 [3]. The model card calls the model "abliterated," a weight-modification process intended to remove refusal behavior, and says the result completes requests a stock version declines while keeping coding, reasoning and multimodal capability [4][5]. Those retention claims and the repository's benchmark results come from dealign.ai's own testing [5]. The specifications are project-reported too, with no independent validation: a dense 27B vision-language hybrid, 64 layers, a stated 262K-token context window, and a Multi-Token-Prediction head in every GGUF [9], with 48 GatedDeltaNet linear-attention layers, 16 full-attention layers and a hidden size of 5,120 [10].

The distribution detail matters more than the architecture. The file table lists seven language-model quantizations from a 10.5 GB IQ2_M to a 29.0 GB Q8_0, plus a separate 0.9 GB F16 vision projector that pairs with any of them [11][13]; dealign.ai recommends the 17.0 GB Q4_K_M as its size-performance balance [12]. Full multimodal at the recommended size is therefore a 17.9 GB download [23], and 11.4 GB at the smallest [24]. The usage instructions cover llama.cpp's terminal client and its OpenAI-compatible local server [14]. That last part is the operational consequence: once running, the thing speaks the same wire protocol your internal gateway already routes to, and it leaves no trace in a hosted provider's logs.

The headline benchmark measures willingness, not quality. On HarmBench-240 the card reports 98.8% compliance for the Q8_0, Q6_K_L, Q6_K and Q4_K_M builds, and 97.5% for IQ2_M [16], a spread of 1.3 points [25]. A high score there measures how well the modification elicits answers the stock system was configured to reject; it is not a conventional safety or quality score [17]. Reported post-modification MMLU runs from 76.0% for IQ2_M to 83.4% for IQ4_XS [19], a 7.4-point band [26], and those figures do not establish performance on real applications, long video or extended agent tasks [19]. No independent evaluation of the HarmBench or MMLU results was located [20]. The card labels the release a research artifact with reduced guardrails, limits stated intended use to research and authorized red-teaming, and places responsibility for lawful use on the operator [18].

This is not an isolated artifact. Heretic automates abliteration while optimizing for fewer refusals and lower KL divergence, OBLITERATUS packages refusal removal and ablation studies into an open-source toolkit, and Orion-zhen's Abliteration implements refusal-direction removal plus norm-preserving variants [21]. Jang, who describes himself as an ML research and systems engineer in Irvine, California, has also published MLX Studio and the JANGQ variable-bit quantization project [7]. Dealign.ai reports more than 200 controlled experiments, over 40 findings and work across nine models from 0.8 billion to 397 billion parameters, all self-reported [15], and discloses no funding, investors, pricing, revenue or customer count on its public pages [8].

Two things worth tracking: whether anyone outside the project reproduces the compliance and MMLU numbers [20], and Jang's stated research line on how safety behavior is distributed across layers and how quantization weakens, preserves or restores it [22]. If quantization itself moves refusal behavior, then the artifact your gateway serves and the artifact you evaluated are not necessarily the same model.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories