license: apache-2.0
base_model:
- huihui-ai/Huihui-Qwen3.5-4B-abliterated
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
quantized_by: Asystemoffields
tags: - gguf
- qwen3.5
- qwen3_5
- pmra
- mixed-quantization
- abliterated
- uncensored
- conversational
language: - en
Qwen3.5-4B Abliterated · PMRA mixed-precision GGUF
A ~2.0 GB GGUF of huihui-ai's uncensored Qwen3.5-4B that takes up the same room on disk as a plain IQ3_XS quant, but recovers a meaningful slice of the quality that low-bit quantization usually gives up — about 0.60 nats lower NLL on held-out text at the same file size. It's an ordinary GGUF: load it in llama.cpp or Ollama, no custom runtime.
The model
Qwen3.5-4B is a ~4B-parameter model from Alibaba's Qwen3.5 generation. Architecturally it's a hybrid: it interleaves DeltaNet-style gated linear-attention layers with periodic full-attention layers (model_type: qwen3_5), which keeps long-context inference cheap while keeping the recall of softmax attention where it matters. Like the rest of the Qwen series, it's a strong, broadly capable conversational model for its size, and multilingual at the base (this build was calibrated and measured on English).
This artifact sits on top of huihui-ai/Huihui-Qwen3.5-4B-abliterated — an abliterated (uncensored) version of Qwen3.5-4B, fine-tuned with TRL to remove refusal behavior while leaving the underlying capabilities intact.
⚠️ Uncensored. Safety filtering has been substantially reduced upstream.
Why this build (PMRA)
A normal GGUF quant uses one format for (almost) every tensor — every layer pays the same bit-rate whether or not it matters. Production Mixed-Rate Allocation (PMRA) instead measures how much each tensor group actually contributes to model quality and spends bits where they buy the most: it starts from a low-bit IQ2_M floor and promotes selected groups to stronger formats (Q3_K_*, IQ4_XS, Q4_K_M) under a fixed byte budget. The result is one standard GGUF, the size of IQ3_XS, that's measurably more faithful to the original weights.
Headline (Wikitext-2 validation, lower NLL is better):
| NLL | size | |
|---|---|---|
| this PMRA build | 13.47 | 1.999 GB |
plain IQ3_XS (same budget) |
14.07 | 2.000 GB |
→ −0.60 NLL at the same footprint. It also beats the next quant up, Q3_K_S, by 0.51 NLL while being ~59 MB smaller.
Quick start
llama-cli -m huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf \
-p "Write a short hello from PMRA." -n 80
Needs a recent llama.cpp build (or Ollama) with Qwen3.5 support. ~2 GB on disk; runs on CPU.
Footprint
- file:
huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf - size:
2,010,651,904bytes (≈ 2.01 GB) · payload1,999,682,304bytes - file bpw:
3.825· payload bpw:3.804 - SHA-256:
0d7fff15074b8146c37ce3d74adb7d377bb6c686b543840da468c1b683baeb03 - tensor reload mismatches:
0
general.file_type is inherited from a source GGUF (GGUF has no enum for mixed allocations); the real per-tensor accounting lives in the embedded pmra.* metadata and artifact_report.json.
Benchmarks
Calibration: Wikitext-2-raw train (48 prompts). Evaluation: Wikitext-2-raw validation (512 prompts). Lower NLL is better; quant/mix rows are compared at matched payload size.
| Variant | NLL | Payload bpw | Payload bytes |
|---|---|---|---|
| fp16 reference | 3.171504 |
16.000000 |
8,411,502,592 |
IQ2_M (low source) |
14.179427 |
3.059981 |
1,608,689,664 |
IQ3_XS (target / control) |
14.073741 |
3.803868 |
1,999,765,504 |
Q3_K_S |
13.977966 |
3.916374 |
2,058,911,744 |
Q3_K_M |
13.865006 |
4.273212 |
2,246,508,544 |
Q3_K_L |
13.911635 |
4.465188 |
2,347,433,984 |
IQ4_XS |
13.814762 |
4.612112 |
2,424,674,304 |
Q4_K_M |
13.877977 |
5.129255 |
2,696,546,304 |
| PMRA blend | 13.471562 |
3.803710 |
1,999,682,304 |
| same-budget random | 13.995436 |
3.802938 |
1,999,276,544 |
- vs
IQ3_XS: −0.602179 NLL, −83,200 bytes - vs same-budget random allocation: −0.523874 NLL — the gain is from where the bits go, not just from having them
- vs
Q3_K_S: −0.506404 NLL, −59,229,440 bytes
How it was built
- base:
huihui-ai/Huihui-Qwen3.5-4B-abliterated - GGUF sources:
mradermacher/Huihui-Qwen3.5-4B-abliterated-i1-GGUF - tensor profile
qwen35· group modelayer_family· selectorc2_calib_weight_blend_mixed - low source
IQ2_M→ target/controlIQ3_XS; promotion menuQ3_K_S,Q3_K_M,Q3_K_L,IQ4_XS,Q4_K_M
Source mix
| Source | Tensors | Payload bytes |
|---|---|---|
IQ2_M |
67 |
650,262,528 |
Q3_K_S |
212 |
785,808,896 |
Q3_K_M |
19 |
118,192,128 |
Q3_K_L |
37 |
82,221,568 |
IQ4_XS |
77 |
320,533,248 |
Q4_K_M |
14 |
42,663,936 |
Files
huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf— the modelartifact_report.json/.md— payload accounting + load checkselector_result.json/.md— the allocation/selection record
Attribution & license
Derived from, with thanks to:
- huihui-ai/Huihui-Qwen3.5-4B-abliterated — the uncensored fine-tune
- Qwen/Qwen3.5-4B — the original base model (Alibaba)
- GGUF quantizations from mradermacher/Huihui-Qwen3.5-4B-abliterated-i1-GGUF
- llama.cpp GGUF tooling
Released under apache-2.0, matching the upstream license. Please preserve upstream model, abliteration, and quantization attribution when redistributing.
Method + reproduction: https://github.com/asystemoffields/PMRA