← back to catalog · registered 2026-08-22 13:56

Asystemoffields/Huihui-Qwen3.5-4B-Abliterated-PMRA-GGUF

Asystemoffields Qwen 4B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Asystemoffields%2FHuihui-Qwen3.5-4B-Abliterated-PMRA-GGUF"
Response includes
  • classification m8
  • files 8
  • hub_downloads_all_time 2,434
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
112 last 30d - cooling
Likes
0
Model age
4mo ago
created 2026-05-18
Downloads over time
Now2.4K→from757↑223%
6731.3K2K2.6K757 on May 202.4K on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
llama.cpp gguf qwen3.5 qwen3_5 pmra mixed-quantization abliterated uncensored conversational text-generation en base_model:huihui-ai/Huihui-Qwen3.5-4B-abliterated

Related

Total size
1.87 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 21:25

Files by quantization

Auxiliary files 8 files 1.87 GB
huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf 1.87 GB 0d7fff15 download
selector_result.json 304 KB 0c4e7f17 download
artifact_report.json 105 KB 84e7ceb4 download
README.md 5.82 KB 5aee7a0d download
QWEN35_ABLITERATED_PMRA.md 4.84 KB e0c09af3 download
selector_result.md 2.17 KB 11e66331 download
.gitattributes 1.61 KB 002173a6 download
artifact_report.md 1.08 KB 980971ee download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.5-4B-abliterated
    base_model_relation: quantized
    library_name: llama.cpp
    pipeline_tag: text-generation
    quantized_by: Asystemoffields
    tags:
  • gguf
  • qwen3.5
  • qwen3_5
  • pmra
  • mixed-quantization
  • abliterated
  • uncensored
  • conversational
    language:
  • en

Qwen3.5-4B Abliterated · PMRA mixed-precision GGUF

A ~2.0 GB GGUF of huihui-ai's uncensored Qwen3.5-4B that takes up the same room on disk as a plain IQ3_XS quant, but recovers a meaningful slice of the quality that low-bit quantization usually gives up — about 0.60 nats lower NLL on held-out text at the same file size. It's an ordinary GGUF: load it in llama.cpp or Ollama, no custom runtime.

The model

Qwen3.5-4B is a ~4B-parameter model from Alibaba's Qwen3.5 generation. Architecturally it's a hybrid: it interleaves DeltaNet-style gated linear-attention layers with periodic full-attention layers (model_type: qwen3_5), which keeps long-context inference cheap while keeping the recall of softmax attention where it matters. Like the rest of the Qwen series, it's a strong, broadly capable conversational model for its size, and multilingual at the base (this build was calibrated and measured on English).

This artifact sits on top of huihui-ai/Huihui-Qwen3.5-4B-abliterated — an abliterated (uncensored) version of Qwen3.5-4B, fine-tuned with TRL to remove refusal behavior while leaving the underlying capabilities intact.

⚠️ Uncensored. Safety filtering has been substantially reduced upstream.

Why this build (PMRA)

A normal GGUF quant uses one format for (almost) every tensor — every layer pays the same bit-rate whether or not it matters. Production Mixed-Rate Allocation (PMRA) instead measures how much each tensor group actually contributes to model quality and spends bits where they buy the most: it starts from a low-bit IQ2_M floor and promotes selected groups to stronger formats (Q3_K_*, IQ4_XS, Q4_K_M) under a fixed byte budget. The result is one standard GGUF, the size of IQ3_XS, that's measurably more faithful to the original weights.

Headline (Wikitext-2 validation, lower NLL is better):

NLL size
this PMRA build 13.47 1.999 GB
plain IQ3_XS (same budget) 14.07 2.000 GB

→ −0.60 NLL at the same footprint. It also beats the next quant up, Q3_K_S, by 0.51 NLL while being ~59 MB smaller.

Quick start

llama-cli -m huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf \
  -p "Write a short hello from PMRA." -n 80

Needs a recent llama.cpp build (or Ollama) with Qwen3.5 support. ~2 GB on disk; runs on CPU.

Footprint

  • file: huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf
  • size: 2,010,651,904 bytes (≈ 2.01 GB) · payload 1,999,682,304 bytes
  • file bpw: 3.825 · payload bpw: 3.804
  • SHA-256: 0d7fff15074b8146c37ce3d74adb7d377bb6c686b543840da468c1b683baeb03
  • tensor reload mismatches: 0

general.file_type is inherited from a source GGUF (GGUF has no enum for mixed allocations); the real per-tensor accounting lives in the embedded pmra.* metadata and artifact_report.json.

Benchmarks

Calibration: Wikitext-2-raw train (48 prompts). Evaluation: Wikitext-2-raw validation (512 prompts). Lower NLL is better; quant/mix rows are compared at matched payload size.

Variant NLL Payload bpw Payload bytes
fp16 reference 3.171504 16.000000 8,411,502,592
IQ2_M (low source) 14.179427 3.059981 1,608,689,664
IQ3_XS (target / control) 14.073741 3.803868 1,999,765,504
Q3_K_S 13.977966 3.916374 2,058,911,744
Q3_K_M 13.865006 4.273212 2,246,508,544
Q3_K_L 13.911635 4.465188 2,347,433,984
IQ4_XS 13.814762 4.612112 2,424,674,304
Q4_K_M 13.877977 5.129255 2,696,546,304
PMRA blend 13.471562 3.803710 1,999,682,304
same-budget random 13.995436 3.802938 1,999,276,544
  • vs IQ3_XS: −0.602179 NLL, −83,200 bytes
  • vs same-budget random allocation: −0.523874 NLL — the gain is from where the bits go, not just from having them
  • vs Q3_K_S: −0.506404 NLL, −59,229,440 bytes

How it was built

  • base: huihui-ai/Huihui-Qwen3.5-4B-abliterated
  • GGUF sources: mradermacher/Huihui-Qwen3.5-4B-abliterated-i1-GGUF
  • tensor profile qwen35 · group mode layer_family · selector c2_calib_weight_blend_mixed
  • low source IQ2_M → target/control IQ3_XS; promotion menu Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_M

Source mix

Source Tensors Payload bytes
IQ2_M 67 650,262,528
Q3_K_S 212 785,808,896
Q3_K_M 19 118,192,128
Q3_K_L 37 82,221,568
IQ4_XS 77 320,533,248
Q4_K_M 14 42,663,936

Files

  • huihui_qwen35_4b_abliterated_pmra_calib_weight_blend.gguf — the model
  • artifact_report.json / .md — payload accounting + load check
  • selector_result.json / .md — the allocation/selection record

Attribution & license

Derived from, with thanks to:

Released under apache-2.0, matching the upstream license. Please preserve upstream model, abliteration, and quantization attribution when redistributing.

Method + reproduction: https://github.com/asystemoffields/PMRA

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02Update README.md7da60b25.8 KB
    Loading...
  2. 2026-06-02Rewrite model card: lead with base model identity + value prop; keep full PMR...a3be74c6.1 KB
    Loading...
  3. 2026-05-23Update README.mdcbeeb743.1 KB
    Loading...
  4. 2026-05-22Sanitized public PMRA release5ea34df3.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration