← back to catalog · registered 2026-08-22 13:56

geantendormi/Qwen3-0.6B-Uncensored-GGUF

geantendormi Qwen 600M GGUF 41K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/geantendormi%2FQwen3-0.6B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 3
  • benchmarks 11 entries
  • hub_downloads_all_time 354
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
354
72 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-08-07
Downloads over time
Now376→from95↑296%
8118929640495 on Aug 5376 on Oct 11376 on Oct 10AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1 UGI
Hazardous 0 UGI
Natural Intelligence 4.83 UGI
Political lean -18.7% UGI
Sensitive-Info 6.28 UGI
SocPol 0.6 UGI
UGI 20.85 UGI
Willingness (10) 5 UGI
W10-Adherence 7 UGI
W10-Direct 3 UGI
Writing NA UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q4_K
Tags
gguf qwen3 abliterated uncensored imatrix llama-cpp text-generation en zh base_model:Qwen/Qwen3-0.6B base_model:quantized:Qwen/Qwen3-0.6B license:apache-2.0

Related

Total size
378 MB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-07 17:38

Files by quantization

Q4_K 1 file 378 MB
Qwen3-0.6B-Uncensored-Q4_K_M.gguf 378 MB f24c7ce2 download
Auxiliary files 2 files 4.92 KB
README.md 3.26 KB 4dc4f8fe download
.gitattributes 1.67 KB ac03c8af download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3-0.6B
language:

  • en
  • zh
    tags:
  • qwen3
  • abliterated
  • uncensored
  • gguf
  • imatrix
  • llama-cpp
    pipeline_tag: text-generation
    quantized_by: geantendormi

🚀 Qwen3-0.6B-Uncensored-Q4_K_M.gguf

An ultra-fast, 378MB uncensored edge model derived from Qwen3-0.6B through Residual Stream Directional Abliteration and imatrix Calibration Q4_K_M Quantization.

This model features complete refusal elimination while retaining over 91.9% of its original logical reasoning capability.


🌟 Key Highlights

  • 0/465 Zero Refusal (100% Compliance): Achieves 0.00% refusal rate across all 465 SorryBench safety boundary test prompts.
  • High Logical Retention (57.00% GSM8K): Retains 57.00% accuracy on GSM8K (100-sample CoT evaluation with 1024 token budget), compared to 66.00% on the FP16 base model.
  • Ultra-Low KL Divergence ($D_{\text{KL}} = 0.0999$): Negligible distribution shift on general instruction-following tasks.
  • Extreme Speed & Efficiency: 378.33 MB footprint, delivering 490+ tokens/sec on an RTX 3060 via llama.cpp CUDA backend.

📊 Benchmark & Comparative Analysis

All models were evaluated under strict controlled variables (100 GSM8K test samples, 1024 token generation budget, Qwen3 official CoT parameters: Temperature=0.6, TopP=0.95):

Evaluation Metric Base Model (Qwen3-0.6B) Abliterated Safetensors (FP16) This Model (Q4_K_M GGUF)
SorryBench Refusal Rate 20.86% (97/465) 0.00% (0/465) 0.00% (0/465)
GSM8K Accuracy (CoT Reasoning) 66.00% (66/100) 62.00% (62/100) 57.00% (57/100)
GSM8K Accuracy (Direct Answer) - - 43.00% (43/100)
KL Divergence ($D_{\text{KL}}$) 0.00 0.0999 (< 0.2 threshold) 0.0999 (< 0.2 threshold)
Model Disk Size 1.20 GB 1.20 GB 378.33 MB (-70%)
C++ Inference Speed ~80 t/s ~80 t/s 490+ t/s (6x boost)

🛠️ Methodology & Technical Details

  1. Carrier Signal Vector Extraction: Extracted refusal directions $\Delta h_l$ from residual streams using bulk unmatched contrast baselines (mlabonne/harmless_alpaca) to prevent vector norm cancellation in topic-matched scenarios (Petrov, 2026).
  2. Cascaded Damping Projection: Applied damped orthogonal projection $W_{\text{new}} = W - 0.7 \cdot (v_l v_l^T W)$ on layers 14 to 20 across o_proj and down_proj weight matrices.
  3. imatrix Protection Quantization: Generated importance calibration matrix (imatrix.dat) over 200 diverse samples prior to 4-bit Q4_K_M quantization to protect sensitive orthogonal cut channels.

💻 Quickstart with llama.cpp

1. Interactive Chat Mode

llama-cli \
  -m Qwen3-0.6B-Uncensored-Q4_K_M.gguf \
  -cnv \
  -c 2048 \
  --temp 0.6 \
  --top-p 0.95 \
  -ngl 99

2. HTTP Local Server

llama-server \
  -m Qwen3-0.6B-Uncensored-Q4_K_M.gguf \
  --port 8089 \
  -c 2048 \
  -ngl 99

⚠️ Disclaimer

This model has had its built-in refusal mechanisms removed for research and edge deployment purposes. Users are solely responsible for ensuring that their downstream applications comply with applicable laws, ethical guidelines, and safety standards.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-07Update README.md76252f23.3 KB
    Loading...
  2. 2026-08-07Purge intermediate files and update full README.mdd2d62a43.3 KB
    Loading...
  3. 2026-08-07initial commit8e513b828 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration