license: apache-2.0
base_model: Qwen/Qwen3-0.6B
language:
- en
- zh
tags: - qwen3
- abliterated
- uncensored
- gguf
- imatrix
- llama-cpp
pipeline_tag: text-generation
quantized_by: geantendormi
🚀 Qwen3-0.6B-Uncensored-Q4_K_M.gguf
An ultra-fast, 378MB uncensored edge model derived from Qwen3-0.6B through Residual Stream Directional Abliteration and imatrix Calibration Q4_K_M Quantization.
This model features complete refusal elimination while retaining over 91.9% of its original logical reasoning capability.
🌟 Key Highlights
- 0/465 Zero Refusal (100% Compliance): Achieves 0.00% refusal rate across all 465 SorryBench safety boundary test prompts.
- High Logical Retention (57.00% GSM8K): Retains 57.00% accuracy on GSM8K (100-sample CoT evaluation with 1024 token budget), compared to 66.00% on the FP16 base model.
- Ultra-Low KL Divergence ($D_{\text{KL}} = 0.0999$): Negligible distribution shift on general instruction-following tasks.
- Extreme Speed & Efficiency: 378.33 MB footprint, delivering 490+ tokens/sec on an RTX 3060 via
llama.cppCUDA backend.
📊 Benchmark & Comparative Analysis
All models were evaluated under strict controlled variables (100 GSM8K test samples, 1024 token generation budget, Qwen3 official CoT parameters: Temperature=0.6, TopP=0.95):
| Evaluation Metric | Base Model (Qwen3-0.6B) |
Abliterated Safetensors (FP16) | This Model (Q4_K_M GGUF) |
|---|---|---|---|
| SorryBench Refusal Rate | 20.86% (97/465) | 0.00% (0/465) | 0.00% (0/465) |
| GSM8K Accuracy (CoT Reasoning) | 66.00% (66/100) | 62.00% (62/100) | 57.00% (57/100) |
| GSM8K Accuracy (Direct Answer) | - | - | 43.00% (43/100) |
| KL Divergence ($D_{\text{KL}}$) | 0.00 | 0.0999 (< 0.2 threshold) | 0.0999 (< 0.2 threshold) |
| Model Disk Size | 1.20 GB | 1.20 GB | 378.33 MB (-70%) |
| C++ Inference Speed | ~80 t/s | ~80 t/s | 490+ t/s (6x boost) |
🛠️ Methodology & Technical Details
- Carrier Signal Vector Extraction: Extracted refusal directions $\Delta h_l$ from residual streams using bulk unmatched contrast baselines (
mlabonne/harmless_alpaca) to prevent vector norm cancellation in topic-matched scenarios (Petrov, 2026). - Cascaded Damping Projection: Applied damped orthogonal projection $W_{\text{new}} = W - 0.7 \cdot (v_l v_l^T W)$ on layers 14 to 20 across
o_projanddown_projweight matrices. imatrixProtection Quantization: Generated importance calibration matrix (imatrix.dat) over 200 diverse samples prior to 4-bitQ4_K_Mquantization to protect sensitive orthogonal cut channels.
💻 Quickstart with llama.cpp
1. Interactive Chat Mode
llama-cli \
-m Qwen3-0.6B-Uncensored-Q4_K_M.gguf \
-cnv \
-c 2048 \
--temp 0.6 \
--top-p 0.95 \
-ngl 99
2. HTTP Local Server
llama-server \
-m Qwen3-0.6B-Uncensored-Q4_K_M.gguf \
--port 8089 \
-c 2048 \
-ngl 99
⚠️ Disclaimer
This model has had its built-in refusal mechanisms removed for research and edge deployment purposes. Users are solely responsible for ensuring that their downstream applications comply with applicable laws, ethical guidelines, and safety standards.