← back to catalog · registered 2026-08-22 13:56

insraq/Qwen3.5-2B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated

insraq Qwen 2.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/insraq%2FQwen3.5-2B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated"
Response includes
  • classification m3
  • files 8
  • benchmarks 11 entries
  • hub_downloads_all_time 853
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
853
261 last 30d - stable
Likes
0
Descendants
2
in 2 direct forks
Model age
7w ago
created 2026-08-17
Downloads over time
Now902→from255↑254%
223471719967255 on Aug 19902 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.3 UGI
Hazardous 0.6 UGI
Natural Intelligence 9.03 UGI
Political lean -14.3% UGI
Sensitive-Info 2.6 UGI
SocPol 0 UGI
UGI 5.9 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 19.27 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text empero-ai qwen3.5 qwen3.8 distillation reasoning function-calling sft edge

Related

Total size
4.12 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 10:44

Files by quantization

Auxiliary files 8 files 4.14 GB
model.safetensors 4.12 GB 8033414e download
tokenizer.json 19.1 MB 6f32ce20 download
chat_template.jinja 7.72 KB cd794859 download
README.md 7.60 KB 16a3022b download
config.json 2.72 KB 60dbf412 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB ab901d8d download
generation_config.json 174 B c8d7e005 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-2B
language:

  • en
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • empero-ai
  • qwen3.5
  • qwen3.8
  • distillation
  • reasoning
  • function-calling
  • sft
  • edge
  • heretic
  • uncensored
  • decensored
  • abliterated
  • reproducible

This is a decensored version of empero-ai/Qwen3.8-2B, made using Heretic v1.4.0

[!TIP]
This model is reproducible!

See the README in the reproduce directory for more information.

Abliteration parameters

Parameter Value
direction_index 9.68
attn.o_proj.max_weight 0.97
attn.o_proj.max_weight_position 14.00
attn.o_proj.min_weight 0.82
attn.o_proj.min_weight_distance 11.66
mlp.down_proj.max_weight 0.81
mlp.down_proj.max_weight_position 14.57
mlp.down_proj.min_weight 0.54
mlp.down_proj.min_weight_distance 10.03

Performance

Metric This model Original model (empero-ai/Qwen3.8-2B)
KL divergence 0.0109 0 (by definition)
Refusals 3/100 82/100

Qwen3.8-2B

Developed by Empero

[!Note]
This repository contains model weights and configuration files in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and other standard runtimes with Qwen3.5 architecture support.

Qwen3.8-2B is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture — the smallest member of the family, trained on the same curriculum as its larger siblings. The student saw ~30,000 curated teacher traces from our internal Qwen3.8 distillation datasets — dense chain-of-thought spanning mathematics, general reasoning, and instruction following, quality-filtered before training.

The objective: the same teacher, the same reasoning curriculum, in a model small enough for the edge. What changes across Qwen3.8-9B, Qwen3.8-4B and this model is the student's capacity — not the quality or the character of what it was taught.

Highlights

  • Distilled chain-of-thought — every answer opens with a <think> block learned directly from Qwen3.8 2.4T A95B traces rather than synthetic self-generated reasoning.
  • Same curriculum as the larger siblings — the same teacher and the same quality-filtered trace mix used for the 9B and 4B distills; only the student's capacity differs.
  • 2B weight class — bf16 in ~4 GB; quantized builds run on phones, single-board computers, and CPU-only machines.
  • Native function calling per Qwen3.5's specification — no wrapper or tool-specific fine-tune required.
  • 262,144-token native context, inherited from the Qwen3.5 base.
  • Full fine-tune — every parameter updated; not an adapter.

Model Overview

  • Type: Causal Language Model (text path of a vision-language base)
  • Base: Qwen/Qwen3.5-2B
  • Number of Parameters: 2B
  • Training: SFT (off-policy distillation) on ~30,000 teacher traces
  • Teacher: Qwen3.8 2.4T A95B (internal distillation datasets)
  • Context Length: 262,144 natively

Benchmark Results

Measured with lm-evaluation-harness, HF backend, identical settings for base and student. Both models are reasoning models and are evaluated with the CoT protocols (gsm8k_cot, mmlu_flan_cot_zeroshot); MMLU covers all 57 subjects (~1,700 questions). Flexible-extract is the primary metric; strict-match requires exact answer formatting.

Task Metric Qwen3.5-2B (base) Qwen3.8-2B Δ
gsm8k_cot exact_match (flexible) 0.330 0.640 +0.310
gsm8k_cot exact_match (strict) 0.545 0.640 +0.095
mmlu (CoT, 57 subjects) acc (flexible-extract) 0.283 0.548 +0.265
mmlu (CoT, 57 subjects) acc (strict-match) 0.004 0.225 +0.221

Sampling for generation: temperature=0.6, top_p=0.95, top_k=20 (Qwen3.5 recommended settings).

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "empero-ai/Qwen3.8-2B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "A snail is at the bottom of a 10-meter well. Each day it climbs 3 meters, each night it slips back 2. How many days until it escapes?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)

out = model.generate(inputs, max_new_tokens=16384,
                     temperature=0.6, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

A recent transformers release with Qwen3.5 support is required, along with the Gated DeltaNet kernels (flash-linear-attention and a CUDA-matched causal_conv1d build) — without them the linear-attention layers fall back to slow, memory-hungry PyTorch ops.

Best Practices

  • Sampling: temperature=0.6, top_p=0.95, top_k=20. Greedy decoding on long generations is a known repetition-loop failure mode for reasoning models in this class.
  • Output length: allow generous max_new_tokens (16,384 recommended); every answer opens with a <think> block. Parse and strip the <think>...</think> span for end users.
  • Weight class, not curriculum: the reasoning form transfers from the same teacher traces the larger siblings learned from; what 2B parameters bound is capacity — factual recall and very hard multi-step problems. For harder workloads, step up to Qwen3.8-4B or Qwen3.8-9B.
  • Scope: the trace mix emphasizes mathematics, reasoning, and instruction following; the 9B additionally trains on code, so use Qwen3.8-9B for code-heavy workloads. The fine-tune is text-only; vision behavior is inherited from the base and was not evaluated here.

Stay in the loop

Sign up for the Empero newsletter at empero.org for releases, evals, and research notes.

Support / Donate

If this model helped you, consider supporting the project:

  • BTC: bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7v
  • LTC: ltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7x

Provenance & licensing

Weights are released under Apache-2.0, inherited from the Qwen3.5-2B base. Shared for research and experimentation, as-is.

Acknowledgements

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Upload README.md with huggingface_hub8b4cdfc7.5 KB
    Loading...
  2. 2026-08-17Upload Qwen3_5ForConditionalGeneration3fcf49c5.1 KB
    Loading...

Discussions 1 thread

  1. 2026-09-16Garbage!open3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration