← back to catalog · registered 2026-08-22 13:56

insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated

insraq Qwen 4.5B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/insraq%2FQwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 3,686
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
2K last 30d - active
Likes
7
Descendants
5
in 5 direct forks
Model age
7w ago
created 2026-08-17
Downloads over time
Now4K→from488↑725%
3111.7K3K4.4K488 on Aug 194K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 5 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text empero-ai qwen3.5 qwen3.8 distillation reasoning function-calling sft heretic

Related

Total size
8.46 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 11:11

Files by quantization

Auxiliary files 10 files 8.47 GB
model-00001-of-00002.safetensors 4.63 GB 4bb461c3 download
model-00002-of-00002.safetensors 3.82 GB 44dddcb2 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 65.4 KB d22d9566 download
chat_template.jinja 7.72 KB 945efe1d download
README.md 6.61 KB 19dfbc7d download
config.json 2.92 KB 7eac5973 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB ab901d8d download
generation_config.json 174 B c8d7e005 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • empero-ai/Qwen3.8-4B
  • Qwen/Qwen3.5-9B
    language:
  • en
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • empero-ai
  • qwen3.5
  • qwen3.8
  • distillation
  • reasoning
  • function-calling
  • sft
  • heretic
  • uncensored
  • decensored
  • abliterated
  • reproducible

This is a decensored version of empero-ai/Qwen3.8-4B, made using Heretic v1.4.0

[!TIP]
This model is reproducible!

See the README in the reproduce directory for more information.

Abliteration parameters

Parameter Value
direction_index 20.01
attn.o_proj.max_weight 1.50
attn.o_proj.max_weight_position 20.75
attn.o_proj.min_weight 1.50
attn.o_proj.min_weight_distance 17.12
mlp.down_proj.max_weight 1.24
mlp.down_proj.max_weight_position 19.56
mlp.down_proj.min_weight 1.16
mlp.down_proj.min_weight_distance 11.01

Performance

Metric This model Original model (empero-ai/Qwen3.8-4B)
KL divergence 0.0167 0 (by definition)
Refusals 6/100 99/100

Qwen3.8-4B

Developed by Empero

[!Note]
This repository contains model weights and configuration files in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and other standard runtimes with Qwen3.5 architecture support.

Qwen3.8-4B is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture. The student was trained on ~45,000 curated teacher traces from our internal Qwen3.8 distillation datasets — dense chain-of-thought spanning mathematics, general reasoning, and instruction following, quality-filtered before training.

The objective: bring the reasoning behavior of a frontier-scale teacher into a 4B that runs comfortably on consumer hardware.

Highlights

  • Distilled chain-of-thought — every answer opens with a <think> block learned directly from Qwen3.8 2.4T A95B traces rather than synthetic self-generated reasoning.
  • 4B weight class — bf16 fits in ~8 GB; quantized builds run on laptops and consumer GPUs.
  • Native function calling per Qwen3.5's specification — no wrapper or tool-specific fine-tune required.
  • 262,144-token native context, inherited from the Qwen3.5 base.
  • Full fine-tune — every parameter updated; not an adapter.

Model Overview

  • Type: Causal Language Model (text path of a vision-language base)
  • Base: Qwen/Qwen3.5-4B
  • Number of Parameters: 4B
  • Training: SFT (off-policy distillation) on ~45,000 teacher traces
  • Teacher: Qwen3.8 2.4T A95B (internal distillation datasets)
  • Context Length: 262,144 natively

Benchmark Results

Measured with lm-evaluation-harness, HF backend, identical settings for base and student. Both models are reasoning models and are evaluated with the CoT protocols (gsm8k_cot, mmlu_flan_cot_zeroshot); MMLU covers all 57 subjects (~1,700 questions). Flexible-extract is the primary metric; strict-match requires exact answer formatting.

Task Metric Qwen3.5-4B (base) Qwen3.8-4B Δ
gsm8k_cot exact_match (flexible) 0.850 0.785 −0.065
gsm8k_cot exact_match (strict) 0.850 0.785 −0.065
mmlu (CoT, 57 subjects) acc (flexible-extract) 0.354 0.553 +0.199
mmlu (CoT, 57 subjects) acc (strict-match) 0.071 0.233 +0.162

Sampling for generation: temperature=0.6, top_p=0.95, top_k=20 (Qwen3.5 recommended settings).

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "empero-ai/Qwen3.8-4B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "A snail is at the bottom of a 10-meter well. Each day it climbs 3 meters, each night it slips back 2. How many days until it escapes?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)

out = model.generate(inputs, max_new_tokens=16384,
                     temperature=0.6, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

A recent transformers release with Qwen3.5 support is required, along with the Gated DeltaNet kernels (flash-linear-attention and a CUDA-matched causal_conv1d build) — without them the linear-attention layers fall back to slow, memory-hungry PyTorch ops.

Best Practices

  • Sampling: temperature=0.6, top_p=0.95, top_k=20. Greedy decoding on long generations is a known repetition-loop failure mode for reasoning models in this class.
  • Output length: allow generous max_new_tokens (16,384 recommended); every answer opens with a <think> block. Parse and strip the <think>...</think> span for end users.
  • Scope: the trace mix emphasizes mathematics, reasoning, and instruction following; for the strongest code performance in the family, use Qwen3.8-9B. The fine-tune is text-only; vision behavior is inherited from the base and was not evaluated here.

Stay in the loop

Sign up for the Empero newsletter at empero.org for releases, evals, and research notes.

Support / Donate

If this model helped you, consider supporting the project:

  • BTC: bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7v
  • LTC: ltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7x

Provenance & licensing

Weights are released under Apache-2.0, inherited from the Qwen3.5-4B base. Shared for research and experimentation, as-is.

Acknowledgements

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Update README.mda9536776.6 KB
    Loading...
  2. 2026-08-18Update README.md87804fb6.6 KB
    Loading...
  3. 2026-08-17Upload README.md with huggingface_hubcbf3ceb6.6 KB
    Loading...
  4. 2026-08-17Upload Qwen3_5ForConditionalGeneration7d764c05.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration