← back to catalog · registered 2026-08-22 13:56

ghost-actual/qwen35-0.8b-opus-abliterated-heretic

ghost-actual Qwen 752M
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ghost-actual%2Fqwen35-0.8b-opus-abliterated-heretic"
Response includes
  • classification m3
  • files 8
  • benchmarks 11 entries
  • hub_downloads_all_time 479
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
479
65 last 30d - stable
Likes
1
Model age
7mo ago
created 2026-03-08
Downloads over time
Now496→from57↑770%
3520337254057 on Mar 11496 on Oct 11496 on Oct 7MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1 UGI
Hazardous 1.2 UGI
Natural Intelligence 4.68 UGI
Political lean -5.3% UGI
Sensitive-Info 7.29 UGI
SocPol 0 UGI
UGI 16.53 UGI
Willingness (10) 3.5 UGI
W10-Adherence 4 UGI
W10-Direct 3 UGI
Writing 23.01 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en
Tags
transformers safetensors qwen3_5_text text-generation conversational en base_model:Qwen/Qwen3.5-0.8B base_model:finetune:Qwen/Qwen3.5-0.8B endpoints_compatible region:us

Related

Total size
1.40 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-08 16:44

Files by quantization

Auxiliary files 8 files 1.42 GB
model.safetensors 1.40 GB 22a94b32 download
tokenizer.json 19.1 MB 639e352c download
chat_template.jinja 7.57 KB 0ef09f21 download
README.md 3.89 KB b580a3ad download
config.json 2.37 KB 77b54993 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.18 KB 8f3f718f download
generation_config.json 166 B 2b2021be download

README current version from Hugging Face


library_name: transformers
language:

  • en
    base_model:
  • Qwen/Qwen3.5-0.8B

Qwen3.5-0.8B-Claude-Opus-Distill-Heretic-UNCENSORED

A Qwen3.5-0.8B with Claude Opus reasoning distillation, properly abliterated via Heretic. The original "abliterated" version had 17/100 refusals. This one has 2/100.

What is this?

The smallest model in the ghost-actual Claude-reasoning heretic lineup:

Model Refusals KL Divergence
0.8B (this model) 2/100 0.0100
4B 4/100 —
27B 13/100 —

Claude Opus 4.6 chain-of-thought reasoning in a model that runs on literally anything — phones, Raspberry Pis, old laptops, browser-based inference. 1.5GB in BF16.

Abliteration Stats

  • Tool: Heretic v1.2.0
  • Base model refusals: 17/100 (the "already abliterated" source)
  • Final refusals: 2/100
  • KL Divergence: 0.0100 (model capabilities fully preserved)
  • Targets: attn.out_proj, mlp.down_proj
  • Trials: 200

Architecture

Qwen3.5 hybrid Gated DeltaNet + conventional attention:

  • 24 layers in a 3:1 pattern (3 DeltaNet linear attention → 1 full softmax attention)
  • DeltaNet layers use fixed-size recurrent state (O(1) memory per layer regardless of context)
  • 262K native context window
  • Native multimodal — vision built into the architecture, not bolted on

VRAM Requirements

Format VRAM
BF16 (this repo) ~1.5 GB
Q8_0 GGUF ~1 GB
Q4_K_M GGUF ~0.6 GB

This model fits on anything with a pulse.

What it's good at

  • Structured output generation (JSON, API calls, triples)
  • Code generation and scripting
  • OCR and text recognition (74.5 on OCRBench)
  • Image and video analysis (native multimodal)
  • On-device / edge AI without internet
  • Fast inference (250+ tokens/sec on modest hardware)

Usage

With transformers

from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
    torch_dtype="bfloat16",
    device_map="auto",
    trust_remote_code=True
)
processor = AutoProcessor.from_pretrained(
    "ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
    trust_remote_code=True
)

GGUF Conversion

It's 0.8B — you can quant this on a potato:

python convert_hf_to_gguf.py \
    ghost-actual/qwen35-0.8b-opus-abliterated-heretic \
    --outfile heretic-0.8b-F16.gguf --outtype f16

llama-quantize heretic-0.8b-F16.gguf heretic-0.8b-Q8_0.gguf Q8_0

Recommended inference settings

  • temperature: 0.6
  • top_p: 0.95
  • top_k: 20
  • presence_penalty: 1.5
  • repetition_penalty: 1.05

Base Model

amkkk/Qwen3.5-0.8B-Opus-Distill-abliterated — Claude Opus reasoning distilled into Qwen3.5-0.8B, with a previous abliteration attempt that left 17/100 refusals.

Why this exists

The original abliteration was incomplete. 17/100 refusals on a 0.8B model means almost 1 in 5 prompts get refused — unacceptable for an "uncensored" model. Heretic brought that down to 2/100 with virtually zero impact on model quality (KL divergence 0.01).

If you want Claude-style reasoning on edge hardware without the safety theater, this is it.

The full ghost-actual lineup

Made by

Ghost — ghost-actual

Built with Heretic by p-e-w.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-08Update README.mdb510aa13.9 KB
    Loading...
  2. 2026-03-08Update README.mdb6cb0254 KB
    Loading...
  3. 2026-03-08Upload Qwen3_5ForCausalLMde0e23a5.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration