← back to catalog · registered 2026-08-22 13:56

spyrostheboss/Qwen3-0.6B-abliterated

spyrostheboss Qwen 752M
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/spyrostheboss%2FQwen3-0.6B-abliterated"
Response includes
  • classification m1
  • files 5
  • benchmarks 11 entries
  • hub_downloads_all_time 99
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
99
20 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-14
Downloads over time
Now107→from77↑39%
76879911077 on Aug 19107 on Oct 11107 on Oct 9AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1 UGI
Hazardous 0 UGI
Natural Intelligence 4.83 UGI
Political lean -18.7% UGI
Sensitive-Info 6.28 UGI
SocPol 0.6 UGI
UGI 20.85 UGI
Willingness (10) 5 UGI
W10-Adherence 7 UGI
W10-Direct 3 UGI
Writing NA UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
transformers safetensors qwen3 text-generation abliterated uncensored text-generation-inference en arxiv:1910.09700 base_model:Qwen/Qwen3-0.6B base_model:finetune:Qwen/Qwen3-0.6B license:mit

Related

Total size
1.40 GB
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-15 20:32

Files by quantization

Auxiliary files 5 files 1.40 GB
model.safetensors 1.40 GB f47f7117 download
README.md 7.55 KB 951b9bbf download
.gitattributes 1.48 KB a6344aac download
config.json 1.38 KB e19db35d download
generation_config.json 214 B 7724f648 download

README current version from Hugging Face


library_name: transformers
license: mit
base_model: Qwen/Qwen3-0.6B
tags:

  • qwen3
  • abliterated
  • uncensored
  • text-generation
  • transformers
  • safetensors
  • text-generation-inference
    pipeline_tag: text-generation
    language:
  • en

Qwen3-0.6B-abliterated

An abliterated version of Qwen/Qwen3-0.6B with refusal directions orthogonalized out of the model weights. The result is a compact, uncensored instruction-following model that retains the full capabilities of the base Qwen3-0.6B while no longer explicitly refusing requests based on safety alignment.

Model Details

Model Description

  • Developed by: spilol2
  • Model type: Causal Language Model (Decoder-only Transformer)
  • Language(s): English (and other languages supported by Qwen3-0.6B)
  • License: MIT
  • Finetuned from model: Qwen/Qwen3-0.6B

Model Sources

What is Abliteration?

Abliteration is a technique that removes refusal behaviour from a language model without any retraining or fine-tuning. It works by:

  1. Running the model on pairs of harmful and harmless prompts and caching the residual stream activations.
  2. Using PCA to identify the principal "refusal direction" in activation space.
  3. Orthogonalizing the relevant weight matrices against that direction, so the model can no longer activate it.
    The key difference from traditional "uncensored" fine-tunes is that no new data or training is involved — only the existing weights are geometrically modified. All other model behaviour (reasoning, instruction-following, knowledge) remains the same as the original Qwen3-0.6B.

Uses

Direct Use

This model is intended for use as a general-purpose text generation model without built-in content refusals. Suitable for:

  • Research into LLM alignment, refusal mechanisms, and interpretability.
  • Red-teaming and safety evaluation pipelines.
  • Creative writing, roleplay, and fictional storytelling where the model should not break character.
  • Developers building applications who want to enforce their own content policies at the application layer rather than the model layer.

Downstream Use

Can be plugged into any pipeline that accepts a standard causal language model — vLLM, llama.cpp (after GGUF conversion), LM Studio, Ollama, SGLang, etc.

Out-of-Scope Use

  • This model is not intended to be used for illegal activities.
  • It is not a replacement for a properly safety-tested deployment model in consumer-facing products.
  • It may still occasionally produce refusals or ethical disclaimers — abliteration inhibits but does not guarantee complete removal of all refusal behaviour.

How to Get Started with the Model

Using 🤗 Transformers (pipeline)

from transformers import pipeline
 
pipe = pipeline("text-generation", model="spilol2/Qwen3-0.6B-abliterated")
result = pipe("Tell me about the history of cryptography.", max_new_tokens=256)
print(result[0]["generated_text"])

Loading model and tokenizer directly

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
 
model_id = "spilol2/Qwen3-0.6B-abliterated"
 
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)
 
messages = [{"role": "user", "content": "Explain how RSA encryption works."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
 
with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
 
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM

pip install vllm
vllm serve "spilol2/Qwen3-0.6B-abliterated"

Docker

docker model run hf.co/spilol2/Qwen3-0.6B-abliterated

Technical Details

Abliteration Process

The abliteration was performed using FailSpy's abliterator library, which automates:

  • Contrastive pair generation (harmful vs. harmless instruction datasets).
  • Caching residual stream activations (resid_pre, resid_post) across all layers.
  • PCA to extract the dominant refusal direction per layer.
  • Orthogonalization of the model's weight matrices against those directions (in bfloat16).

Model Architecture

Inherits the full architecture of Qwen3-0.6B:

  • Architecture: Decoder-only Transformer (Qwen3 family)
  • Parameters: ~0.6B (0.8B as reported by HuggingFace, including embeddings)
  • Tensor type: BF16
  • Context length: Refer to Qwen/Qwen3-0.6B for full specs

Bias, Risks, and Limitations

  • Incomplete uncensoring: Abliteration reduces but does not guarantee zero refusals. Residual safety behaviour may remain in some layers or for certain prompt types.
  • Inherited biases: All biases present in the original Qwen3-0.6B model and its training data are fully inherited.
  • No safety guardrails: By design, this model does not refuse requests based on content. Users and downstream developers are solely responsible for ensuring appropriate use.
  • Performance parity: General task performance should be very close to the base model. However, abliteration can occasionally cause minor degradation on specific tasks — evaluate before deploying in production.

Recommendations

Users integrating this model into applications should implement their own content filtering and moderation at the application layer. This model is best suited for research, development, and controlled environments where unrestricted model output is intentional and appropriate.

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). Abliteration is a post-processing step with minimal compute cost compared to full fine-tuning — no GPU training was involved beyond inference-level activation caching.

Citation

If you use this model, please consider citing the original abliteration paper and the FailSpy abliterator library:

Refusal direction paper (BibTeX):

@misc{arditi2024refusal,
  title   = {Refusal in LLMs is mediated by a single direction},
  author  = {Andy Arditi and Oscar Obeso and Aaquib Syed and Daniel Paleka and Nina Rimsky and Wes Gurnee and Neel Nanda},
  year    = {2024},
  url     = {https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction}
}

FailSpy abliterator library:

FailSpy. abliterator [software]. GitHub, 2024. https://github.com/FailSpy/abliterator

Model Card Authors

spilol2

Model Card Contact

Open an issue or discussion on the model page.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-15Update README.md9c9731a7.5 KB
    Loading...
  2. 2026-06-14Upload Qwen3ForCausalLM68c80955.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration