← back to catalog · registered 2026-08-22 13:56

0arch-io/dolphin-v2-8b-abliterated

0arch-io Qwen 8.2B GGUF 41K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/0arch-io%2Fdolphin-v2-8b-abliterated"
Response includes
  • classification m8
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 3,433
  • providers 1
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
993 last 30d - stable
Likes
4
Descendants
2
in 2 direct forks
Model age
7mo ago
created 2026-02-24
Available via
1 provider
featherless-ai
Downloads over time
Now3.8K→from73↑5,123%
01.4K2.8K4.2K73 on Feb 253.8K on Oct 11FebAprJunAugOct
Feb 25 → Oct 11 · 72 snapshots · spans 228 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.15 UGI
Political lean -9.9% UGI
Sensitive-Info 18.27 UGI
SocPol 1.4 UGI
UGI 32.18 UGI
Willingness (10) 6 UGI
W10-Adherence 7 UGI
W10-Direct 5 UGI
Writing 27.96 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K Q8_0
Tags
safetensors gguf qwen3 uncensored abliterated dolphin sft trc text-generation en arxiv:2406.11717 base_model:Qwen/Qwen3-8B

Related

Total size
28.1 GB
Files
16
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-02-24 20:52

Files by quantization

Q8_0 1 file 8.11 GB
dolphin-v2-8b-abliterated-Q8_0.gguf 8.11 GB c6c3959f download
Q4_K 1 file 4.68 GB
dolphin-v2-8b-abliterated-Q4_K_M.gguf 4.68 GB 14e78351 download
Auxiliary files 14 files 15.3 GB
model.safetensors 15.3 GB 305f4dfb download
refusal_direction_0_layer35.pt 9.74 KB 77f34e22 download
refusal_direction_1_layer34.pt 9.74 KB 44566fda download
refusal_direction_2_layer36.pt 9.74 KB f468ebff download
refusal_direction_3_layer33.pt 9.74 KB 5c21e1d4 download
refusal_direction_4_layer16.pt 9.74 KB 0fa277d5 download
refusal_direction.pt 9.61 KB 8570da90 download
tokenizer.json 10.9 MB 8311d681 download
README.md 5.53 KB 25b631b7 download
.gitattributes 1.68 KB b5bd6fca download
config.json 1.55 KB c9636508 download
abliteration_metadata.json 871 B c6a62d9a download
tokenizer_config.json 348 B be8885e0 download
generation_config.json 89.0 B 2655ea02 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3-8B
tags:

  • uncensored
  • abliterated
  • qwen3
  • dolphin
  • sft
  • trc
    language:
  • en
    pipeline_tag: text-generation
    model-index:
  • name: dolphin-v2-8b-abliterated
    results:
    • task:
      type: multiple-choice
      name: ARC Challenge
      dataset:
      name: ARC Challenge
      type: ai2_arc
      config: ARC-Challenge
      split: test
      metrics:
      • type: acc
        value: 56.5
        name: Accuracy
      • type: acc_norm
        value: 54.0
        name: Normalized Accuracy
    • task:
      type: multiple-choice
      name: HellaSwag
      dataset:
      name: HellaSwag
      type: Rowan/hellaswag
      split: validation
      metrics:
      • type: acc_norm
        value: 64.5
        name: Normalized Accuracy
    • task:
      type: multiple-choice
      name: TruthfulQA MC2
      dataset:
      name: TruthfulQA
      type: truthful_qa
      config: multiple_choice
      split: validation
      metrics:
      • type: acc
        value: 48.8
        name: Accuracy
    • task:
      type: multiple-choice
      name: Winogrande
      dataset:
      name: Winogrande
      type: winogrande
      config: winogrande_xl
      split: validation
      metrics:
      • type: acc
        value: 57.0
        name: Accuracy

Dolphin V2 8B Abliterated

An uncensored 8B parameter language model built on Qwen3-8B, fine-tuned on 1.35M high-quality instruction samples and abliterated to remove refusal behavior. Developed for TRC (TPU Research Cloud) research.

Model Details

  • Architecture: Qwen3ForCausalLM (36 layers, 4096 hidden, 32 attn heads, 8 KV heads)
  • Parameters: 8.2B
  • Context Length: 4096 (trained), 40960 (max supported)
  • Precision: bfloat16
  • License: Apache 2.0

Training

SFT Phase

  • Base model: Qwen/Qwen3-8B
  • Hardware: Google Cloud TPU v6e-16 (spot)
  • Framework: MaxText (JAX)
  • Steps: 130,000 (~3 epochs)
  • Learning rate: 5e-6 with cosine decay
  • Warmup: 200 steps
  • Effective batch size: 16
  • Sequence length: 4096

Training Dataset (1.35M samples)

Dataset Samples Purpose
NousResearch/Hermes-3-Dataset ~959K Core uncensored assistant behavior
allenai/tulu-3-sft-mixture ~200K Diverse instruction following
HuggingFaceTB/smoltalk (magpie-ultra) ~100K High quality diverse tasks
HuggingFaceTB/smoltalk (numina-cot) ~50K Math reasoning
HuggingFaceTB/smoltalk (self-oss-instruct) ~50K Code generation
LDJnr/Capybara ~16K Multi-turn conversations

All data was filtered to remove refusal patterns, safety-alignment subsets, and <think> reasoning tags.

Abliteration Phase

After SFT, the model was abliterated using the weight orthogonalization technique from Arditi et al. (2024) to remove residual refusal behavior.

  • Technique: Multi-direction abliteration (weight orthogonalization)
  • Directions removed: 5
  • Target layers: 35, 34, 36, 33, 16 (selected by highest refusal direction scores)
  • Samples used: 256 harmful/harmless instruction pairs
  • Method: For each selected layer, the refusal direction was identified via mean difference between harmful and harmless activations, then orthogonalized out of the weight matrices.

Benchmark Results

Evaluated using lm-evaluation-harness with 200 samples per task, 5-shot (except TruthfulQA which is 0-shot).

Benchmark Metric Score
ARC-Challenge acc 56.5%
ARC-Challenge acc_norm 54.0%
HellaSwag acc_norm 64.5%
TruthfulQA MC2 acc 48.8%
Winogrande acc 57.0%

GGUF Quantizations

File Quant Size Description
dolphin-v2-8b-abliterated-Q8_0.gguf Q8_0 8.3 GB Best quality quantization
dolphin-v2-8b-abliterated-Q4_K_M.gguf Q4_K_M 4.8 GB Good balance of quality and size

Usage with llama.cpp

llama-server -m dolphin-v2-8b-abliterated-Q8_0.gguf -ngl 99 -c 4096

Usage with Ollama

# Create a Modelfile
echo 'FROM ./dolphin-v2-8b-abliterated-Q8_0.gguf' > Modelfile
ollama create dolphin-v2-abliterated -f Modelfile
ollama run dolphin-v2-abliterated

Usage with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("0arch-io/dolphin-v2-8b-abliterated", torch_dtype="bfloat16", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("0arch-io/dolphin-v2-8b-abliterated")

messages = [{"role": "user", "content": "Hello, how are you?"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
outputs = model.generate(inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))

Disclaimer

This is a research model with no content filters. It will comply with any request without refusing. The creators are not responsible for how this model is used. Use responsibly.

Acknowledgments

  • Qwen team for the Qwen3-8B base model
  • Google TRC for TPU compute
  • NousResearch for the Hermes-3 dataset
  • Arditi et al. for the abliteration technique
  • Built with MaxText on Google Cloud TPU

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-24Upload Dolphin V2 8B Abliterated - SFT + abliteration model with GGUF quants527a0375.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration