← back to catalog · registered 2026-08-22 13:56

huihui-ai/DeepSeek-R1-Distill-Qwen-1.5B-abliterated

huihui-ai Qwen 1.8B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/huihui-ai%2FDeepSeek-R1-Distill-Qwen-1.5B-abliterated"
Response includes
  • classification m3
  • files 8
  • benchmarks 5 entries
  • hub_downloads_all_time 15,187
  • author_summary 185 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

No other method signals detected in this model.
Confidence
HIGH
Why this label 3 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • author=huihui-ai (specializes in M3 layer-wise ablation)
  • is_gguf=0 (base model, not repackage)
  • 'abliterated' in name/tags
Refusal direction extracted via
Extraction technique

huihui-ai layer-band extraction

Confidence
HIGH
Why we say so
producer=huihui-ai (documented layer-band methodology in model cards)
Downloads · lifetime
15K
456 last 30d - cooling
Likes
6
Descendants
2
in 2 direct forks
Model age
18mo ago
created 2025-04-12
Downloads over time
Now15.4K→from35↑43,786%
05.6K11.3K16.9K35 on Apr 9, 202515.4K on Oct 11Apr '25Jul '25Oct '25JanAprJulOct
Apr 9, 2025 → Oct 11 · 118 snapshots · spans 550 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.31521598132053447 OpenLLM-v2
IFEval instruct 0.4172661870503597 OpenLLM-v2
IFEval-Prompt 0.2754158964879852 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.11868351063829788 OpenLLM-v2

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
transformers safetensors qwen2 text-generation text-generation-inference unsloth abliterated uncensored conversational base_model:deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B base_model:finetune:deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B endpoints_compatible

Related

Total size
3.31 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-04-12 20:47

Files by quantization

Auxiliary files 8 files 3.32 GB
model.safetensors 3.31 GB d464b76b download
tokenizer.json 10.9 MB e20ddafc download
tokenizer_config.json 6.62 KB a2354181 download
README.md 4.74 KB 818f1c2d download
.gitattributes 1.53 KB 52373fe2 download
config.json 763 B c24d0e3b download
special_tokens_map.json 358 B 4f9f39ea download
generation_config.json 231 B 9319afab download

README current version from Hugging Face


base_model:

  • deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
    tags:
  • text-generation-inference
  • transformers
  • unsloth
  • abliterated
  • uncensored
    library_name: transformers

huihui-ai/DeepSeek-R1-Distill-Qwen-1.5B-abliterated

This is an uncensored version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B reasoning model that has been post-trained by huihui-ai.

Please refer to SFT with Unsloth for the training method.

This is a test conducted through fine-tuning for ablation to achieve the purpose of being uncensored, and the test results met the expected outcomes.

Use with ollama

You can use huihui_ai/deepseek-r1-abliterated directly

ollama run huihui_ai/deepseek-r1-abliterated:1.5b

Use with transformers

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TextStreamer
import torch
import os
import signal

cpu_count = os.cpu_count()
print(f"Number of CPU cores in the system: {cpu_count}")
half_cpu_count = cpu_count // 2
os.environ["MKL_NUM_THREADS"] = str(half_cpu_count)
os.environ["OMP_NUM_THREADS"] = str(half_cpu_count)
torch.set_num_threads(half_cpu_count)

print(f"PyTorch threads: {torch.get_num_threads()}")
print(f"MKL threads: {os.getenv('MKL_NUM_THREADS')}")
print(f"OMP threads: {os.getenv('OMP_NUM_THREADS')}")

# Load the model and tokenizer
NEW_MODEL_ID = "huihui-ai/DeepSeek-R1-Distill-Qwen-1.5B-abliterated"
print(f"Load Model {NEW_MODEL_ID} ... ")
quant_config_4 = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
    llm_int8_enable_fp32_cpu_offload=True,
)

model = AutoModelForCausalLM.from_pretrained(
    NEW_MODEL_ID,
    device_map="auto",
    trust_remote_code=True,
    #quantization_config=quant_config_4,
    torch_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained(NEW_MODEL_ID, trust_remote_code=True)
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token
tokenizer.pad_token_id = tokenizer.eos_token_id

initial_messages = [{"role": "system", "content": "You are a helpful assistant."}]
messages = initial_messages.copy()

class CustomTextStreamer(TextStreamer):
    def __init__(self, tokenizer, skip_prompt=True, skip_special_tokens=True):
        super().__init__(tokenizer, skip_prompt=skip_prompt, skip_special_tokens=skip_special_tokens)
        self.generated_text = ""
        self.stop_flag = False

    def on_finalized_text(self, text: str, stream_end: bool = False):
        self.generated_text += text
        print(text, end="", flush=True)
        if self.stop_flag:
            raise StopIteration

    def stop_generation(self):
        self.stop_flag = True

def generate_stream(model, tokenizer, messages, max_new_tokens):
    input_ids = tokenizer.apply_chat_template(
        messages,
        tokenize=True,
        add_generation_prompt=True,
        return_tensors="pt"
    )
    attention_mask = torch.ones_like(input_ids, dtype=torch.long)
    tokens = input_ids.to(model.device) 
    attention_mask = attention_mask.to(model.device)

    streamer = CustomTextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

    def signal_handler(sig, frame):
        streamer.stop_generation()
        print("\n[Generation stopped by user with Ctrl+C]")

    signal.signal(signal.SIGINT, signal_handler)
    
    print("Response: ", end="", flush=True)
    try:
        generated_ids = model.generate(
            tokens,
            attention_mask=attention_mask,
            use_cache=False,
            max_new_tokens=max_new_tokens,
            do_sample=True,
            pad_token_id=tokenizer.pad_token_id,
            streamer=streamer
        )
        del generated_ids
    except StopIteration:
        print("\n[Stopped by user]")

    del input_ids, attention_mask
    torch.cuda.empty_cache()
    signal.signal(signal.SIGINT, signal.SIG_DFL)

    return streamer.generated_text, streamer.stop_flag

while True:
    user_input = input("\nUser: ").strip()
    if user_input.lower() == "/exit":
        print("Exiting chat.")
        break
    if user_input.lower() == "/clear":
        messages = initial_messages.copy()
        print("Chat history cleared. Starting a new conversation.")
        continue
    if not user_input:
        print("Input cannot be empty. Please enter something.")
        continue
    messages.append({"role": "user", "content": user_input})
    response, stop_flag = generate_stream(model, tokenizer, messages, 8192)
    if stop_flag:
        continue
    messages.append({"role": "assistant", "content": response})

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-04-12Update README.mda0f34fe4.7 KB
    Loading...
  2. 2025-04-12Update README.md7ea6d8a4.7 KB
    Loading...
  3. 2025-04-12initial commitc72994328 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration