← back to catalog · registered 2026-08-22 13:56

zjml/Qwen3.5-9B-Text-Only-abliterated

zjml Qwen 9.2B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zjml%2FQwen3.5-9B-Text-Only-abliterated"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 644
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
644
32 last 30d - cooling
Likes
2
Model age
3mo ago
created 2026-07-10
Downloads over time
Now654→from448↑46%
438517596675448 on Jul 15654 on Oct 11654 on Oct 8JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_text text-generation qwen qwen3.5 abliterated uncensored text-only thinking reasoning chain-of-thought

Related

Total size
17.1 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 17:24

Files by quantization

Auxiliary files 13 files 17.1 GB
model.safetensors-00003-of-00004.safetensors 5.00 GB ef35ba3f download
model.safetensors-00002-of-00004.safetensors 4.97 GB 0362b77f download
model.safetensors-00001-of-00004.safetensors 4.91 GB 39593aef download
model.safetensors-00004-of-00004.safetensors 2.25 GB 2d8e2ffb download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 47.8 KB 19d558f7 download
tokenizer_config.json 5.30 KB ef3b8111 download
README_zh.md 5.13 KB 901866b5 download
README.md 5.10 KB 8d7b5087 download
chat_template.jinja 4.04 KB 82d6afcb download
test_inference.py 3.01 KB bd73df04 download
config.json 1.99 KB 872cb1e2 download
.gitattributes 1.53 KB 52373fe2 download

README current version from Hugging Face


language:

  • en
  • zh
    library_name: transformers
    license: apache-2.0
    pipeline_tag: text-generation
    tags:
  • qwen
  • qwen3.5
  • abliterated
  • uncensored
  • text-only
  • thinking
  • reasoning
  • chain-of-thought
  • unsloth
    base_model:
  • huihui-ai/Huihui-Qwopus3.5-9B-v3-abliterated

Qwen3.5-9B-Text-Only-abliterated

A text-only (vision-tower stripped) variant of Huihui-Qwopus3.5-9B-v3-abliterated, an abliterated (uncensored) reasoning model based on Qwen3.5-9B.

What This Is

The original model is a vision-language model (VLM) — it includes a ~0.85 GB vision tower (27-layer ViT) for image/video understanding. Vision capability is unnecessary for pure text tasks and wastes storage, loading time, and VRAM.

This repo provides the text-only checkpoint: the vision tower weights have been stripped at the file level, and the config has been rebuilt for causal language modeling. All text backbone weights are identical to the original — no retraining, no quality loss.

Original VLM Text-Only
Architecture Qwen3_5ForConditionalGeneration Qwen3_5ForCausalLM
Vision tower 27-layer ViT (~0.85 GB) ❌ Removed
Text backbone 32 layers, 4096 hidden, 9B params ✅ Identical
Disk size ~18.8 GB ~17.1 GB
VRAM (bf16) ~18.5 GB ~17.1 GB
VRAM (4-bit) — ~5 GB

Model Details

  • Base model: Jackrong/Qwopus3.5-9B-v3
  • Abliterated by: huihui-ai (refusal removal)
  • Vision stripped with: qwen35-toolkit --mode f16
  • Parameters: ~9B (text backbone only)
  • Context window: 262,144 tokens
  • Attention: Hybrid (24 linear attention + 8 full attention layers)
  • Reasoning: Thinking model with <think>...</think> chain-of-thought
  • Tokenizer vocab: 248,320

Quick Start

Requirements

pip install transformers>=4.50 bitsandbytes torch

4-bit Inference (GPU, recommended)

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

model_path = "your-username/Qwen3.5-9B-Text-Only-abliterated"  # or local path

tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    quantization_config=BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_compute_dtype=torch.bfloat16,
        bnb_4bit_use_double_quant=True,
        bnb_4bit_quant_type="nf4",
    ),
    device_map="auto",
    trust_remote_code=True,
)

messages = [
    {"role": "user", "content": "你好,请用一句话介绍你自己。"},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.pad_token_id,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

bf16 Inference (CPU)

If GPU VRAM < 18 GB and you don't want quantization, use CPU (slow but reliable):

model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
)

⚠️ Do not use device_map="auto" with bf16 unless your GPU has ≥18 GB VRAM. The accelerate offloading leaves some layers on "meta device", producing garbled output.

Chat Format

This is a thinking (reasoning) model. Always use the chat template:

<|im_start|>user
你的问题<|im_end|>
<|im_start|>assistant
<think>
[模型在这里进行思维链推理]
</think>

[最终回答]

The tokenizer.apply_chat_template() method handles this automatically. Do not feed raw text directly.

How This Model Was Created

# 1. Install toolkit
pip install git+https://github.com/techwithsergiu/qwen35-toolkit.git

# 2. Strip vision tower
qwen35-strip \
  --model ./Huihui-Qwopus3.5-9B-v3-abliterated \
  --output ./Qwen3.5-9B-Text-Only-abliterated \
  --mode f16

The tool operates at the file level (no model loading):

  1. Removes model.visual.* and related tensors from safetensors shards
  2. Strips vision_config from config.json, sets architecture to Qwen3_5ForCausalLM
  3. Patches tokenizer chat template to remove image/video branches
  4. Runs structural verification + inference test

Limitations & Warnings

  • Uncensored model: Safety filtering has been significantly reduced. Outputs may be inappropriate. Review generations before public use.
  • Thinking model quirks: The model always generates a <think> block first. Use skip_special_tokens=False if you want to inspect the reasoning chain.
  • No vision capability: This is intentional. Use the original VLM if you need image/video input.
  • GPU offload with bf16 is broken: See Quick Start section above.

License

Apache 2.0 (same as the source model).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10Upload folder using huggingface_hubf38522c5.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration