← back to catalog · registered 2026-08-22 13:56

maci0/Qwopus3.6-27B-v2-abliterated-NVFP4

maci0 24B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/maci0%2FQwopus3.6-27B-v2-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 606
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
606
45 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-04
Downloads over time
Now622→from0↑0%
02284566840 on Jul 1622 on Oct 11622 on Oct 9JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 w4a4 gptq quantized compressed-tensors llm-compressor vllm

Related

Total size
19.1 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 01:51

Files by quantization

Auxiliary files 12 files 19.2 GB
model.safetensors 17.5 GB 65828328 download
model-towers.safetensors 1.65 GB 3126fb89 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 235 KB b5d11b7e download
README.md 13.8 KB 6ee03d04 download
config.json 8.79 KB 7671e5e7 download
chat_template.jinja 6.83 KB ea146234 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.30 KB 32f01452 download
tokenizer_config.json 1.17 KB 8943d027 download
recipe.yaml 326 B 0ed8315b download
generation_config.json 168 B c57ea5c1 download

README current version from Hugging Face


base_model: Jackrong/Qwopus3.6-27B-v2
base_model_relation: quantized
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • nvfp4
  • fp4
  • w4a4
  • gptq
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
  • qwen3_5
  • vision-language
  • thinking
  • uncensored
  • abliterated
  • heretic

RQ-27B-V2-AAbliterated
Qwopus3.6-27B-v2 abliterated · NVFP4
27B 256K general-purpose reasoner · refusals removed (2-round Heretic).
Params27B
Active27B (dense)
Size18 GB
Perplexity6.92
Refusals99 → 9 / 100
Context256K
MTP headbf16

TL;DR: Qwopus3.6-27B-v2 abliterated, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 18 GB, wikitext-2 PPL 6.92, 256K general-purpose reasoner, refusals removed (2-round Heretic).

Qwopus3.6-27B-v2 abliterated NVFP4

Jackrong/Qwopus3.6-27B-v2,
abliterated (refusal direction removed) with Heretic in
two iterative rounds, then quantized to NVFP4 (W4A4) in the compressed-tensors
nvfp4-pack-quantized format with llm-compressor
(GPTQ + MSE, shared fused-layer scales).

Near-lossless and decensored. Two Heretic rounds cut refusals from 99/100 to 9/100 of
held-out harmful prompts while keeping a KL divergence of 0.0160 to the original model
(well under the 0.5 line that signals capability damage). NVFP4 then compresses to ~18 GB with a
wikitext-2 perplexity of 6.92.

  • Built for vLLM on NVIDIA Blackwell (4-bit weight + 4-bit activation). Pre-Blackwell GPUs
    run it weight-only.
  • Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).

Uncensored / abliterated model. It follows instructions without refusal guardrails. The
abliteration only removes refusals; all other behaviour comes from the base model. You
are responsible for how you use it.

Fidelity

Near-lossless versus the bf16 source, 18 GB vs 55.6 GB bf16 (~33%), at wikitext-2 perplexity 6.92 and KL divergence 0.0160 to the original. GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

NVFP4 is auto-detected from config.json (compressed-tensors); no quantization flag
needed. --reasoning-parser qwen3 splits the <think> block into reasoning_content.

vllm serve maci0/Qwopus3.6-27B-v2-abliterated-NVFP4 \
  --served-model-name qwopus-27b-v2-abliterated-nvfp4 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder
  • Supports up to 262144 tokens; keep at least 128K to preserve thinking quality.
    --max-model-len 131072 is a safe default; raise it if memory allows.
  • Add --language-model-only to skip the vision tower and free KV cache for text use.
  • The parser flags are not auto-detected; pass them explicitly. Drop the tool-call line
    if you do not need tool calling.

Python (OpenAI client)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
r = client.chat.completions.create(
    model="qwopus-27b-v2-abliterated-nvfp4",
    messages=[{"role": "user", "content": "Explain, step by step, why the sky is blue."}],
)
print(r.choices[0].message.content)

curl

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwopus-27b-v2-abliterated-nvfp4",
  "messages": [{"role": "user", "content": "Explain, step by step, why the sky is blue."}]
}'

About the base model

A 27B Qwen3.5-family vision-language model (Qwopus 3.6 v2), a general-purpose
reasoning and instruction-following model with thinking-mode reasoning and a 256K
context window.

  • 64 decoder layers: hybrid gated delta-net linear attention plus full attention, dense
    MLP, plus a vision tower for image and video input.
  • 256K context (max_position_embeddings 262144).
  • Thinking mode by default, with an instruct toggle.

Abliteration

Heretic removes the refusal direction with a TPE-optimized
search over per-component ablation strength, jointly minimizing refusal rate and KL divergence
from the original model, then merges the best trial. This model was abliterated in two iterative
rounds
: round 1 removed the dominant refusal direction, then Heretic was re-run on the round-1
model to remove the residual refusal direction that surfaced once the first was gone. Because Qwopus
is a thinking model, evaluation ran in non-thinking mode so each judged response is a real answer
rather than an unfinished <think> block; the shipped model restores the original thinking chat
template.

Round Refusals KL divergence Note
Baseline 99/100 original model
Round 1 32/100 0.0235 dominant refusal direction removed
Round 2 9/100 0.0160 residual direction removed (shipped)
  • Datasets: mlabonne/harmless_alpaca (good) vs mlabonne/harmful_behaviors (bad).
  • This checkpoint required Heretic with --row-normalization NONE. The default (FULL)
    degenerated this model, producing a broken output; run with --row-normalization NONE
    to reproduce.

Quantization

Scheme NVFP4, W4A4
Weight rounding GPTQ (Hessian-based error compensation), MSE observer
Weights FP4 (E2M1), group_size=16, tensor_group, FP8 (E4M3) group scales, shared across fused layers
Activations FP4, dynamic per-group, FP8 (E4M3) scales
Quantized all language-model Linear layers
Kept in bf16 vision tower (model.visual.*), lm_head, MTP head
Untouched gated delta-net Conv1d and SSM params (A_log, dt_bias), never Linear

GPTQ is a quantization-time cost only; inference speed and format are identical to
plain round-to-nearest NVFP4, but it chooses better 4-bit values.

Calibration: 512 domain-matched samples (long reasoning + general chat + code),
max_seq_len=2048, text-only path through the VL model.

Recommended sampling

Thinking mode is the default.

  • Thinking, precise: temperature=0.6, top_p=0.95, top_k=20
  • Thinking, general: temperature=1.0, top_p=0.95, top_k=20
  • Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20
  • To run non-thinking, set {%- set enable_thinking = false %} in the chat template, or
    pass extra_body={"chat_template_kwargs": {"enable_thinking": false}}.

Related

Notes

  • Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • --reasoning-parser and --tool-call-parser are not auto-detected; pass them explicitly.
  • Thinking mode is on by default; toggle it via the chat template or chat_template_kwargs.
  • No refusal guardrails; you are responsible for how you use it.

License

Apache-2.0, following the base model. Intended use and all responsibility for use follow
the base model.

Credits

Part of Rogue Quants · NVFP4 component datasheets · collection. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).

README history 9 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04uniform schema, no em-dashes18d3a3513.8 KB
    Loading...
  2. 2026-08-04card: uniform characteristics schema (params/active/experts/etc)cb3cdd313.8 KB
    Loading...
  3. 2026-08-03trim redundant Fidelity table (covered by header)5e9841012.2 KB
    Loading...
  4. 2026-08-03card: readability - sans labels, mono values, more spacingb036daa12.3 KB
    Loading...
  5. 2026-08-03card: fix table borders + readabilityb60c0e612.3 KB
    Loading...
  6. 2026-08-03card: die sigil + a11y table semantics23fc27b12 KB
    Loading...
  7. 2026-08-03card: DATASHEET design language9fa078411.4 KB
    Loading...
  8. 2026-07-07Add Ornith-35B MoE abliterated to Related siblings0c8a5f313.9 KB
    Loading...
  9. 2026-07-04Add model card (2-round abliteration, row-norm NONE)8d34d2413.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration