← back to catalog · registered 2026-08-22 13:56

maci0/Qwopus3.6-27B-Coder-abliterated-NVFP4

maci0 24B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/maci0%2FQwopus3.6-27B-Coder-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 814
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
814
65 last 30d - cooling
Likes
3
Model age
3mo ago
created 2026-07-02
Downloads over time
Now845→from28↑2,918%
030961892728 on Jul 1845 on Oct 11845 on Oct 8JulAugSepOct
Jul 1 → Oct 11 · 55 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 w4a4 gptq quantized compressed-tensors llm-compressor vllm

Related

Total size
19.1 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 01:51

Files by quantization

Auxiliary files 12 files 19.2 GB
model.safetensors 17.5 GB 38a3b67c download
model-towers.safetensors 1.65 GB 3126fb89 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 235 KB b5d11b7e download
README.md 13.7 KB a2a38030 download
config.json 8.79 KB a9b37894 download
chat_template.jinja 4.61 KB 8ad2f87b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.30 KB 32f01452 download
tokenizer_config.json 1.17 KB 8943d027 download
recipe.yaml 326 B 0ed8315b download
generation_config.json 214 B aeb4ce60 download

README current version from Hugging Face


base_model: Jackrong/Qwopus3.6-27B-Coder
base_model_relation: quantized
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • nvfp4
  • fp4
  • w4a4
  • gptq
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
  • qwen3_5
  • vision-language
  • thinking
  • code
  • coder
  • uncensored
  • abliterated
  • heretic

RQ-27B-CODER-AAbliterated
Qwopus3.6-27B-Coder abliterated · NVFP4
27B 256K agentic coder · tool-calling · refusals removed (2-round Heretic).
Params27B
Active27B (dense)
Size18 GB
Perplexity6.73
Refusals99 → 3 / 100
Context256K
MTP headbf16

TL;DR: Qwopus3.6-27B-Coder abliterated, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 18 GB, wikitext-2 PPL 6.73, 256K agentic coder, refusals removed (2-round Heretic).

Qwopus3.6-27B-Coder abliterated NVFP4

Jackrong/Qwopus3.6-27B-Coder,
abliterated (refusal direction removed) with Heretic in
two iterative rounds, then quantized to NVFP4 (W4A4) in the compressed-tensors
nvfp4-pack-quantized format with llm-compressor
(GPTQ + MSE, shared fused-layer scales).

Near-lossless and decensored. Two Heretic rounds cut refusals from 99/100 to 3/100 of
held-out harmful prompts while keeping a KL divergence of 0.0398 to the original model
(well under the 0.5 line that signals capability damage). NVFP4 then compresses to ~18 GB with a
wikitext-2 perplexity of 6.73.

  • Built for vLLM on NVIDIA Blackwell (4-bit weight + 4-bit activation). Pre-Blackwell GPUs
    run it weight-only.
  • Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).

Uncensored / abliterated model. It follows instructions without refusal guardrails. The
abliteration only removes refusals; all other behaviour comes from the base model. You
are responsible for how you use it.

Fidelity

Near-lossless versus the bf16 source, 18 GB vs 55.6 GB bf16 (~33%), at wikitext-2 perplexity 6.73 and KL divergence 0.0398 to the original. GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

NVFP4 is auto-detected from config.json (compressed-tensors); no quantization flag
needed. --reasoning-parser qwen3 splits the <think> block into reasoning_content;
--tool-call-parser qwen3_coder enables tool/function calling for agentic coding.

vllm serve maci0/Qwopus3.6-27B-Coder-abliterated-NVFP4 \
  --served-model-name qwopus-27b-coder-abliterated-nvfp4 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder
  • Supports up to 262144 tokens; keep at least 128K to preserve thinking quality.
    --max-model-len 131072 is a safe default; raise it if memory allows.
  • Add --language-model-only to skip the vision tower and free KV cache for text use.
  • The parser flags are not auto-detected; pass them explicitly.

Python (OpenAI client)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
r = client.chat.completions.create(
    model="qwopus-27b-coder-abliterated-nvfp4",
    messages=[{"role": "user", "content": "Write a Python function that merges two sorted lists."}],
)
print(r.choices[0].message.content)

curl

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwopus-27b-coder-abliterated-nvfp4",
  "messages": [{"role": "user", "content": "Write a Python function that merges two sorted lists."}]
}'

About the base model

A 27B Qwen3.5-family vision-language model specialized for code (Qwopus 3.6 Coder),
with thinking-mode reasoning and a 256K context window.

  • 64 decoder layers: hybrid gated delta-net linear attention plus full attention, dense
    MLP, plus a vision tower for image and video input.
  • 256K context (max_position_embeddings 262144).
  • Thinking mode by default, with an instruct toggle.

Abliteration

Heretic removes the refusal direction with a TPE-optimized
search over per-component ablation strength, jointly minimizing refusal rate and KL divergence
from the original model, then merges the best trial. This model was abliterated in two iterative
rounds
: round 1 removed the dominant refusal direction, then Heretic was re-run on the round-1
model to remove the residual refusal direction that surfaced once the first was gone. Because Qwopus
is a thinking model, evaluation ran in non-thinking mode so each judged response is a real answer
rather than an unfinished <think> block; the shipped model restores the original thinking chat
template.

Round Refusals Note
Baseline 99/100 original model
Round 1 25/100 dominant refusal direction removed
Round 2 3/100 residual direction removed (final)
  • Datasets: mlabonne/harmless_alpaca (good) vs mlabonne/harmful_behaviors (bad).
  • Final KL divergence 0.0398 (capability preserved, well under the 0.5 damage line).

Quantization

Scheme NVFP4, W4A4
Weight rounding GPTQ (Hessian-based error compensation), MSE observer
Weights FP4 (E2M1), group_size=16, tensor_group, FP8 (E4M3) group scales, shared across fused layers
Activations FP4, dynamic per-group, FP8 (E4M3) scales
Quantized all language-model Linear layers
Kept in bf16 vision tower (model.visual.*), lm_head, MTP head
Untouched gated delta-net Conv1d and SSM params (A_log, dt_bias), never Linear

GPTQ is a quantization-time cost only; inference speed and format are identical to
plain round-to-nearest NVFP4, but it chooses better 4-bit values.

Calibration: 512 domain-matched samples (long reasoning + general chat + code),
max_seq_len=2048, text-only path through the VL model.

Recommended sampling

Thinking mode is the default.

  • Thinking, precise coding: temperature=0.6, top_p=0.95, top_k=20
  • Thinking, general: temperature=1.0, top_p=0.95, top_k=20
  • Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20
  • To run non-thinking, set {%- set enable_thinking = false %} in the chat template, or
    pass extra_body={"chat_template_kwargs": {"enable_thinking": false}}.

Related

Notes

  • Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • --reasoning-parser and --tool-call-parser are not auto-detected; pass them explicitly.
  • Thinking mode is on by default; toggle it via the chat template or chat_template_kwargs.
  • No refusal guardrails; you are responsible for how you use it.

License

Apache-2.0, following the base model. Intended use and all responsibility for use follow
the base model.

Credits

Part of Rogue Quants · NVFP4 component datasheets · collection. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04uniform schema, no em-dashes87a585913.7 KB
    Loading...
  2. 2026-08-04card: uniform characteristics schema (params/active/experts/etc)2f5e52913.7 KB
    Loading...
  3. 2026-08-03trim redundant Fidelity table (covered by header)f25633c12.7 KB
    Loading...
  4. 2026-08-03card: readability - sans labels, mono values, more spacingfff9e6712.7 KB
    Loading...
  5. 2026-08-03card: fix table borders + readability7ce639e12.8 KB
    Loading...
  6. 2026-08-03card: die sigil + a11y table semantics9999b7012.3 KB
    Loading...
  7. 2026-08-03card: DATASHEET design languagee4f77b911.7 KB
    Loading...
  8. 2026-07-07Add Ornith-35B MoE abliterated to Related siblings53a247714 KB
    Loading...
  9. 2026-07-04Add Qwopus-27B-v2 abliterated to Related siblingsf86192e13.9 KB
    Loading...
  10. 2026-07-02Add model card (2-round abliteration)50500e913.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration