← back to catalog · registered 2026-08-22 13:56

Lord-H4D3ZS/Qwen3.8-Distill-35B-A3B-Coder-Abliterated

Lord-H4D3ZS Qwen 35B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Lord-H4D3ZS%2FQwen3.8-Distill-35B-A3B-Coder-Abliterated"
Response includes
  • classification m8
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 60,818
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
61K
21K last 30d - stable
Likes
25
Model age
8w ago
created 2026-08-15
Downloads over time
Now65.3K→from17.4K↑275%
15K33.4K51.7K70.1K17.4K on Aug 1965.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0 UGI
Natural Intelligence 25.43 UGI
Political lean -19.6% UGI
Sensitive-Info 14.03 UGI
SocPol 2.6 UGI
UGI 16.02 UGI
Willingness (10) 2 UGI
W10-Adherence 0 UGI
W10-Direct 4 UGI
Writing 35.83 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf qwen3_5_moe_text qwen3 moe a3b mtp rocmfpx code text-generation conversational base_model:Qwen/Qwen3.6-35B-A3B base_model:quantized:Qwen/Qwen3.6-35B-A3B

Related

Total size
11.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 13:57

Files by quantization

Auxiliary files 10 files 11.5 GB
Qwen3.8-Distill-35B-A3B-Coder-Abliterated-Q2KXL_ROCMFPX.gguf 11.4 GB 6e2208e5 download
tokenizer.json 19.1 MB 87a7830d download
chat_template.jinja 7.58 KB a8755d82 download
README.md 3.08 KB 00c78302 download
BUILD.md 3.01 KB ca26458a download
build_rocmfpx.sh 2.52 KB 48b7a938 download
config.json 2.26 KB 8cf8cc78 download
.gitattributes 1.63 KB 192a58c1 download
tokenizer_config.json 1.07 KB 6be6ce17 download
generation_config.json 213 B 79c8cce3 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
tags:

  • qwen3
  • moe
  • a3b
  • mtp
  • gguf
  • rocmfpx
  • code
    pipeline_tag: text-generation

Qwen3.8-Distill-35B-A3B-Coder-Abliterated (Q2 ROCmFPX, PoC)

A 2-bit ROCmFPX GGUF of a distilled Qwen3.6-35B-A3B (MoE, 256 experts / ~3B active), sized to
run on a 16GB consumer GPU. Ships with the MTP (nextn) head and build instructions for the
matching runtime.

Honest status: this is a proof-of-concept. On the internal 10-task smoke eval the distilled
model tied its base (6/10 vs 6/10) — no regression, no measurable gain yet — and it is now
quantized to 2-bit, which trades quality for fit. Publishing it as a reproducible artifact of the
pipeline (distill → graft MTP → ROCmFPX 2-bit GGUF), not as a benchmark-winning coder.
The quality fix is a larger, tool-calling-heavy corpus — a separate follow-up run.

What this is

  • Base / architecture: Qwen/Qwen3.6-35B-A3B (Qwen3_5MoeForCausalLM, 256 experts, ~3B
    active). The "3.8" in the name refers to the teacher, not the base.
  • Teacher: abliterated Qwen3.8-27B (GGUF Q8_0) via llama.cpp — sequence-level reasoning
    distillation (teacher <think> chains as SFT targets).
  • Method: Unsloth 4-bit QLoRA, completion_only_loss, 1 epoch / 850 teacher completions,
    merged to bf16, MTP head grafted back from base, converted + quantized with ROCmFPX.
  • Quant (the interesting part): a hand-built role-aware mix — 2-bit experts
    (Q2_0_ROCMFPX, the ~90% bulk) + Q6 attention / embeddings / shared-experts / output
    (Q6_0_ROCMFPX, the coherence-critical ~10%), norms in F32. 12GB total, fits a 16GB card
    with ~4GB left for KV/context. This is the llama.cpp/ROCmFPX analogue of the eschamoe/OTQ
    role-aware idea: pure 2-bit-everywhere collapses the model; keeping attention precise while
    2-bit'ing the experts preserves coherence. See the exact --tensor-type recipe in
    BUILD.md.
  • "Abliterated": transferred over the training corpus (teacher was abliterated) — corpus-scoped,
    NOT a globally abliterated model.

Run it

You need a llama-server built from the pinned ROCmFPX source — see BUILD.md.

llama-server -m *-Q2_ROCMFPX.gguf --host 127.0.0.1 --port 8080 \
  -ngl 99 -c 16384 -fa on --jinja --alias qwen38-distill-a3b
# OpenAI-compatible API at http://127.0.0.1:8080/v1

16GB card: context and concurrency share one KV pool — pick single-stream long context
(-c 32768 -np 1) or many short sessions (-c 8192 -np 8).

Known limitations (measured)

  • No accuracy gain over base yet; 2-bit lowers quality further.
  • Weak on tool-calling/agentic tasks (thin PoC corpus) — the first thing the next run must fix.
  • MTP nextn tensors are present but speculative decoding depends on your runtime's support
    (see BUILD.md). Text-only; no vision.

Files

  • *-Q2_ROCMFPX.gguf — the model (~16GB-card fit)
  • BUILD.md — build the ROCmFPX runtime (pinned commit b2f5829)
  • build_rocmfpx.sh — exact build script used

Apache-2.0, inheriting the base model's terms.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Upload README.md with huggingface_hubd8633a93.1 KB
    Loading...
  2. 2026-08-15Add files using upload-large-folder tool4a923401.8 KB
    Loading...
  3. 2026-08-15initial commit0c8ba1f28 B
    Loading...

Discussions 4 threads

  1. 2026-08-29Support NVIDIA cardsopen1 💬#5
    Loading...
  2. 2026-08-21MXFP4 or Q4_1?open4 💬#4
    Loading...
  3. 2026-08-17🚩 Report: Spamopen2 💬#3
    Loading...
  4. 2026-08-15🚩 Report: Spamclosed5 💬#2
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration