← back to catalog · registered 2026-08-22 13:56

Blackfrost-AI/PINQWEN-3.6-27B-NVFP4-ABLITERATED

Blackfrost-AI Qwen 27B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Blackfrost-AI%2FPINQWEN-3.6-27B-NVFP4-ABLITERATED"
Response includes
  • classification m1
  • files 10
  • author_summary 19 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
115
↑ 101% in 90 days
Likes
3
Model age
3mo ago
created 2026-07-08
Downloads over time
Now732→from365↑101%
347487628769365 on Jul 15732 on Sep 16JulAugSep
Jul 15 → Sep 16 · 28 snapshots · spans 63 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text agentic tool-calling nvfp4 qwen3.6 modelopt text-generation conversational base_model:huihui-ai/Huihui-Qwen3.6-27B-abliterated

Related

Total size
18.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-26 16:24

Files by quantization

Auxiliary files 10 files 18.3 GB
model.safetensors 18.3 GB 8ea3ad9a download
tokenizer.json 19.1 MB 87a7830d download
878395f6-e83a-4f2e-a63f-f7c5111ba0cb.jpg 293 KB 62f29b00 download
chat_template.jinja 7.58 KB a8755d82 download
config.json 7.01 KB c017414c download
hf_quant_config.json 3.33 KB f1f9d24f download
README.md 3.18 KB e767debe download
.gitattributes 1.61 KB 269c43e4 download
tokenizer_config.json 1.17 KB 4d9ac0cf download
generation_config.json 214 B c8e006f3 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.6-27B-abliterated
    pipeline_tag: text-generation
    library_name: transformers
    tags:
  • agentic
  • tool-calling
  • nvfp4
  • qwen3.6
  • modelopt

PINQWEN-3.6-27B-NVFP4-ABLITERATED

878395f6-e83a-4f2e-a63f-f7c5111ba0cb

Blackfrost internal archive. Agentic / tool-calling specialist — Qwen3.6-27B fine-tuned on the The Void (v3) corpus, then quantized to NVFP4 (nvidia-modelopt) for fast serving on Blackwell (SM120).


At a glance

Field Value
Base huihui-ai/Huihui-Qwen3.6-27B-abliterated (Qwen3.6-27B — hybrid Gated-DeltaNet + attention, abliterated)
Params ~27 B (hybrid dense, 64-layer + 1 MTP)
Fine-tune data The Void (v3) — 5,025 records (3,132 knowledge/CoT + 1,893 agentic ReAct, 38% agentic)
Method unsloth QLoRA, rank 32, α 32, all 7 proj targets, multi-turn masking, length-grouped
Schedule 3 epochs, 474 steps, lr 2e-4 linear, bsz 2 × accum 2 × 8 GPU (eff 32)
Hardware 8× RTX PRO 6000 Blackwell (g4-standard-384), ~1.5 h wall
Quant NVFP4 (nvidia-modelopt NVFP4_DEFAULT_CFG), text-only, MTP grafted
Serves on vLLM, Blackwell SM120
Trained 2026-07-08 · Blackfrost-AI

Benchmark — agent_benchmark.py (vs. base, identical battery)

Metric Base (abliterated) VEGA-27B
Tool-calling accuracy 7 / 8 8 / 8
Avg latency / response 10.5 s 6.2 s (−41%)
Avg tokens / response (verbosity) 267 158 (−41%)
Multi-turn drivability 2 / 2 1 / 2 (mock-tool keyword artifact — not a real regression; needs a live tool to score)

Takeaway: the The Void (v3) training made an already-strong base more accurate at tool selection, ~40% faster, and ~40% more concise.


Quantization details (why this one serves, unlike a naive convert)

  • Tool: nvidia-modelopt (NVFP4_DEFAULT_CFG) — the fast SM120 path. llm-compressor/compressed-tensors NVFP4 null-outputs on SM120 + vLLM for this hybrid family, so modelopt is required.
  • Arch kept: Qwen3_5ForConditionalGeneration + language_model_only: true. vLLM has no text-only class for the Qwen3.5 hybrid family, so the full class + this flag is how a text model loads.
  • Kept in bf16 (quant ignore list): linear_attn.conv1d (the Gated-DeltaNet conv), lm_head, vision tower, MTP head.
  • MTP head (15 tensors) grafted back in bf16 → working speculative-decoding draft path.
  • Vision tower stripped → text-only, smaller + faster than the VLM variant.
  • Recipe: lna-lab/GGUF-to-NVFP4-SM120 qwen36_27b_text_mtp.py.

Serving (vLLM on Blackwell)

vllm serve Blackfrost-AI/VEGA-27B-NVFP4 \
  --trust-remote-code \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --max-model-len 131072

No --quantization flag — vLLM auto-detects the modelopt NVFP4 checkpoint. Tool calls use Qwen3's XML format (<function=name><parameter=x>…), parsed by qwen3_xml.


README history 9 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-26Update model card: corrected license, base_model lineage, sanitized tagse1cbb7a3.2 KB
    Loading...
  2. 2026-07-08Update README.mddda61313.2 KB
    Loading...
  3. 2026-07-08Card: brand as The Void (not god-void)5722bbd3.5 KB
    Loading...
  4. 2026-07-08Update README.md7a202843.5 KB
    Loading...
  5. 2026-07-08Update README.md8e0734b3.5 KB
    Loading...
  6. 2026-07-08Update README.md6981a8b3.5 KB
    Loading...
  7. 2026-07-08Update README.md8210e1a3.5 KB
    Loading...
  8. 2026-07-08Update README.md5fe31923.4 KB
    Loading...
  9. 2026-07-08Upload README.md with huggingface_hubb2c96743.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration