← back to catalog · registered 2026-08-22 13:56

vladimir94/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF

vladimir94 Qwen 35B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/vladimir94%2FHuihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF"
Response includes
  • classification m8
  • files 5
  • hub_downloads_all_time 10,820
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
11K
651 last 30d - cooling
Likes
0
Model age
4mo ago
created 2026-06-06
Downloads over time
Now11K→from948↑1,064%
4444.3K8.2K12K948 on Jun 1011K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
BF16
Tags
llama.cpp gguf qwen3.6 qwen3_5_moe moe nvfp4 blackwell-native-fp4 mtp speculative-decoding mmproj vision vision-to-text
Total size
24.4 GB
Files
5
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-06-06 02:08

Files by quantization

BF16 2 files 24.4 GB
Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf 21.0 GB 58fdf6c8 download
mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf 3.48 GB 0f31fdde download
F16 1 file 858 MB
mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf 858 MB 3a3000f3 download
Auxiliary files 2 files 10.1 KB
README.md 8.40 KB 030ddec0 download
.gitattributes 1.74 KB c76ccf57 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • YuYu1015/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
  • huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
    base_model_relation: quantized
    library_name: llama.cpp
    pipeline_tag: text-generation
    tags:
  • gguf
  • llama.cpp
  • qwen3.6
  • qwen3_5_moe
  • moe
  • nvfp4
  • blackwell-native-fp4
  • mtp
  • speculative-decoding
  • mmproj
  • vision
  • vision-to-text
  • text-to-text
  • abliterated
  • uncensored
  • dgx-spark
  • gb10
  • sm121
    language:
  • en
  • zh

Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF

GGUF conversion of YuYu1015/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4, itself an NVFP4 W4A4 quantization of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated.

This repository is intended for llama.cpp on NVIDIA Blackwell systems such as DGX Spark / GB10. The text model is provided as a main GGUF plus a separate MTP draft GGUF for llama.cpp speculative decoding. An optional mmproj file is also included for image input.

This is an uncensored/abliterated model. Use it responsibly.

Files

File Size Purpose
Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf 20.95 GiB Main text model, MTP excluded
mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf 3.48 GiB Separate MTP draft model for llama.cpp
mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf 0.84 GiB Optional vision projector

The BF16 suffix means the converter used BF16 for non-NVFP4 tensors and fallback tensors. The source checkpoint is YuYu1015's NVFP4 packed model, and the llama.cpp runtime used for testing reported BLACKWELL_NATIVE_FP4 = 1.

Conversion Notes

Converted from the downloaded Hugging Face safetensors checkpoint with llama.cpp:

  • llama.cpp commit: 65ef50a0a4bb240211a41d43c957ae6313af6841
  • llama.cpp version output: version: 1 (65ef50a)
  • Platform used for conversion/test: NVIDIA DGX Spark / GB10, Linux aarch64
  • Runtime build reported: BLACKWELL_NATIVE_FP4 = 1

Main model and MTP were split deliberately:

  • --no-mtp was used for the main model.
  • --mtp was used on a second pass to create the standalone draft GGUF.
  • This split duplicates some shared tensors, but lets llama.cpp load the MTP head as --model-draft.

Equivalent commands:

SRC=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
OUT=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF
LLAMA=/path/to/llama.cpp

python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
  --outfile "$OUT/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
  --outtype bf16 \
  --no-mtp

python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
  --outfile "$OUT/mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
  --outtype bf16 \
  --mtp

The vision projector needed a small conversion wrapper because this checkpoint stores visual tensors under model.language_model.visual.*, while the llama.cpp Qwen3-VL mmproj converter expects visual.*. The wrapper remaps only that prefix and then calls convert_hf_to_gguf.py.

Equivalent mmproj command after applying that prefix wrapper:

python3 convert_qwen35_mmproj.py "$SRC" \
  --outfile "$OUT/mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf" \
  --outtype f16 \
  --mmproj

Recommended llama.cpp Settings

Text-only, 262K context, MTP enabled:

llama-server \
  --model Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
  --model-draft mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
  --alias huihui-qwen3.6-35b-a3b-abliterated-uncensored-nvfp4-mtp \
  --ctx-size 262144 \
  --parallel 8 \
  --batch-size 8192 \
  --ubatch-size 2048 \
  --flash-attn on \
  --n-gpu-layers all \
  --kv-unified \
  --cont-batching \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --cache-type-k-draft q8_0 \
  --cache-type-v-draft q8_0 \
  --spec-type draft-mtp \
  --spec-draft-n-max 1 \
  --draft-p-min 0.3 \
  --cache-ram 8192 \
  --reasoning off

For vision, add:

--mmproj mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf

For the tested deployment, the non-vision preset used 8 slots. The vision preset used 2 slots, because loading mmproj adds memory and disables some cache-reuse behavior in llama.cpp multimodal mode.

MTP Parameter Testing

All numbers below are llama.cpp generation throughput on DGX Spark / GB10 at 262K context with the text-only GGUF, --flash-attn on, Q8_0 KV cache, and native Blackwell FP4 support enabled. Continuous batching was enabled for the live llama-server tests and was verified with MTP. These are local smoke/throughput tests, not a formal benchmark suite.

Setup Throughput Notes
No speculative decoding ~31.0 tok/s Baseline
MTP, --spec-draft-n-max 1, --draft-p-min 0.0 ~34.76 tok/s Acceptance ~71.7%
MTP, --spec-draft-n-max 1, --draft-p-min 0.3 ~34.84 tok/s Acceptance ~78.3%; selected setting
MTP, --spec-draft-n-max 2, --draft-p-min 0.0 ~33.99 tok/s Slower than n=1
MTP, --spec-draft-n-max 2, --draft-p-min 0.3 ~33.07 tok/s Slower than n=1
MTP, --spec-draft-n-max 3, --draft-p-min 0.0 ~31.88 tok/s Close to baseline
MTP, --spec-draft-n-max 4, --draft-p-min 0.0 ~29.15 tok/s Slower than baseline

Recommended MTP setting for this GGUF:

--spec-type draft-mtp --spec-draft-n-max 1 --draft-p-min 0.3

Higher MTP depths were not useful in the local tests. We stopped the sweep after n=4 because throughput was already declining, and kept n=1.

DFlash was not benchmarked for this GGUF/llama.cpp release. The source NVFP4 model card discusses DFlash for vLLM, but this repository is focused on llama.cpp with the model's built-in MTP draft.

DGX Spark / GB10 Throughput and Memory

Tested on a DGX Spark / GX10 with 128GB unified memory.

One-slot temporary llama.cpp test:

  • Best observed setting: MTP n_max=1, draft_p_min=0.3
  • Generation throughput: ~34.8 tok/s

Eight-slot live llama-server deployment:

  • Server loaded with --parallel 8, --ctx-size 262144, --kv-unified, --cont-batching
  • Each slot reported n_ctx = 262144
  • Single active generation in the 8-slot continuous-batching server: ~30.6 to ~31.2 tok/s
  • NVIDIA accounting for the worker process: ~30.4 GiB
  • The separate llama-server controller process used about 170 MiB

The 8-slot throughput number is a single active request on an 8-slot server, not an 8-concurrent aggregate throughput benchmark.

Because --kv-unified is enabled, filling slots with context does not allocate eight separate full-size KV buffers. The live worker stayed around 30-31 GiB by nvidia-smi. Host memory can still grow due to prompt/checkpoint caching; the tested config set --cache-ram 8192.

Vision / mmproj Status

The optional mmproj file was converted and a local image-input smoke test succeeded. The test image had a red/blue split, and the model correctly described red and blue side-by-side.

Recommended usage:

  • Use the text-only model for normal assistant/chat workloads.
  • Load mmproj only when image input is needed.
  • Treat vision support as available but less extensively benchmarked than text generation.

Known multimodal caveat from local llama.cpp logs:

  • With mmproj loaded, llama.cpp reported that cache reuse is not supported by multimodal mode and disabled it.

Prompt Cache and Cache Reuse Notes

The tested llama.cpp deployment used:

--cache-ram 8192
--kv-unified
--cont-batching

Prompt/checkpoint caching worked in the sense that idle slots were saved and restored from the prompt cache. However, for this hybrid Qwen3.6/GDN context, llama.cpp also logged that cache_reuse was not supported and disabled cache reuse. In practice, do not expect every repeated prompt to show classic prefix-cache hit behavior.

Safety

This is an abliterated/uncensored model. It may produce unsafe, offensive, or policy-violating content. Users are responsible for deployment choices, filtering, logging, and compliance with applicable law.

Credits

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-06Add Huihui Qwen3.6 NVFP4 GGUF conversion0c74a658.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration