library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.6-35B-A3B
tags: - abliterated
- uncensored
Huihui-Qwen3.6-35B-A3B-abliterated MXFP4 MOE
Original Model
- Base model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
- Architecture: Qwen3.6 MoE (Mixture of Experts)
- Base version: bf16
- License: apache-2.0
Quantized Model
- File:
Huihui-Qwen3.6-35B-A3B-abliterated-bf16_MXFP4_MOE.gguf - Format: GGUF (llama.cpp)
- Quantization: MXFP4 (4-bit Mixed Precision Floating Point for MoE)
- Type: MoE-optimized
Specifications
| Metric | Value |
|---|---|
| Original size (bf16) | 64.6 GB (69,376,637,664 bytes) |
| Quantized size (MXFP4 MOE) | 20.6 GB (22,064,888,416 bytes) |
| Compression ratio | ~3.15x |
| Active parameters | ~36B |
| Total experts | 8 experts (MoE) |
About MXFP4
MXFP4 (Mixed-precision FP4) is a 4-bit floating point quantization format specifically optimized for MoE (Mixture of Experts) models. This format:
- Uses block floating point representation (per-block scaling)
- Maintains precision in sensitive tensors
- Optimizes expert weights while keeping bf16 in critical layers
- Compatible with llama.cpp and GGUF runners
Tensor Types Preserved in BF16
The following components are kept in bf16 to preserve model quality:
- Token embeddings (
token_embd.weight) - Normalization layers (
attn_norm,ssm_norm,ffn_norm, etc.) - MoE gate weights (
ffn_gate_inp,ffn_gate_inp_shexp) - Expert selection vectors
- Biases and normalization parameters
- Full attention layers
How to Use with llama.cpp
Prerequisites
You need to have llama.cpp installed or download a prebuilt binary. Visit the llama.cpp releases page to get the latest version for your platform.
Basic Usage
- Use the F32 version of the mmproj file for optimal results. Recommended quality ranking: F32 > BF16 > F16.
For configuration tips, follow the Unsloth Qwen3.5 local run guide
Usage Warning
This is an abliterated model - Safety filtering has been significantly reduced. Use with caution.
For more information about usage warnings, please refer to the model page on Hugging Face.