license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3-Coder-30B-A3B-Instruct-abliterated
tags:
- gguf
- uncensored
- abliterated
- code
- moe
- ollama
Huihui-Qwen3-Coder-30B-A3B-Instruct-abliterated-i1-GGUF — AI-conn fork
This is a fork, not our model. The abliterated weights are by
huihui-ai;
this GGUF conversion is by
mradermacher;
the base model is Qwen/Qwen3-Coder-30B-A3B-Instruct.
This fork exists so the fleet's copy cannot be altered or pulled upstream,
and to carry the fit notes below. License: Apache-2.0 (per the source repo).
Uncensored coder pick: MoE, 30B total / ~3B active per token — but note that
all expert weights must sit in VRAM for GPU residency; "3B active" does
not shrink the memory bill.
Recommended file for a 20 GB card: Huihui-Qwen3-Coder-30B-A3B-Instruct-abliterated.i1-Q4_K_M.gguf (17.28 GiB).
Ollama pull:
ollama pull hf.co/AI-conn/Huihui-Qwen3-Coder-30B-A3B-Instruct-abliterated-i1-GGUF:i1-Q4_K_M
Prime fit (measured budget, RX 7900 XT — 19.98 GiB VRAM, 32 GB RAM)
- i1-Q4_K_M 17.28 GiB + small MoE KV + ~0.7 GiB reserve ≈ 18.3 GiB —
resident, tight - i1-IQ4_XS (15.24 GiB) is the roomy option if you want more context headroom
at slightly lower quant fidelity. - Independent datapoint (community harness): the stock base scores 60.4% on
HumanEval Pro. No independent coding eval of any abliterated build exists
yet — our fleet settles coder rankings with a local bake-off, not cards.