license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.5-2B-abliterated
tags:
- iron
- amd-npu
- ryzen-ai
- xdna
- image-text-to-text
- qwen3_5
- conversational
- abliterated
- uncensored
- npu2
Huihui-Qwen3.5-2B-abliterated chat (text and images) on the AMD NPU
An IRON export of huihui-ai/Huihui-Qwen3.5-2B-abliterated for AMD Ryzen AI NPUs: the compiled NPU kernels (.xclbin + instruction streams) and the packed weights an IRON Rust runtime replays.
The kernels are compiled for NPU2 (AIE2P: Strix Point, Strix Halo, Krackan) and will not load on NPU1 (Phoenix, Hawk Point).
Download
hf download brishen/iron-huihui-qwen3.5-2b-abliterated-npu2 --local-dir qwen3.5-2b-abliterated
# or, from an IRON checkout:
python scripts/hf_models.py download qwen3.5-2b-abliterated --repo brishen/iron-huihui-qwen3.5-2b-abliterated-npu2 --out qwen3.5-2b-abliterated
Qwen/Qwen3.5-2B with its refusal direction projected out of the weights (huihui-ai's abliteration): the same architecture, tokenizer and chat template, exported exactly as brishen/iron-qwen3.5-2b-npu2 is. The text model and the vision tower (the multi-token prediction head is not exported). Every projection runs on the NPU: flm.GEMMs over the prompt and the image's patches, GEMVbfp16s for each generated token (~10 tokens/s on a Ryzen AI 9 HX 370), and the taconite-qwen35 Rust runtime replays it over XRT or directly over the amdxdna driver. ~2.5 GB.
Usage warnings
Abliteration removes the model's refusal behaviour, so it will answer
requests the original declines and its output is not safety-filtered.
Review what it generates, keep it to research, testing and controlled
settings rather than public-facing products, and follow your local laws;
see the upstream card's
warnings, which apply here unchanged.
Requirements
- An AMD Ryzen AI NPU2 (Strix Point, Strix Halo, Krackan) on Linux
with theamdxdnadriver and its firmware (/dev/accel/accel0), and
an unlimited locked-memory limit (ulimit -l unlimited): the weights
live in ~2.5 GB of NPU-visible buffers. - XRT, unless the runtime is built with
--features direct(straight to
the driver's ioctls, no XRT at all). - The model uses 3 of the NPU's 16 hardware-context slots (shared by all
processes). When other programs hold the rest, the runtime swaps its
contexts rather than failing (--timingreports the swaps).
Usage
With the taconite-qwen35
runtime (docs):
cargo install taconite-qwen35 # over XRT
# or, with no XRT at all:
cargo install taconite-qwen35 --no-default-features --features cli,direct
qwen35 qwen3.5-2b-abliterated --prompt "Why is the sky blue?" # streams the answer
qwen35 qwen3.5-2b-abliterated --image photo.jpg --prompt "Describe this image." # --image repeats
qwen35 qwen3.5-2b-abliterated --thinking --max-new 1024 --prompt "..." # reason in <think> first
qwen35 qwen3.5-2b-abliterated --interactive # one prompt a line
qwen35 check qwen3.5-2b-abliterated # verify this bundle
As a library:
use taconite_qwen35::{ChatOptions, Qwen35, RgbImage};
let mut q = Qwen35::load(std::path::Path::new("qwen3.5-2b-abliterated"), 8192)?;
let opts = ChatOptions { images: vec![RgbImage { width, height, rgb }], ..Default::default() };
let (text, stats) = q.chat("Describe this image.", &opts, |s| print!("{s}"))?;
How it runs
Every weight matrix runs on the NPU; the host does the rest:
| NPU | host (f32) | |
|---|---|---|
| the prompt | every projection as an flm.GEMM over 256-row chunks, one hardware context |
tokenizer, embedding, norms, the Gated DeltaNet's causal conv and recurrence, partial M-RoPE, GQA attention |
| each generated token | every projection and the tied LM head as a GEMVbfp16, a second context |
the same, one row; greedy sampling |
| an image | the 24-block vision tower's projections as flm.GEMMs, a third context |
resize and patchify (torchvision-exact), LayerNorms, 2D RoPE, attention |
Weights are stored once, as bfp16 (8-bit mantissas, one exponent a block
of 8: ~9 bits a weight): the decode GEMVs read the prefill GEMMs' packed
streams, and the input embedding is decoded from the tied LM head's. The
vision tower's weights are int8 per output channel, held exactly in bfp16
blocks with each channel's leftover scale factor applied on the host (a
single bfp16 block of the float weights loses too much over its 24
blocks). Activations enter the multiplies as bfp16 hi + lo (~16 bits).
Accuracy
qwen35 check qwen3.5-2b-abliterated on a Ryzen AI 9 HX 370, against transformers float32
(these weights) on the ROCm iGPU, the next token teacher-forced on the
reference's:
| prompt | prompt tokens | steps | top-1 agrees | disagreements (the reference's top-2 margin) | max |dlogit| (reference's top 16) |
|---|---|---|---|---|---|
| "Explain in three sentences why the sky is blue." | 23 | 95 | 92/95 | 3, all near-ties (<= 0.20) | 0.40 |
| a model-card summary | 778 | 134 | 132/134 | 2, both exact ties (0.00) | 0.55 |
| a kitchen photo + "Describe this image in detail." | 280 | 160 | 150/160 | 10, all near-ties (<= 0.22) | 0.86 |
No disagreement where float32's top two are 0.5 or more apart. The
tokenizer and chat template are Qwen3.5-2B's and reproduce Hugging Face's
ids exactly, the image preprocessing matches the Python reference bit for
bit, and the vision tower's output has cosine 0.992 against float32. The
bundle carries these references, so qwen35 check repeats the comparison
on your machine.
Speed
Ryzen AI 9 HX 370 (NPU2): setup ~2 s; prompt 0.2 s for 23 tokens, 1.3 s
for 778; ~93 ms a generated token (~10 tokens/s), bound by the ~2.1 GB of
weights a step streams from memory; an image of ~1M pixels (1040 patches)
~2-2.5 s through the vision tower (NPU ~0.25 s, host attention ~1.7 s).
Limits
- Greedy decoding only.
- A prompt of up to 2048 tokens (image tokens included); context up to
--max-ctx(default 8192). - Images are resized as the checkpoint's processor does, capped at ~1M
pixels (4096 patches, 1024 image tokens); no video. - The checkpoint's multi-token-prediction head is not used.
Files
manifest.txt (the kernels and model constants), tensors.txt /tensors.bin (packed weights and references), kernels/ (xclbins and
instruction streams), tokenizer.json (the checkpoint's). Exported by
IRON's iron/applications/qwen3_5/export_qwen35.py.
Provenance
- Upstream model:
huihui-ai/Huihui-Qwen3.5-2B-abliterated(license:apache-2.0; its terms apply to these weights) - IRON commit:
55484ce - Uploaded: 2026-09-30
- Files: 141, 2.5 GB