← back to catalog · registered 2026-09-30 09:58

brishen/iron-huihui-qwen3.5-2b-abliterated-npu2

brishen 2B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/brishen%2Firon-huihui-qwen3.5-2b-abliterated-npu2"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-30

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
iron amd-npu ryzen-ai xdna image-text-to-text qwen3_5 conversational abliterated uncensored npu2 base_model:huihui-ai/Huihui-Qwen3.5-2B-abliterated base_model:finetune:huihui-ai/Huihui-Qwen3.5-2B-abliterated

Related

Total size
2.47 GB
Files
6
Quantizations
1
Registered
2026-09-30 09:58
Last updated on HF
2026-09-30 09:50

Files by quantization

Auxiliary files 6 files 2.48 GB
tensors.bin 2.47 GB 7e30f1ca download
tokenizer.json 12.2 MB 5f9e4d49 download
tensors.txt 26.4 KB bf8e5f0b download
manifest.txt 20.6 KB f8d589fe download
README.md 6.96 KB 7b5fd3af download
.gitattributes 1.64 KB c62d725f download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.5-2B-abliterated
tags:

  • iron
  • amd-npu
  • ryzen-ai
  • xdna
  • image-text-to-text
  • qwen3_5
  • conversational
  • abliterated
  • uncensored
  • npu2

Huihui-Qwen3.5-2B-abliterated chat (text and images) on the AMD NPU

An IRON export of huihui-ai/Huihui-Qwen3.5-2B-abliterated for AMD Ryzen AI NPUs: the compiled NPU kernels (.xclbin + instruction streams) and the packed weights an IRON Rust runtime replays.

The kernels are compiled for NPU2 (AIE2P: Strix Point, Strix Halo, Krackan) and will not load on NPU1 (Phoenix, Hawk Point).

Download

hf download brishen/iron-huihui-qwen3.5-2b-abliterated-npu2 --local-dir qwen3.5-2b-abliterated
# or, from an IRON checkout:
python scripts/hf_models.py download qwen3.5-2b-abliterated --repo brishen/iron-huihui-qwen3.5-2b-abliterated-npu2 --out qwen3.5-2b-abliterated

Qwen/Qwen3.5-2B with its refusal direction projected out of the weights (huihui-ai's abliteration): the same architecture, tokenizer and chat template, exported exactly as brishen/iron-qwen3.5-2b-npu2 is. The text model and the vision tower (the multi-token prediction head is not exported). Every projection runs on the NPU: flm.GEMMs over the prompt and the image's patches, GEMVbfp16s for each generated token (~10 tokens/s on a Ryzen AI 9 HX 370), and the taconite-qwen35 Rust runtime replays it over XRT or directly over the amdxdna driver. ~2.5 GB.

Usage warnings

Abliteration removes the model's refusal behaviour, so it will answer
requests the original declines and its output is not safety-filtered.
Review what it generates, keep it to research, testing and controlled
settings rather than public-facing products, and follow your local laws;
see the upstream card's
warnings, which apply here unchanged.

Requirements

  • An AMD Ryzen AI NPU2 (Strix Point, Strix Halo, Krackan) on Linux
    with the amdxdna driver and its firmware (/dev/accel/accel0), and
    an unlimited locked-memory limit (ulimit -l unlimited): the weights
    live in ~2.5 GB of NPU-visible buffers.
  • XRT, unless the runtime is built with --features direct (straight to
    the driver's ioctls, no XRT at all).
  • The model uses 3 of the NPU's 16 hardware-context slots (shared by all
    processes). When other programs hold the rest, the runtime swaps its
    contexts rather than failing (--timing reports the swaps).

Usage

With the taconite-qwen35
runtime (docs):

cargo install taconite-qwen35   # over XRT
# or, with no XRT at all:
cargo install taconite-qwen35 --no-default-features --features cli,direct

qwen35 qwen3.5-2b-abliterated --prompt "Why is the sky blue?"                       # streams the answer
qwen35 qwen3.5-2b-abliterated --image photo.jpg --prompt "Describe this image."    # --image repeats
qwen35 qwen3.5-2b-abliterated --thinking --max-new 1024 --prompt "..."             # reason in <think> first
qwen35 qwen3.5-2b-abliterated --interactive                                        # one prompt a line
qwen35 check qwen3.5-2b-abliterated                                                # verify this bundle

As a library:

use taconite_qwen35::{ChatOptions, Qwen35, RgbImage};

let mut q = Qwen35::load(std::path::Path::new("qwen3.5-2b-abliterated"), 8192)?;
let opts = ChatOptions { images: vec![RgbImage { width, height, rgb }], ..Default::default() };
let (text, stats) = q.chat("Describe this image.", &opts, |s| print!("{s}"))?;

How it runs

Every weight matrix runs on the NPU; the host does the rest:

NPU host (f32)
the prompt every projection as an flm.GEMM over 256-row chunks, one hardware context tokenizer, embedding, norms, the Gated DeltaNet's causal conv and recurrence, partial M-RoPE, GQA attention
each generated token every projection and the tied LM head as a GEMVbfp16, a second context the same, one row; greedy sampling
an image the 24-block vision tower's projections as flm.GEMMs, a third context resize and patchify (torchvision-exact), LayerNorms, 2D RoPE, attention

Weights are stored once, as bfp16 (8-bit mantissas, one exponent a block
of 8: ~9 bits a weight): the decode GEMVs read the prefill GEMMs' packed
streams, and the input embedding is decoded from the tied LM head's. The
vision tower's weights are int8 per output channel, held exactly in bfp16
blocks with each channel's leftover scale factor applied on the host (a
single bfp16 block of the float weights loses too much over its 24
blocks). Activations enter the multiplies as bfp16 hi + lo (~16 bits).

Accuracy

qwen35 check qwen3.5-2b-abliterated on a Ryzen AI 9 HX 370, against transformers float32
(these weights) on the ROCm iGPU, the next token teacher-forced on the
reference's:

prompt prompt tokens steps top-1 agrees disagreements (the reference's top-2 margin) max |dlogit| (reference's top 16)
"Explain in three sentences why the sky is blue." 23 95 92/95 3, all near-ties (<= 0.20) 0.40
a model-card summary 778 134 132/134 2, both exact ties (0.00) 0.55
a kitchen photo + "Describe this image in detail." 280 160 150/160 10, all near-ties (<= 0.22) 0.86

No disagreement where float32's top two are 0.5 or more apart. The
tokenizer and chat template are Qwen3.5-2B's and reproduce Hugging Face's
ids exactly, the image preprocessing matches the Python reference bit for
bit, and the vision tower's output has cosine 0.992 against float32. The
bundle carries these references, so qwen35 check repeats the comparison
on your machine.

Speed

Ryzen AI 9 HX 370 (NPU2): setup ~2 s; prompt 0.2 s for 23 tokens, 1.3 s
for 778; ~93 ms a generated token (~10 tokens/s), bound by the ~2.1 GB of
weights a step streams from memory; an image of ~1M pixels (1040 patches)
~2-2.5 s through the vision tower (NPU ~0.25 s, host attention ~1.7 s).

Limits

  • Greedy decoding only.
  • A prompt of up to 2048 tokens (image tokens included); context up to
    --max-ctx (default 8192).
  • Images are resized as the checkpoint's processor does, capped at ~1M
    pixels (4096 patches, 1024 image tokens); no video.
  • The checkpoint's multi-token-prediction head is not used.

Files

manifest.txt (the kernels and model constants), tensors.txt /
tensors.bin (packed weights and references), kernels/ (xclbins and
instruction streams), tokenizer.json (the checkpoint's). Exported by
IRON's iron/applications/qwen3_5/export_qwen35.py.

Provenance

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.