← back to catalog · registered 2026-08-22 13:56

darkmaniac7/Huihui-Qwen3.6-27B-abliterated-MNN

darkmaniac7 Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/darkmaniac7%2FHuihui-Qwen3.6-27B-abliterated-MNN"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 95
  • author_summary 21 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
95
27 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-04-30
Downloads over time
Now115→from20↑475%
15528812520 on Apr 29115 on Oct 11115 on Oct 10AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
BF16
Tags
chat qwen qwen3.6 qwen3_5 image-text-to-text vlm multimodal abliterated uncensored mnn tokforge en

Related

Total size
2.37 GB
Files
13
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-10-05 08:42

Files by quantization

BF16 1 file 2.37 GB
embeddings_bf16.bin 2.37 GB 8e363559 download
Auxiliary files 12 files 15.2 GB
llm.mnn.weight 14.9 GB c1c08714 download
visual.mnn.weight 249 MB 8f2a00ce download
llm.mnn.json 21.4 MB 732d7018 download
llm.mnn 8.46 MB c9457f0c download
tokenizer.txt 6.17 MB 30439bef download
visual.mnn 805 KB fce22af1 download
tokforge_benchmarks.json 19.6 KB 7184f90a download
llm_config.json 8.46 KB e48c551b download
README.md 7.90 KB 715285bc download
.gitattributes 1.72 KB e23c58dd download
export_args.json 1.09 KB 4a5c9bfe download
config.json 387 B 2d94fea9 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    pipeline_tag: image-text-to-text
    tags:
  • chat
  • qwen
  • qwen3.6
  • qwen3_5
  • image-text-to-text
  • vlm
  • multimodal
  • abliterated
  • uncensored
  • mnn
  • tokforge
    base_model:
  • huihui-ai/Huihui-Qwen3.6-27B-abliterated
  • Qwen/Qwen3.6-27B

TokForge

Runs on-device in the TokForge app.

Huihui-Qwen3.6-27B-abliterated-MNN

MNN-format 4-bit TokForge export of huihui-ai/Huihui-Qwen3.6-27B-abliterated, packaged for Android on-device inference with TokForge.

What this is

The bundle includes the visual MNN sidecar, but the validation below is text-generation only through TokForge's /test-prompt endpoint.

Bundle contents

  • config.json - TokForge/MNN runtime defaults.
  • llm_config.json - model geometry and tokenizer/chat settings.
  • llm.mnn / llm.mnn.weight - quantized LLM graph and external weights.
  • embeddings_bf16.bin - separated BF16 embedding table.
  • tokenizer.txt - tokenizer exported by the MNN pipeline.
  • visual.mnn / visual.mnn.weight - Qwen3.6 vision sidecar.
  • export_args.json - conversion settings.
  • llm.mnn.json - exported graph metadata.
  • tokforge_benchmarks.json - raw benchmark summary from the device validation run.

Quantization scheme

Flag Value
--quant_bit 4
--quant_block 64
--lm_quant_bit 4
--lm_quant_block 64
--embed_bit 16
--hqq enabled
--seperate_embed enabled

Runtime defaults in config.json: CPU backend, 4 threads, low precision, low memory, separated BF16 embeddings.

TokForge device validation

Validated on April 30 and May 1, 2026 over wireless ADB with TokForge's authenticated /test-prompt control endpoint. The prompts were intentionally independent rather than sequential chat turns: arithmetic, short sentence, paragraph generation, and a small Python function.

Generation config for every prompt:

Setting Value
Backend mnn
Context 8192
Threads 4
Precision low
thinking_enabled false
temperature 0.0
top_p 1.0
top_k 1
seed 42
Path reported by TokForge full_history
QNN/NPU not requested, not attempted, CPU effective

Summary

Device Model SoC RAM Load time Avg decode tok/s Weighted decode tok/s Avg prefill Completed Coherent completed Timeouts
RedMagic NX809J SM8850 24 GB ~31-33 s 2.49 2.25 12.06 s 4/4 4/4 -
Redmi 2407FRK8EC MT6989 24 GB 55 s 2.63 2.33 25.66 s 4/4 4/4 -
OnePlus PLC110 MT6991 16 GB 51.80 s 0.28 0.24 25.19 s 3/4 3/3 paragraph
Lenovo TB520FU SM8650 16 GB 62.90 s 0.30 0.27 20.07 s 3/4 3/3 paragraph
S26 SM-S948U1 SM8850 16 GB 29.47 s 0.25 0.22 17.93 s 2/4 2/2 paragraph, code
Pixel Pixel 9 Pro XL Tensor G4 16 GB 76.55 s 0.12 0.11 33.08 s 2/4 2/2 paragraph, code

Weighted decode tok/s is computed as total generated decode tokens divided by total decode wall time. For devices with timeouts, decode averages are computed across completed prompts only.

The 24 GB devices completed the full four-prompt matrix and are practical for short local text generation with this export. The 16 GB devices can load the model, but longer paragraph/code prompts are not practical in this CPU MNN configuration: OnePlus and Lenovo completed short/code prompts but timed out on paragraph generation; S26 and Pixel completed only arithmetic and short sentence before timing out on longer prompts.

Per-prompt decode results

Device Prompt class Prompt tokens Decode tokens Decode tok/s Result
RedMagic arithmetic 73 3 3.07 12
RedMagic short sentence 72 14 2.58 coherent
RedMagic paragraph 83 122 2.24 coherent 5-sentence paragraph
RedMagic Python code 87 20 2.06 valid clamp function
Redmi arithmetic 92 3 3.42 12
Redmi short sentence 91 14 2.43 coherent
Redmi paragraph 102 103 2.29 coherent 5-sentence paragraph
Redmi Python code 106 20 2.37 valid clamp function
OnePlus arithmetic 92 3 0.36 12
OnePlus short sentence 91 14 0.25 coherent
OnePlus paragraph - - timeout no final response captured; server history showed partial decode in the 56-60 token range
OnePlus Python code 106 20 0.23 valid clamp function
Lenovo arithmetic 92 3 0.37 12
Lenovo short sentence 91 14 0.27 coherent
Lenovo paragraph - - timeout HTTP 504 after roughly 300 s; server generated 71 tokens before timeout
Lenovo Python code 106 20 0.26 valid clamp function
S26 arithmetic 92 3 0.29 12
S26 short sentence 91 14 0.20 coherent
S26 paragraph - - timeout HTTP 504 / no final response
S26 Python code - - timeout HTTP 504 / no final response
Pixel arithmetic 92 3 0.14 12
Pixel short sentence 91 14 0.10 coherent
Pixel paragraph - - timeout HTTP 504 / no final response
Pixel Python code - - timeout HTTP 504 / no final response

The paragraph and code prompts were added to reduce benchmark bias from tiny outputs. Completed code prompts produced the expected Python function:

def clamp(value, low, high):
    return max(low, min(value, high))

Download

pip install -U huggingface_hub
hf download darkmaniac7/Huihui-Qwen3.6-27B-abliterated-MNN --local-dir Huihui-Qwen3.6-27B-abliterated-MNN

Usage with upstream MNN llm_demo

git clone https://github.com/alibaba/MNN.git
cd MNN
mkdir build && cd build
cmake .. -DMNN_LOW_MEMORY=true -DMNN_CPU_WEIGHT_DEQUANT_GEMM=true \
         -DMNN_BUILD_LLM=true -DMNN_SUPPORT_TRANSFORMER_FUSE=true
make -j

./llm_demo /path/to/Huihui-Qwen3.6-27B-abliterated-MNN/config.json prompt.txt

Host smoke validation with the TokForge MNN build loaded successfully, detected 48 linear-attention state layers out of 64 total layers, and emitted the expected one-token OK response for a no-thinking prompt.

Attribution

License and safety

Apache-2.0, inherited from the upstream Qwen and huihui-ai model cards. This is a safety-reduced / uncensored model; deploy with appropriate product policy, user controls, and local-law awareness.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-05Card metadata: license, base_model, base_model_relationfbc6f497.9 KB
    Loading...
  2. 2026-07-04Add TokForge app linkscb113037.9 KB
    Loading...
  3. 2026-05-02Update TokForge device benchmark matrix788eb207.6 KB
    Loading...
  4. 2026-04-30Add files using upload-large-folder tool6dc29e35.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration