← back to catalog · registered 2026-08-22 13:56

Wondernutts/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16-int4-ov

Wondernutts Qwen 27B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Wondernutts%2FQwen3.6-27B-AEON-Ultimate-Uncensored-BF16-int4-ov"
Response includes
  • classification m-uncensored
  • files 22
  • hub_downloads_all_time 159
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
159
19 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-07-04
Downloads over time
Now163→from92↑77%
8811614317092 on Jul 15163 on Oct 11163 on Oct 8JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
openvino qwen3_5 int4 intel-arc roleplay uncensored en base_model:AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 base_model:finetune:AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 license:apache-2.0 region:us

Related

Total size
14.6 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-26 08:47

Files by quantization

Auxiliary files 22 files 14.6 GB
openvino_language_model.bin 13.0 GB 35474cce download
openvino_text_embeddings_model.bin 1.18 GB f02bb129 download
openvino_vision_embeddings_merger_model.bin 436 MB b3986801 download
openvino_tokenizer.bin 9.18 MB b67896f9 download
openvino_detokenizer.bin 3.65 MB 48fc2c6a download
openvino_vision_embeddings_pos_model.bin 2.54 MB 736234ab download
openvino_vision_embeddings_model.bin 1.69 MB 7af806e6 download
tokenizer.json 19.1 MB 639e352c download
openvino_language_model.xml 9.23 MB f7194c3d download
openvino_vision_embeddings_merger_model.xml 1.26 MB 378c1629 download
openvino_tokenizer.xml 31.9 KB 8543f29d download
openvino_detokenizer.xml 15.2 KB cdd52b34 download
openvino_vision_embeddings_model.xml 8.44 KB c87abd48 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 5.83 KB 3c577399 download
openvino_text_embeddings_model.xml 5.77 KB e9f29c4d download
openvino_vision_embeddings_pos_model.xml 5.74 KB 5f4a702b download
config.json 3.59 KB cc5e5a59 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.20 KB 7cd4b692 download
openvino_config.json 1.18 KB 21b49df6 download
generation_config.json 213 B a0d4001b download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
tags:

  • openvino
  • int4
  • intel-arc
  • qwen3_5
  • roleplay
  • uncensored
    language:
  • en

Qwen3.6-27B AEON Ultimate, OpenVINO INT4. The dense Qwen: richest prose, on ONE Intel Arc B70.

This is AEON-7's "Ultimate Uncensored" tune of the dense Qwen3.6-27B, converted to OpenVINO INT4
with Intel's published recipe for this model (int4 asymmetric, group size 128). It completes my
Qwen3.6 family: pick the
35B MoE siblings
for speed (~108 tok/s), pick this dense 27B when you want every token computed by the full
27B parameters: noticeably richer, more deliberate prose at dense speeds.

Runs on the stock 2026.2 runtime. No patches, no workarounds, straight out of the converter.

Family and toolkit: OpenVino-For-Gemma-4 toolkit,
35B MoE heretic,
Ornith 35B MoE,
plus the Gemma-4 side: 26B MoE,
31B,
12B.

Measured performance (single Arc Pro B70, OpenVINO 2026.2)

Metric Value
Decode, short context 32.1 tok/s
Decode at 6K context ~30.7 tok/s
Prefill, 512 tokens 1,382 tok/s
Prefill, 2K tokens 1,898 tok/s
Prefill, 6K tokens 2,018 tok/s
Prefill, 24K coherence prompt, plain pipeline 1,562 tok/s
Model load ~19-21 s
Weights ~15 GB (24 GB+ card recommended for real context headroom)

These numbers use DYNAMIC_QUANTIZATION_GROUP_SIZE: 128. The matched DQGS sweep was:

Input DQGS 0 DQGS 128 Gain
512 834 tok/s 1,382 tok/s +65.7%
2,048 1,400 tok/s 1,898 tok/s +35.6%
6,144 1,555 tok/s 2,018 tok/s +29.8%

Decode stayed at about 32 tok/s. Both settings passed the 6,741-token retrieval gate. DQGS 128
also passed the plain-pipeline 24K gate with all four facts and the requested writing style.

Long-context capability and the honest single-card limit

Needle retrieval test: a password fact planted early in the prompt, retrieved at the end.

Context Result
8K PASS (6.4 s total)
16K PASS (13.6 s)
24K PASS (29.4 s)
32K out of GPU memory on a 32 GB card

The dense 27B carries 64 layers of KV cache; around 32K context that plus ~15 GB of weights
exceeds a 32 GB card. Practical guidance on one B70: treat this as a 24K model. The MoE
siblings verify to 32K on the same card because their KV footprint leaves more headroom. The
architecture itself is rated to 262K positions; more VRAM extends it.

That 24K result is for the ordinary single-stream pipeline. Continuous batching reserves a fixed
KV cache and has a lower limit on this card. With an 8 GB cache it passed at 12K, truncated at
13K, and returned zero tokens from 14K upward. A 12 GB cache passed the full 16K gate, but still
returned zero tokens at 24K. For a Discord or OpenAI-compatible server, use a 12 GB prefix cache
and cap requests at 16K. For 24K, use the plain pipeline without SchedulerConfig.

Sample output

Asked, in character as a dry-witted housecarl, about a sweetroll dropped off a cliff:

You have the navigational instincts of a blind horse, my lord. If you had looked down for
five seconds, we would still be warm and well-fed.

How to run

Identical usage to the 35B MoE sibling:
standard Qwen <|im_start|> format, pre-closed <think>\n\n</think> block for fast no-think
replies (remove it and give at least 1024 tokens to enable reasoning),
and {"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128} for the measured fast path. Full prompt examples
are on the sibling card; point them at this folder.

pip install openvino-genai==2026.2.0 huggingface_hub
huggingface-cli download Wondernutts/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16-int4-ov --local-dir ./qwen36-27b-aeon-ov

Plain pipeline, verified through 24K:

import openvino_genai as genai

pipe = genai.VLMPipeline(
    "./qwen36-27b-aeon-ov",
    "GPU",
    **{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)

Prefix-cached serving, verified through 16K on a 32 GB B70:

import openvino_genai as genai

scheduler = genai.SchedulerConfig()
scheduler.enable_prefix_caching = True
scheduler.cache_size = 12

pipe = genai.VLMPipeline(
    "./qwen36-27b-aeon-ov",
    "GPU",
    scheduler_config=scheduler,
    **{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)

Conversion recipe (reproducible)

Intel's published recipe for the dense 27B: INT4 asymmetric, group size 128, no AWQ, no
ignored_scope, exported as image-text-to-text with optimum-intel (git main) and transformers
5.2.0. The finetune repo's config was saved by a newer transformers than the export toolchain
reads; the conversion scripts in the toolkit swap in the base model's structurally identical
config automatically.

Provenance

Qwen/Qwen3.6-27B, "AEON Ultimate Uncensored" tune by
AEON-7,
OpenVINO INT4 conversion (this repo).

Intended use and content notice

Uncensored general model, built and tested for roleplay and creative writing on local Intel
hardware. The tune removes refusal behavior and outputs are unfiltered; you are responsible for
lawful and appropriate use. Licensed under
Apache 2.0, same as the upstream Qwen3.6 release.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-26Document Arc B70 DQGS=128 benchmarks7c352595.8 KB
    Loading...
  2. 2026-07-06Upload README.md with huggingface_hubc3d1f244.4 KB
    Loading...
  3. 2026-07-06Upload README.md with huggingface_hubd78f7a24.3 KB
    Loading...
  4. 2026-07-04Upload README.md with huggingface_hub981baf84.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration