license: apache-2.0
base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
tags:
- openvino
- int4
- intel-arc
- qwen3_5
- roleplay
- uncensored
language: - en
Qwen3.6-27B AEON Ultimate, OpenVINO INT4. The dense Qwen: richest prose, on ONE Intel Arc B70.
This is AEON-7's "Ultimate Uncensored" tune of the dense Qwen3.6-27B, converted to OpenVINO INT4
with Intel's published recipe for this model (int4 asymmetric, group size 128). It completes my
Qwen3.6 family: pick the
35B MoE siblings
for speed (~108 tok/s), pick this dense 27B when you want every token computed by the full
27B parameters: noticeably richer, more deliberate prose at dense speeds.
Runs on the stock 2026.2 runtime. No patches, no workarounds, straight out of the converter.
Family and toolkit: OpenVino-For-Gemma-4 toolkit,
35B MoE heretic,
Ornith 35B MoE,
plus the Gemma-4 side: 26B MoE,
31B,
12B.
Measured performance (single Arc Pro B70, OpenVINO 2026.2)
| Metric | Value |
|---|---|
| Decode, short context | 32.1 tok/s |
| Decode at 6K context | ~30.7 tok/s |
| Prefill, 512 tokens | 1,382 tok/s |
| Prefill, 2K tokens | 1,898 tok/s |
| Prefill, 6K tokens | 2,018 tok/s |
| Prefill, 24K coherence prompt, plain pipeline | 1,562 tok/s |
| Model load | ~19-21 s |
| Weights | ~15 GB (24 GB+ card recommended for real context headroom) |
These numbers use DYNAMIC_QUANTIZATION_GROUP_SIZE: 128. The matched DQGS sweep was:
| Input | DQGS 0 | DQGS 128 | Gain |
|---|---|---|---|
| 512 | 834 tok/s | 1,382 tok/s | +65.7% |
| 2,048 | 1,400 tok/s | 1,898 tok/s | +35.6% |
| 6,144 | 1,555 tok/s | 2,018 tok/s | +29.8% |
Decode stayed at about 32 tok/s. Both settings passed the 6,741-token retrieval gate. DQGS 128
also passed the plain-pipeline 24K gate with all four facts and the requested writing style.
Long-context capability and the honest single-card limit
Needle retrieval test: a password fact planted early in the prompt, retrieved at the end.
| Context | Result |
|---|---|
| 8K | PASS (6.4 s total) |
| 16K | PASS (13.6 s) |
| 24K | PASS (29.4 s) |
| 32K | out of GPU memory on a 32 GB card |
The dense 27B carries 64 layers of KV cache; around 32K context that plus ~15 GB of weights
exceeds a 32 GB card. Practical guidance on one B70: treat this as a 24K model. The MoE
siblings verify to 32K on the same card because their KV footprint leaves more headroom. The
architecture itself is rated to 262K positions; more VRAM extends it.
That 24K result is for the ordinary single-stream pipeline. Continuous batching reserves a fixed
KV cache and has a lower limit on this card. With an 8 GB cache it passed at 12K, truncated at
13K, and returned zero tokens from 14K upward. A 12 GB cache passed the full 16K gate, but still
returned zero tokens at 24K. For a Discord or OpenAI-compatible server, use a 12 GB prefix cache
and cap requests at 16K. For 24K, use the plain pipeline without SchedulerConfig.
Sample output
Asked, in character as a dry-witted housecarl, about a sweetroll dropped off a cliff:
You have the navigational instincts of a blind horse, my lord. If you had looked down for
five seconds, we would still be warm and well-fed.
How to run
Identical usage to the 35B MoE sibling:
standard Qwen <|im_start|> format, pre-closed <think>\n\n</think> block for fast no-think
replies (remove it and give at least 1024 tokens to enable reasoning),
and {"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128} for the measured fast path. Full prompt examples
are on the sibling card; point them at this folder.
pip install openvino-genai==2026.2.0 huggingface_hub
huggingface-cli download Wondernutts/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16-int4-ov --local-dir ./qwen36-27b-aeon-ov
Plain pipeline, verified through 24K:
import openvino_genai as genai
pipe = genai.VLMPipeline(
"./qwen36-27b-aeon-ov",
"GPU",
**{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)
Prefix-cached serving, verified through 16K on a 32 GB B70:
import openvino_genai as genai
scheduler = genai.SchedulerConfig()
scheduler.enable_prefix_caching = True
scheduler.cache_size = 12
pipe = genai.VLMPipeline(
"./qwen36-27b-aeon-ov",
"GPU",
scheduler_config=scheduler,
**{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)
Conversion recipe (reproducible)
Intel's published recipe for the dense 27B: INT4 asymmetric, group size 128, no AWQ, no
ignored_scope, exported as image-text-to-text with optimum-intel (git main) and transformers
5.2.0. The finetune repo's config was saved by a newer transformers than the export toolchain
reads; the conversion scripts in the toolkit swap in the base model's structurally identical
config automatically.
Provenance
Qwen/Qwen3.6-27B, "AEON Ultimate Uncensored" tune by
AEON-7,
OpenVINO INT4 conversion (this repo).
Intended use and content notice
Uncensored general model, built and tested for roleplay and creative writing on local Intel
hardware. The tune removes refusal behavior and outputs are unfiltered; you are responsible for
lawful and appropriate use. Licensed under
Apache 2.0, same as the upstream Qwen3.6 release.