← back to catalog · registered 2026-08-22 13:56

KostkaIT/Qwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX

KostkaIT Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/KostkaIT%2FQwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 2,984
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
1K last 30d - stable
Likes
6
Model age
7w ago
created 2026-08-21
Downloads over time
Now3.5K→from154↑2,144%
01.3K2.5K3.8K154 on Aug 193.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 omlx apple-silicon qwen qwen3 qwen3.8 vision mtp native-mtp quantized

Related

Total size
22.1 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-24 11:24

Files by quantization

Auxiliary files 17 files 22.1 GB
model-00001-of-00005.safetensors 4.71 GB 5902b860 download
model-00003-of-00005.safetensors 4.70 GB 6d11bd8f download
model-00004-of-00005.safetensors 4.70 GB 61898a7a download
model-00002-of-00005.safetensors 4.67 GB 43d25fb7 download
model-00005-of-00005.safetensors 3.31 GB a559ecc1 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 207 KB 40de72a8 download
tokenizer_config.json 17.5 KB 5de744b3 download
config.json 12.1 KB 6409a817 download
LICENSE 11.1 KB 8e512de1 download
README.md 8.88 KB 25bebca1 download
chat_template.jinja 8.74 KB c0c686f9 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
    library_name: mlx
    pipeline_tag: image-text-to-text
    tags:
  • mlx
  • omlx
  • apple-silicon
  • qwen
  • qwen3
  • qwen3_5
  • qwen3.8
  • vision
  • mtp
  • native-mtp
  • imatrix
  • safetensors
  • abliterated
  • oQ6e

Huihui Qwen3.8-27B Abliterated — oQ6e MLX with Native MTP

A quality-focused MLX/oMLX quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated.

This release was prepared for local inference on Apple Silicon. The main goals
were quantization quality, preservation of native MTP, and clean integration
with oMLX.

Maintained by KostkaIT, an independent technology project by Łukasz
Frąckowiak (vsnake87).

Release summary

  • MLX safetensors format
  • oQ6e quantization
  • 6-bit affine quantization with group_size: 64
  • Selected sensitive tensors preserved at 8-bit
  • Importance-matrix-assisted quantization
  • Native Qwen3.8 MTP weights preserved
  • Native MTP detected by oMLX
  • MTP exposed directly in oMLX model-level Advanced settings
  • MTP can be enabled or disabled without manually editing
    model_settings.json
  • Vision and processor files included in this release
  • Quality-oriented design rather than maximum compression or raw speed

This is an independent community conversion. It is not an official Qwen or
Huihui release.

Model lineage

Stage Model
Official base model Qwen/Qwen3.8-27B
Abliterated derivative huihui-ai/Huihui-Qwen3.8-27B-abliterated
This release oQ6e MLX quantization with native MTP preserved

The upstream Huihui model is an abliterated derivative of Qwen3.8-27B.
According to the upstream model card, the first 15 layers were retained without
abliteration and the MTP and visual components were not modified in that
release.

The source revision used for this conversion was:

d42ca8978c5a66e92c3446d46e8adfe03ef692ff

This repository contains a quantized format conversion. It is not presented as
a new fine-tune or as an official continuation of Qwen or Huihui.

Quantization approach

This release uses oQ6e quantization created with an importance matrix
(imatrix).

The conversion was designed to prioritize:

  • preservation of model quality;
  • stability of the original behavior;
  • retention of sensitive model information;
  • native MTP compatibility;
  • predictable operation in oMLX.

The target was not the smallest possible model or the highest possible decode
speed.

The oQ6e name is the release label used for this conversion. It is not an
official Qwen quantization name. The exact tensor layout and quantization
metadata included in this repository are authoritative.

Quantization can change the output distribution compared with the original BF16
model. This release should not be considered bit-identical or lossless.

The local imatrix calibration cache and generation report are intentionally not
included. They are not required for inference and may contain machine-specific
paths and private calibration metadata.

Native MTP and oMLX integration

Native MTP is a central feature of this release.

The checkpoint contains native MTP tensors under the
language_model.mtp.* namespace and declares one MTP hidden layer in its model
configuration.

With a compatible oMLX version, the expected workflow is:

  1. Load the model in oMLX.
  2. Open the model's Advanced settings.
  3. Enable or disable MTP directly from the model settings.
  4. Reload the model if requested by oMLX.
  5. Confirm the selected MTP state in the runtime log.

Manual editing of model_settings.json should not be required for normal use.

This is different from a model that only contains:

{
  "mtp_enabled": true
}

A configuration flag alone does not create or restore missing MTP weights. This
release is intended to expose native MTP because the required checkpoint data is
present and recognized by oMLX.

Native MTP is different from:

  • a configuration flag without MTP tensors;
  • external speculative decoding using a separate draft model;
  • the separate VLM-MTP path;
  • generic speculative decoding implemented by another runtime.

A runtime may load the model while ignoring native MTP. Successful loading alone
is therefore not sufficient evidence that MTP is active.

Runtime target

This release primarily targets:

  • oMLX on Apple Silicon;
  • local or single-user inference;
  • coding and reasoning workloads;
  • research and experimentation;
  • users who prefer quality-oriented MLX quantization;
  • users who want per-model MTP control without manual configuration edits.

Other MLX-compatible runtimes may load the model files, but native MTP support
depends on the runtime.

Download

pip install -U huggingface_hub

hf download KostkaIT/Qwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX \
  --local-dir ./KostkaIT/Qwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX

Keep the complete directory structure intact, including the model shards,
configuration, tokenizer, chat template, safetensors index, and processor files.

Using with oMLX

Place the complete model directory in the location used by oMLX and load it
through the normal oMLX model-management workflow.

After loading:

  • open the model's Advanced settings;
  • enable or disable native MTP as required;
  • reload the model when requested;
  • inspect the runtime log if you need to verify the active path.

This model should not require manually adding mtp_enabled to
model_settings.json when used with a compatible oMLX version.

Do not enable a separate VLM-MTP or external drafter setting as a substitute for
native MTP.

Chat template and reasoning

Qwen3.8 uses a structured chat template with support for reasoning controls.

The final behavior depends on the runtime and request parameters, including:

  • thinking enabled or disabled;
  • reasoning effort;
  • preservation of previous thinking content;
  • tool calling;
  • structured output;
  • context length;
  • sampling configuration.

Use the tokenizer and chat template shipped with this release. Replacing them
with a generic template may change the model's behavior.

Vision support

The upstream Qwen3.8 model is a vision-language model. This release includes the
vision configuration, processor configuration, and vision tensors from the
source model.

Vision and video support still depend on the target runtime. Text generation and
native MTP do not automatically prove that every image or video workflow works.
Verify multimodal input separately in the runtime you intend to use.

Safety and responsible use

This is an abliterated model derivative. Abliteration changes refusal behavior;
it does not make the model more truthful, more reliable, or inherently safe.

Compared with a safety-aligned model, this model may produce content that the
original model would refuse. It can also generate:

  • factually incorrect information;
  • unsafe or harmful instructions;
  • biased, hateful, sexual, or otherwise inappropriate content;
  • insecure or malicious code;
  • confident medical, legal, or financial advice.

Users and deployers are responsible for applying safeguards appropriate to their
use case.

Do not use this model as the sole decision-maker for:

  • medical or dietary treatment;
  • legal or financial decisions;
  • employment, housing, credit, education, or insurance decisions;
  • autonomous actions affecting people or external systems;
  • content moderation without an additional moderation and review layer.

For public or multi-user deployment, implement authentication, authorization,
input and output moderation, rate limiting, prompt and data isolation, logging,
and human review for high-impact outputs.

Local inference can reduce network exposure, but it does not guarantee privacy.
Operating-system services, runtimes, applications, logs, caches, or monitoring
tools may still retain prompts and generated outputs.

Limitations

  • Quantization may change model quality and output behavior.
  • Native MTP may be unavailable in unsupported runtimes.
  • Vision and video support must be tested separately.
  • Long-context behavior depends on memory and runtime configuration.
  • Tool calling and structured output require dedicated validation.
  • Results from one Apple Silicon device should not be generalized to all devices.
  • MTP is an inference optimization, not a safety mechanism.
  • The model can hallucinate and must not be treated as a source of truth.

License and attribution

The upstream Qwen3.8 and Huihui model repositories are marked as Apache-2.0.
The full license text is included in LICENSE.

Relevant sources:

This repository is an independent community conversion and is not affiliated
with or endorsed by Qwen or Huihui.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-24Update README.mdfdce39215.6 KB
    Loading...
  2. 2026-08-21Update README.mde9694d68.9 KB
    Loading...
  3. 2026-08-21model upload163985a8.8 KB
    Loading...
  4. 2026-08-21initial commit8a7216c28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration