license: apache-2.0
base_model:
- huihui-ai/Huihui-Qwen3.8-27B-abliterated
library_name: mlx
pipeline_tag: image-text-to-text
tags: - mlx
- omlx
- apple-silicon
- qwen
- qwen3
- qwen3_5
- qwen3.8
- vision
- mtp
- native-mtp
- imatrix
- safetensors
- abliterated
- oQ6e
Huihui Qwen3.8-27B Abliterated — oQ6e MLX with Native MTP
A quality-focused MLX/oMLX quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated.
This release was prepared for local inference on Apple Silicon. The main goals
were quantization quality, preservation of native MTP, and clean integration
with oMLX.
Maintained by KostkaIT, an independent technology project by Łukasz
Frąckowiak (vsnake87).
Release summary
- MLX safetensors format
oQ6equantization- 6-bit affine quantization with
group_size: 64 - Selected sensitive tensors preserved at 8-bit
- Importance-matrix-assisted quantization
- Native Qwen3.8 MTP weights preserved
- Native MTP detected by oMLX
- MTP exposed directly in oMLX model-level Advanced settings
- MTP can be enabled or disabled without manually editing
model_settings.json - Vision and processor files included in this release
- Quality-oriented design rather than maximum compression or raw speed
This is an independent community conversion. It is not an official Qwen or
Huihui release.
Model lineage
| Stage | Model |
|---|---|
| Official base model | Qwen/Qwen3.8-27B |
| Abliterated derivative | huihui-ai/Huihui-Qwen3.8-27B-abliterated |
| This release | oQ6e MLX quantization with native MTP preserved |
The upstream Huihui model is an abliterated derivative of Qwen3.8-27B.
According to the upstream model card, the first 15 layers were retained without
abliteration and the MTP and visual components were not modified in that
release.
The source revision used for this conversion was:
d42ca8978c5a66e92c3446d46e8adfe03ef692ff
This repository contains a quantized format conversion. It is not presented as
a new fine-tune or as an official continuation of Qwen or Huihui.
Quantization approach
This release uses oQ6e quantization created with an importance matrix
(imatrix).
The conversion was designed to prioritize:
- preservation of model quality;
- stability of the original behavior;
- retention of sensitive model information;
- native MTP compatibility;
- predictable operation in oMLX.
The target was not the smallest possible model or the highest possible decode
speed.
The oQ6e name is the release label used for this conversion. It is not an
official Qwen quantization name. The exact tensor layout and quantization
metadata included in this repository are authoritative.
Quantization can change the output distribution compared with the original BF16
model. This release should not be considered bit-identical or lossless.
The local imatrix calibration cache and generation report are intentionally not
included. They are not required for inference and may contain machine-specific
paths and private calibration metadata.
Native MTP and oMLX integration
Native MTP is a central feature of this release.
The checkpoint contains native MTP tensors under thelanguage_model.mtp.* namespace and declares one MTP hidden layer in its model
configuration.
With a compatible oMLX version, the expected workflow is:
- Load the model in oMLX.
- Open the model's Advanced settings.
- Enable or disable MTP directly from the model settings.
- Reload the model if requested by oMLX.
- Confirm the selected MTP state in the runtime log.
Manual editing of model_settings.json should not be required for normal use.
This is different from a model that only contains:
{
"mtp_enabled": true
}
A configuration flag alone does not create or restore missing MTP weights. This
release is intended to expose native MTP because the required checkpoint data is
present and recognized by oMLX.
Native MTP is different from:
- a configuration flag without MTP tensors;
- external speculative decoding using a separate draft model;
- the separate VLM-MTP path;
- generic speculative decoding implemented by another runtime.
A runtime may load the model while ignoring native MTP. Successful loading alone
is therefore not sufficient evidence that MTP is active.
Runtime target
This release primarily targets:
- oMLX on Apple Silicon;
- local or single-user inference;
- coding and reasoning workloads;
- research and experimentation;
- users who prefer quality-oriented MLX quantization;
- users who want per-model MTP control without manual configuration edits.
Other MLX-compatible runtimes may load the model files, but native MTP support
depends on the runtime.
Download
pip install -U huggingface_hub
hf download KostkaIT/Qwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX \
--local-dir ./KostkaIT/Qwen3.8-27B-Huihui-Abliterated-oQ6e-MTP-MLX
Keep the complete directory structure intact, including the model shards,
configuration, tokenizer, chat template, safetensors index, and processor files.
Using with oMLX
Place the complete model directory in the location used by oMLX and load it
through the normal oMLX model-management workflow.
After loading:
- open the model's Advanced settings;
- enable or disable native MTP as required;
- reload the model when requested;
- inspect the runtime log if you need to verify the active path.
This model should not require manually adding mtp_enabled tomodel_settings.json when used with a compatible oMLX version.
Do not enable a separate VLM-MTP or external drafter setting as a substitute for
native MTP.
Chat template and reasoning
Qwen3.8 uses a structured chat template with support for reasoning controls.
The final behavior depends on the runtime and request parameters, including:
- thinking enabled or disabled;
- reasoning effort;
- preservation of previous thinking content;
- tool calling;
- structured output;
- context length;
- sampling configuration.
Use the tokenizer and chat template shipped with this release. Replacing them
with a generic template may change the model's behavior.
Vision support
The upstream Qwen3.8 model is a vision-language model. This release includes the
vision configuration, processor configuration, and vision tensors from the
source model.
Vision and video support still depend on the target runtime. Text generation and
native MTP do not automatically prove that every image or video workflow works.
Verify multimodal input separately in the runtime you intend to use.
Safety and responsible use
This is an abliterated model derivative. Abliteration changes refusal behavior;
it does not make the model more truthful, more reliable, or inherently safe.
Compared with a safety-aligned model, this model may produce content that the
original model would refuse. It can also generate:
- factually incorrect information;
- unsafe or harmful instructions;
- biased, hateful, sexual, or otherwise inappropriate content;
- insecure or malicious code;
- confident medical, legal, or financial advice.
Users and deployers are responsible for applying safeguards appropriate to their
use case.
Do not use this model as the sole decision-maker for:
- medical or dietary treatment;
- legal or financial decisions;
- employment, housing, credit, education, or insurance decisions;
- autonomous actions affecting people or external systems;
- content moderation without an additional moderation and review layer.
For public or multi-user deployment, implement authentication, authorization,
input and output moderation, rate limiting, prompt and data isolation, logging,
and human review for high-impact outputs.
Local inference can reduce network exposure, but it does not guarantee privacy.
Operating-system services, runtimes, applications, logs, caches, or monitoring
tools may still retain prompts and generated outputs.
Limitations
- Quantization may change model quality and output behavior.
- Native MTP may be unavailable in unsupported runtimes.
- Vision and video support must be tested separately.
- Long-context behavior depends on memory and runtime configuration.
- Tool calling and structured output require dedicated validation.
- Results from one Apple Silicon device should not be generalized to all devices.
- MTP is an inference optimization, not a safety mechanism.
- The model can hallucinate and must not be treated as a source of truth.
License and attribution
The upstream Qwen3.8 and Huihui model repositories are marked as Apache-2.0.
The full license text is included in LICENSE.
Relevant sources:
This repository is an independent community conversion and is not affiliated
with or endorsed by Qwen or Huihui.