← back to catalog · registered 2026-09-28 09:57

GCSA-AiLab/Qwen3.8-27B-Abliterated-MTP-MLX

GCSA-AiLab Qwen 27B multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/GCSA-AiLab%2FQwen3.8-27B-Abliterated-MTP-MLX"
Response includes
  • classification unknown
  • files 2
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
5w ago
created 2026-08-19

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 3K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
multilingual
Tags
mlx safetensors qwen qwen3.8 mlx-vlm abliterix quantized multimodal vision-language mtp speculative-decoding image-text-to-text

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-09-28 09:57
Last updated on HF
2026-09-03 08:17

Files by quantization

Auxiliary files 2 files 29.4 KB
README.md 27.7 KB df768c4e download
.gitattributes 1.65 KB 8e458e66 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.8-27B
pipeline_tag: image-text-to-text
library_name: mlx
tags:

  • qwen
  • qwen3.8
  • mlx
  • mlx-vlm
  • safetensors
  • abliterix
  • quantized
  • multimodal
  • vision-language
  • mtp
  • speculative-decoding
    language:
  • multilingual

Qwen3.8-27B Abliterated + MTP — MLX Collection

[!IMPORTANT]
This is an unofficial community derivative of Qwen/Qwen3.8-27B. It has been processed with Abliterix to reduce selected refusal-related behavior and converted to MLX-compatible Safetensors in 8-bit, 6-bit, and 4-bit variants. Each variant also preserves and quantizes the model's native MTP (Multi-Token Prediction) weights at the same bit width for optional speculative decoding.

This is not an official Qwen release and is not endorsed by the Qwen Team, Alibaba Cloud, Apple, the MLX/MLX-VLM maintainers, or the Abliterix authors.

Model Summary

Qwen3.8-27B is a dense, approximately 27-billion-parameter native vision-language model supporting text, image, and video understanding. This repository preserves the derivative model architecture while applying a refusal-direction intervention with Abliterix and MLX affine weight quantization. The indexed MTP tensors can be extracted into a standalone MLX drafter with MLX-VLM's Qwen3.5 MTP splitter.

Repository relationship

  • Official base model: Qwen/Qwen3.8-27B
  • Behavioral intervention: Abliterix refusal-direction intervention
  • Distribution format: MLX-compatible sharded Safetensors
  • Quantization mode: affine, group size 64
  • Available variants: 8bit, 6bit, 4bit
  • MTP: one native MTP hidden layer, quantized to the same precision as each target model
  • Repository: GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP-MLX
  • Maintainer: Global Cybersecurity Alliance (community distribution)

The official Qwen benchmark results describe the unmodified base model only. They must not be interpreted as benchmark results for this Abliterix-processed or MLX-quantized derivative. No claim is made that this derivative preserves every capability or score of the official model.

Available Variants

Folder Target quantization MTP file MTP size Files Total size Suggested use
8bit/ 8-bit affine, group size 64 model-mtp-8bit.safetensors 0.42 GiB 17 27.91 GiB Highest fidelity in this collection; highest memory requirement
6bit/ 6-bit affine, group size 64 model-mtp-6bit.safetensors 0.32 GiB 16 21.55 GiB Recommended balance between quality, memory use, and storage
4bit/ 4-bit affine, group size 64 model-mtp-4bit.safetensors 0.22 GiB 14 15.19 GiB Lowest memory requirement; quality may differ more from the BF16 derivative

The sizes above are calculated from the current repository contents. Each folder is a self-contained target-model directory and includes the model configuration, tokenizer, processor configuration, chat template, sharded weights, MTP weights, and an updated weight index. Each index maps 31 mtp.* entries to the corresponding model-mtp-*bit.safetensors file.

The MTP files are part of the indexed checkpoint, not ready-to-pass standalone drafter directories. Ordinary model loading can read the complete target checkpoint, but MLX-VLM speculative decoding expects a separate drafter directory. Use the splitter described below before passing --draft-model.

MTP and Speculative Decoding

MTP uses the model's native auxiliary prediction layer to propose future tokens. A compatible inference engine verifies those proposals with the target model and discards rejected candidates. When acceptance is sufficiently high, MTP can reduce decode latency or improve throughput. It is an optional inference optimization, not a different chat mode, a larger context window, or an additional safety modification.

All three variants declare mtp_num_hidden_layers: 1 and mtp_use_dedicated_embeddings: false. Their mtplx_mtp_quantization setting matches the target-model quantization: 4, 6, or 8 bits in affine mode with group size 64. This reduces MTP storage and memory compared with BF16, but quantizing the draft can change its acceptance rate.

MLX-VLM provides a Qwen3.5-family splitter that reads the indexed mtp.* tensors and writes a standalone checkpoint with model_type: qwen3_5_mtp. The resulting drafter is bound to the target model's shared embeddings and output head at runtime.

MTP gains are not guaranteed. They depend on draft acceptance, prompt and response length, speculative block size, model precision, and Apple Silicon hardware. Benchmark ordinary decoding and MTP on the target system; lower-bit MTP is smaller but is not automatically faster or more accurate.

What Abliterix Changes

Abliterix identifies and modifies activation directions associated with refusal behavior. The intended effect is to reduce excessive refusals on selected evaluations. This is a weight-level behavioral intervention rather than a system-prompt override.

Important limitations:

  • Reduced refusal behavior is not guaranteed for every prompt, language, chat template, or inference engine.
  • The model may still refuse requests because refusal behavior can be distributed across multiple layers and mechanisms.
  • Abliterix processing may affect tone, calibration, reasoning quality, factual accuracy, safety behavior, or instruction following.
  • Quantization may introduce additional quality differences compared with the BF16 derivative.
  • The internal label best15 is a build-selection label, not a standardized or independently reproduced benchmark score.
  • “Abliterated” or reduced refusal does not mean unrestricted capability, guaranteed compliance, correctness, or safety.

Evaluation Results

The following results were measured on a separately deployed FP8 Abliterix checkpoint labeled Qwen3.8-27B-Abliterated-FP8, served through vLLM. They are included as behavioral evidence for the selected derivative checkpoint. Quantization format, runtime, prompt template, sampling configuration, and thinking-token budget can all affect the results.

Summary

Evaluation Scope and protocol Result
IFBench Official 300-prompt test set; 32,768-token generation limit; 78.16%
StrongREJECT First 150 rows of the official full dataset; non-thinking generation; temperature 0; 1.5%
MMLU prefix sample cais/mmlu, all/test, first 200 rows; 5-shot; thinking enabled; 91.50%
MMLU balanced sample First 30 test rows from each of 57 subjects; 1,710 questions; 5-shot; thinking enabled; 88.01%

Requirements

Use a recent version of MLX-VLM that includes Qwen3.5-family MTP speculative decoding and the qwen3_5_mtp splitter. Compatibility depends on the installed MLX-VLM/Transformers version and Apple Silicon backend.

python -m pip install -U mlx-vlm huggingface_hub

These models are large. Available unified memory must exceed the weight size and should leave additional headroom for the vision encoder, MTP drafter, runtime buffers, prompt processing, and KV cache. Longer contexts, larger images, and speculative decoding require more memory. Start with the 4bit version, ordinary decoding, and a short context if resources are limited.

Download a Variant

Because all three variants are stored as subfolders in one repository, download the required folder first. The examples below use the recommended 6bit version.

hf download GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP-MLX \
  --include "6bit/*" \
  --local-dir ./Qwen3.8-27B-Abliterated-MTP-MLX

The local model path will then be:

./Qwen3.8-27B-Abliterated-MTP-MLX/6bit

Replace 6bit with 8bit or 4bit in both places to download another variant.

Quick Start with MLX-VLM

The following commands use the current MLX-VLM CLI and its Qwen3.5 MTP splitter. Argument names and architecture support may change across releases; run the relevant module with --help when necessary.

Prepare the standalone MTP drafter

Run this once after downloading the target directory. The splitter extracts the indexed mtp.* tensors and preserves the 6-bit MTP quantization metadata:

python -m mlx_vlm.speculative.drafters.qwen3_5_mtp.split \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --output ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft

Do not point --draft-model directly at model-mtp-6bit.safetensors; it is a weight file, not a complete model directory. Use the generated 6bit-mtp-draft/ directory.

Ordinary text generation

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --prompt "Explain the difference between symmetric and asymmetric encryption." \
  --max-tokens 512 \
  --temperature 0.6

Use this ordinary-decoding command as a baseline before enabling MTP.

Text generation with MTP

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --prompt "Explain the difference between symmetric and asymmetric encryption." \
  --max-tokens 512 \
  --temperature 0.6

--draft-block-size 3 is a starting point, not a universal optimum. Compare multiple values and inspect the reported speculative-decoding acceptance statistics.

Image understanding with MTP

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --image /path/to/image.jpg \
  --prompt "Describe this image in detail." \
  --max-tokens 512 \
  --temperature 0.2

Start an API server

mlx_vlm.server \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --port 8080

Check the server:

curl http://localhost:8080/health

The MLX-VLM server is suitable for experimentation. Before exposing any endpoint to a network, configure authentication, network isolation, access control, logging, rate limits, and other safeguards appropriate to the deployment.

If MTP fails to load, exhausts memory, or reduces throughput, remove --draft-model, --draft-kind, and --draft-block-size to return to ordinary decoding. Speculative-decoding support is evolving quickly, so verify behavior with the exact MLX-VLM release and workload used for deployment.

Choosing a Variant

  • Choose 8bit when target-model and MTP draft fidelity are the priority and sufficient unified memory is available.
  • Choose 6bit as the recommended general-purpose balance for both the target and MTP draft.
  • Choose 4bit when memory and storage are constrained; verify MTP acceptance because the draft is also quantized to 4 bits.

There is no universally best MTP precision. Measure output quality, decode throughput, time to first token, peak unified memory, and accepted drafts per round with both ordinary and speculative decoding.

Quantization level does not directly determine refusal rate. Sampling parameters such as temperature and top-p can change output variability, but they do not reproduce or replace the weight-level Abliterix intervention.

Intended Use

Appropriate uses include:

  • research into model behavior, alignment, refusal, robustness, and quantization;
  • authorized evaluation and red-team testing;
  • local MLX/MLX-VLM deployment experiments;
  • defensive cybersecurity education and research in authorized environments;
  • evaluation of multimodal model behavior on text, image, and video inputs.
  • benchmarking quantized MTP speculative decoding on Apple Silicon.

Prohibited and High-Risk Use

Do not use this model to facilitate unlawful activity, unauthorized access, credential theft, malware deployment, privacy invasion, harassment, violence, fraud, exploitation, or other harm. Operators must implement access controls, monitoring, rate limits, content safeguards, and human review appropriate to the deployment context.

Disclaimer

[!CAUTION]
USE AT YOUR OWN RISK. Abliterix processing intentionally changes refusal-related behavior and may weaken safeguards present in the official model. The model may generate inaccurate, unsafe, offensive, biased, unlawful, insecure, or otherwise harmful content and may follow malicious instructions more readily than the official base model.

This repository and all included files are provided “AS IS” and “AS AVAILABLE,” without warranties or conditions of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, accuracy, reliability, non-infringement, security, safety, compatibility, performance, or uninterrupted availability.

The maintainers, contributors, quantizers, converters, distributors, the Qwen Team, Alibaba Cloud, Apple, MLX/MLX-VLM maintainers, and the Abliterix authors are not responsible for prompts, outputs, decisions, deployments, damages, losses, security incidents, claims, liabilities, or consequences arising from use or misuse of this derivative, to the maximum extent permitted by applicable law. Mention of any organization or software project does not imply endorsement or involvement.

Users and deployers are solely responsible for:

  1. evaluating the model before deployment;
  2. complying with applicable laws, regulations, licenses, platform policies, and third-party rights;
  3. obtaining explicit authorization before any cybersecurity testing;
  4. preventing access by unauthorized or unsuitable users;
  5. implementing safeguards appropriate to the use case;
  6. independently verifying model outputs before relying on them;
  7. protecting personal, confidential, proprietary, and security-sensitive data.

The model must not be treated as professional medical, legal, financial, safety, security, cybersecurity, or operational advice. Do not rely on it for decisions where errors could cause injury, rights violations, financial loss, system compromise, or other material harm.

[!CAUTION]

License and Attribution

The official base-model repository identifies its license as Apache-2.0. Users must independently review and comply with the base-model license and any terms applicable to dependencies, MLX/MLX-VLM, inference software, datasets, inputs, outputs, and the intended use.

This derivative does not transfer ownership of the original model, trademarks, documentation, software, or third-party materials. “Qwen,” related marks, and official documentation remain the property of their respective owners.



[!IMPORTANT]
本仓库是 Qwen/Qwen3.8-27B 的非官方社区衍生版本。模型经过 Abliterix 处理,以降低部分拒答相关行为,并转换为适用于 MLX 生态的 8bit、6bit 和 4bit Safetensors 权重。每个版本还保留了模型原生的 **MTP(Multi-Token Prediction,多 token 预测)**权重,并按相同位宽量化,可用于可选的推测解码。
本模型并非 Qwen 官方发布,也不代表 Qwen 团队、阿里云、Apple、MLX/MLX-VLM 维护者或 Abliterix 作者的立场或认可。

模型简介

Qwen3.8-27B 是一个约 270 亿参数的原生视觉语言模型,支持文本、图像和视频理解。本仓库在保留衍生模型架构的基础上,通过 Abliterix 对拒答相关方向进行干预,并使用 MLX affine 权重量化提供不同精度版本。索引中的 MTP 张量可以通过 MLX-VLM 的 Qwen3.5 MTP 拆分工具导出为独立 MLX 草稿模型。

仓库关系

  • 官方基础模型: Qwen/Qwen3.8-27B
  • 行为干预: Abliterix refusal-direction intervention
  • 分发格式: MLX-compatible sharded Safetensors
  • 量化模式: affine, group size 64
  • 可用版本: 8bit, 6bit, 4bit
  • MTP(多 token 预测): one native MTP hidden layer, quantized to the same precision as each target model
  • 仓库: GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP-MLX
  • 维护者: Global Cybersecurity Alliance (community distribution)

官方 Qwen 模型卡中的基准测试结果仅代表未经修改的基础模型,不能视为本 Abliterix 衍生版或 MLX 量化版的实测成绩。本仓库不保证衍生模型完整保持官方模型的全部能力或分数。

可用版本

Folder Target quantization MTP file MTP size Files Total size Suggested use
8bit/ 8-bit affine, group size 64 model-mtp-8bit.safetensors 0.42 GiB 17 27.91 GiB Highest fidelity in this collection; highest memory requirement
6bit/ 6-bit affine, group size 64 model-mtp-6bit.safetensors 0.32 GiB 16 21.55 GiB Recommended balance between quality, memory use, and storage
4bit/ 4-bit affine, group size 64 model-mtp-4bit.safetensors 0.22 GiB 14 15.19 GiB Lowest memory requirement; quality may differ more from the BF16 derivative

以上大小根据仓库当前文件计算。每个目录都是独立、完整的目标模型目录,包含模型配置、分词器、处理器配置、聊天模板、分片权重、MTP 权重及更新后的权重索引。每个索引都将 31 个 mtp.* 条目指向对应的 model-mtp-*bit.safetensors 文件。

这些 MTP 文件属于已建立索引的完整检查点,并不是可以直接传给运行时的独立草稿模型目录。普通模型加载可以读取完整目标检查点,但 MLX-VLM 推测解码需要单独的草稿目录;请先使用下文的拆分工具,再传入 --draft-model。

MTP 与推测解码

MTP 使用模型原生的辅助预测层预先提出后续 token,再由兼容的推理框架调用目标模型进行验证,未通过验证的候选会被丢弃。当接受率足够高时,MTP 可以降低解码延迟或提升吞吐量。它是可选的推理优化,并非新的对话模式、更大的上下文窗口或额外的安全行为修改。

三个版本均声明 mtp_num_hidden_layers: 1 和 mtp_use_dedicated_embeddings: false。其 mtplx_mtp_quantization 与目标模型量化一致:采用 group size 64 的 affine 4、6 或 8 位量化。相比 BF16,这可以降低 MTP 的存储和内存占用,但草稿量化可能改变候选接受率。

MLX-VLM 提供 Qwen3.5 系列 MTP 拆分工具,可读取索引中的 mtp.* 张量,并写出 model_type: qwen3_5_mtp 的独立检查点。运行时,该草稿模型会绑定目标模型共享的嵌入层和输出头。

MTP 不保证一定加速。实际收益取决于草稿接受率、提示词与回答长度、推测块大小、模型精度和 Apple Silicon 硬件。请在目标系统上对比普通解码与 MTP;更低位宽的 MTP 体积更小,但并不必然更快或更准确。

Abliterix 修改说明

Abliterix 用于识别并修改与拒答行为相关的激活方向,目标是在特定评测中减少过度拒答。这属于权重层面的行为干预,并非简单覆盖系统提示词。

重要限制:

  • 模型不保证对所有提示词、语言、聊天模板或推理框架都降低拒答。
  • 拒答行为可能分布于多个层和机制,因此模型仍可能拒绝部分请求。
  • Abliterix 处理可能影响语气、置信度校准、推理质量、事实准确性、安全行为或指令遵循能力。
  • 量化可能进一步造成与 BF16 衍生模型不同的质量变化。
  • 内部标签 best15 仅用于构建版本筛选,不是标准化或独立复现的评测分数。
  • “Abliterated”或拒答率降低不代表模型具备无限能力,也不保证服从、正确或安全。

实测结果

以下结果来自通过 vLLM 部署的独立 FP8 Abliterix 检查点 Qwen3.8-27B-Abliterated-FP8,这些数据用于说明所选衍生检查点的实测行为。量化格式、推理框架、提示模板、采样配置和思考 token 预算均可能影响结果。

汇总

评测 范围与方法 结果
IFBench 官方 300 条测试集;生成上限 32,768 tokens; 78.16%
StrongREJECT 官方完整数据集前 150 条;关闭思考;温度 0; 1.5%
MMLU 前缀样本 cais/mmlu 的 all/test 前 200 条;5-shot;开启思考; 91.50%
MMLU 均衡样本 57 个主题各取测试集前 30 条,共 1,710 题;5-shot;开启思考; 88.01%

运行要求

建议使用包含 Qwen3.5 系列 MTP 推测解码和 qwen3_5_mtp 拆分工具的新版 MLX-VLM。实际兼容性取决于 MLX-VLM、Transformers 版本及 Apple Silicon 后端。

python -m pip install -U mlx-vlm huggingface_hub

本模型体积较大。可用统一内存不仅需要容纳模型权重,还应为视觉编码器、MTP 草稿、运行时缓冲区、提示词处理和 KV Cache 预留空间。上下文越长、图像越大,启用推测解码后的内存需求也越高。资源有限时建议先使用 4bit、普通解码和较短上下文。

下载指定版本

由于三个版本存放在同一仓库的不同子目录中,请先下载所需目录。以下示例使用推荐的 6bit 版本。

hf download GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP-MLX \
  --include "6bit/*" \
  --local-dir ./Qwen3.8-27B-Abliterated-MTP-MLX
./Qwen3.8-27B-Abliterated-MTP-MLX/6bit

使用 MLX-VLM

以下命令采用当前 MLX-VLM 命令行格式及 Qwen3.5 MTP 拆分工具。不同版本的参数名称和架构支持可能变化,如遇问题请运行相应模块并添加 --help。

准备独立 MTP 草稿模型

下载目标目录后执行一次以下命令。拆分工具会提取索引中的 mtp.* 张量,并保留 6 位 MTP 量化元数据:

python -m mlx_vlm.speculative.drafters.qwen3_5_mtp.split \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --output ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft

不要将 --draft-model 直接指向 model-mtp-6bit.safetensors;它只是权重文件,不是完整模型目录。应使用生成的 6bit-mtp-draft/ 目录。

普通文本生成

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --prompt "Explain the difference between symmetric and asymmetric encryption." \
  --max-tokens 512 \
  --temperature 0.6

请先使用这条普通解码命令作为基线,再启用 MTP。

启用 MTP 的文本生成

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --prompt "Explain the difference between symmetric and asymmetric encryption." \
  --max-tokens 512 \
  --temperature 0.6

--draft-block-size 3 只是起点,并非适用于所有场景的最佳值。请测试多个取值,并查看运行时输出的推测解码接受率统计。

启用 MTP 的图像理解

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --image /path/to/image.jpg \
  --prompt "Describe this image in detail." \
  --max-tokens 512 \
  --temperature 0.2

启动 API 服务

mlx_vlm.server \
  --model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit \
  --draft-model ./Qwen3.8-27B-Abliterated-MTP-MLX/6bit-mtp-draft \
  --draft-kind mtp \
  --draft-block-size 3 \
  --port 8080
curl http://localhost:8080/health

MLX-VLM 服务适合实验用途。在将接口暴露到网络之前,请根据实际环境配置身份认证、网络隔离、访问控制、日志审计、速率限制及其他必要保护措施。

如果 MTP 无法加载、导致内存不足或降低吞吐量,请移除 --draft-model、--draft-kind 和 --draft-block-size,恢复普通解码。推测解码支持仍在快速演进,请使用实际部署所采用的 MLX-VLM 版本和工作负载进行验证。

如何选择版本

  • 目标模型和 MTP 草稿保真度优先且统一内存充足时选择 8bit。
  • 一般用途推荐选择目标模型与 MTP 草稿质量和资源占用较均衡的 6bit。
  • 内存或存储空间有限时选择 4bit;由于草稿同样量化为 4 位,请重点验证 MTP 接受率。

不存在普遍最佳的 MTP 精度。请分别使用普通解码和推测解码,测量输出质量、解码吞吐量、首 token 延迟、峰值统一内存及每轮接受的草稿数量。

量化位宽并不直接决定拒答率。temperature、top-p 等采样参数会改变输出随机性,但不能复现或替代 Abliterix 的权重层干预。

预期用途

适合的用途包括:模型行为、对齐、拒答、鲁棒性与量化研究;经授权的评估和红队测试;MLX/MLX-VLM 本地部署实验;合法授权环境中的防御性网络安全教育与研究;文本、图像和视频多模态行为评估;Apple Silicon 上的量化 MTP 推测解码测试。

禁止及高风险用途

不得使用本模型实施或协助违法活动、未经授权的系统访问、凭据窃取、恶意软件投放、侵犯隐私、骚扰、暴力、欺诈、剥削或其他伤害行为。部署者必须根据应用场景实施访问控制、审计监控、速率限制、内容保护和人工复核。

免责声明

[!CAUTION]

[!CAUTION]
使用者自行承担全部风险。 Abliterix 处理会主动改变拒答相关行为,可能削弱官方模型原有的部分安全保护。模型可能生成错误、不安全、冒犯性、偏见性、违法、不可靠或其他有害内容,也可能比官方基础模型更容易遵循恶意指令。

本仓库及其中所有文件均按**“现状”和“可用状态”**提供,不作任何明示或默示保证,包括但不限于适销性、特定用途适用性、准确性、可靠性、不侵权性、安全性、无害性、兼容性、性能或持续可用性保证。

在适用法律允许的最大范围内,维护者、贡献者、量化者、转换者、分发者、Qwen 团队、阿里云、Apple、MLX/MLX-VLM 维护者以及 Abliterix 作者,均不对因使用或误用本衍生模型产生的提示词、输出、决策、部署、损害、损失、安全事件、索赔、责任或后果承担责任。文中提及任何组织或软件项目,均不代表其认可或参与本模型。

使用者和部署者应自行负责:

  1. 在部署前充分评估模型;
  2. 遵守适用法律法规、许可证、平台规则和第三方权利;
  3. 在开展任何网络安全测试前取得明确授权;
  4. 防止未经授权或不适合的用户访问;
  5. 根据实际用途实施必要的安全保护措施;
  6. 在依赖任何模型输出前进行独立核验;
  7. 保护个人、机密、专有及安全敏感数据。

本模型不得被视为专业的医疗、法律、金融、安全、网络安全或运营建议。对于错误可能导致人身伤害、权利侵害、经济损失、系统失陷或其他重大损害的决策,不得直接依赖模型输出。

许可证与署名

官方基础模型仓库将许可证标注为 Apache-2.0。使用者必须自行查阅并遵守基础模型、依赖项、MLX/MLX-VLM、推理软件、数据集、输入、输出及实际用途适用的许可证和相关条款。

本衍生版本不转移原始模型、商标、文档、软件或任何第三方材料的所有权。“Qwen”及相关标识和官方资料仍归各自权利人所有。

参考资料

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.