← back to catalog · registered 2026-09-28 09:57

GCSA-AiLab/Qwen3.8-27B-Abliterated-MTP

GCSA-AiLab Qwen 27B multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/GCSA-AiLab%2FQwen3.8-27B-Abliterated-MTP"
Response includes
  • classification unknown
  • files 2
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
5w ago
created 2026-08-19

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 3K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
multilingual
Tags
transformers safetensors qwen qwen3.8 abliterix compressed-tensors vllm quantized int8 int4 fp8 fp4

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-09-28 09:57
Last updated on HF
2026-09-03 08:16

Files by quantization

Auxiliary files 2 files 35.0 KB
README.md 33.3 KB dfbee706 download
.gitattributes 1.75 KB 3fcd2a4d download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.8-27B
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • qwen
  • qwen3.8
  • abliterix
  • compressed-tensors
  • safetensors
  • vllm
  • quantized
  • int8
  • int4
  • fp8
  • fp4
  • bf16
  • mtp
  • speculative-decoding
  • multimodal
  • vision-language
    language:
  • multilingual

Qwen3.8-27B Abliterated + MTP — BF16, INT8, INT4, FP8 and FP4 Collection

[!IMPORTANT]
This is an unofficial community derivative of Qwen/Qwen3.8-27B. It has been processed with Abliterix to reduce selected refusal-related behavior and is distributed as an unquantized BF16 checkpoint plus four compressed-tensors quantizations: INT8, INT4, FP8, and FP4. The four quantized variants also preserve the model's native MTP (Multi-Token Prediction) draft head in BF16 for optional speculative decoding.

This is not an official Qwen release and is not endorsed by the Qwen Team, Alibaba Cloud, NVIDIA, the vLLM/compressed-tensors/LLM Compressor maintainers, or the Abliterix authors.

Model Summary

Qwen3.8-27B is a dense, approximately 27-billion-parameter native vision-language model supporting text, image, and video understanding. This repository preserves the derivative architecture while applying an Abliterix refusal-direction intervention. It provides the resulting BF16 derivative as the highest-fidelity reference build, together with four post-training quantizations for deployment-oriented experimentation. The quantized variants include the native one-layer MTP head, which can draft tokens for verification by the main model when the inference runtime supports MTP speculative decoding.

Relationship to the official model

  • Official base model: Qwen/Qwen3.8-27B
  • Behavioral intervention: Abliterix refusal-direction intervention
  • Speculative-decoding component: native one-layer MTP head in INT8, INT4, FP8, and FP4; stored in BF16
  • Storage format: sharded Safetensors; quantized variants include compressed-tensors metadata
  • Available variants: BF16, INT8, INT4, FP8, FP4
  • Repository: GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP
  • Maintainer: Global Cybersecurity Alliance (community distribution)

The official Qwen benchmark results describe the unmodified base model only. They must not be interpreted as results for this Abliterix-processed or quantized derivative. No claim is made that this derivative preserves every capability or score of the official model.

Available Variants

The BF16 folder contains the unquantized Abliterix-processed derivative in native BFloat16 Safetensors. The other four variants use compressed-tensors 0.18-compatible metadata and include a separate model-mtp-bf16.safetensors file. In the quantized variants, the MTP head, vision encoder, multimodal projector, token embeddings, lm_head, and selected linear-attention convolution modules were excluded from quantization to preserve compatibility and stability. Consequently, actual repository sizes are larger than a simple parameter-count × bit-width estimate.

Folder Scheme Weight quantization Activation quantization Files Total bytes Approx. size Suggested use
BF16/ Native BF16 Unquantized BFloat16 BFloat16 63 54,733,721,635 51GB Highest-fidelity reference version; currently does not include the MTP weight tensors
INT8/ INT8 W8A8 + BF16 MTP 8-bit integer, per-channel, symmetric Dynamic 8-bit integer, per-token 64 31,240,486,740 29GB Conservative 8-bit integer deployment with optional MTP
INT4/ W4A16 + BF16 MTP 4-bit integer, group size 128, symmetric Unquantized A16 64 19,438,008,114 18GB Lower memory use; recommended integer option with optional MTP
FP8/ FP8 W8A8 + BF16 MTP 8-bit floating point, per-channel Dynamic 8-bit floating point, per-token 64 31,240,489,219 29GB FP8-capable accelerator deployment with optional MTP
FP4/ NVFP4A16 + BF16 MTP NVFP4 weights, group size 16, FP8 E4M3 scales Unquantized A16 64 20,579,438,970 19GB NVFP4-capable deployment with optional MTP

Each folder is a self-contained model directory containing configuration files, tokenizer and processor assets, a chat template, 55 main-model weight shards, and a weight index. The four quantized folders additionally contain quantization metadata and one 849,400,412-byte model-mtp-bf16.safetensors file. Their indexes map 15 mtp.* tensors to that file.

[!NOTE]
The current BF16/ index does not reference any mtp.* tensors and the folder does not contain a separate MTP weight file, although its text configuration declares mtp_num_hidden_layers: 1. Therefore, MTP availability in this repository currently applies to INT8/, INT4/, FP8/, and FP4/, not BF16/.

QUANTIZATION_COMPLETE.json, when present, is build-verification metadata and is not required for inference.

Precision and Quantization Notes

BF16

BF16 is the unquantized Abliterix-processed derivative and serves as the highest-fidelity reference checkpoint in this collection. It avoids additional quantization error and is the preferred source for evaluation, further conversion, fine-tuning experiments, or deployments where sufficient accelerator memory is available. Runtime memory will exceed the approximately 51GB weight size because additional space is required for the vision encoder execution, KV cache, activations, CUDA/runtime buffers, batching, and context processing.

INT8

INT8 uses static per-channel symmetric 8-bit integer weights and dynamic per-token symmetric 8-bit integer activations. It generally requires more memory than the 4-bit variants but offers a less aggressive integer quantization.

INT4

INT4 is a W4A16 weight-only scheme: weights are packed as symmetric 4-bit integers with group size 128, while activations remain at 16-bit precision. This provides the smallest integer-weight variant in this collection.

FP8

FP8 uses per-channel 8-bit floating-point weights and dynamic per-token 8-bit floating-point activations. Efficient execution depends on suitable accelerator and runtime support.

FP4

FP4 uses the NVFP4A16 scheme: 4-bit floating-point weights with group size 16 and FP8 E4M3 local scales, while activations remain at 16-bit precision. Native acceleration is hardware-dependent; unsupported platforms may fail to load the model or fall back to a slower path if the runtime provides one.

MTP and Speculative Decoding

MTP (Multi-Token Prediction) uses the model's native draft head to propose one or more future tokens. A compatible inference engine then verifies those proposals with the main model. Accepted tokens can reduce decode latency or improve decode throughput without changing the target model's output distribution; rejected proposals are discarded. MTP is an optional inference optimization, not a different chat mode, a larger context window, or an additional safety modification.

In INT8/, INT4/, FP8/, and FP4/, the configuration declares one MTP hidden layer (mtp_num_hidden_layers: 1) without dedicated embeddings (mtp_use_dedicated_embeddings: false). The MTP tensors are deliberately kept in BF16 rather than quantized. This adds about 0.79 GiB to each quantized folder and also consumes additional accelerator memory when loaded.

MTP acceleration is not automatic merely because the weights are present. It requires explicit runtime support and usually must be enabled in the server configuration. Actual gains depend on draft-token acceptance rate, request length, batch size, hardware, quantization kernel, and the number of speculative tokens. Benchmark both enabled and disabled on the target deployment; a higher speculative depth can be slower when acceptance is low.

What Abliterix Changes

Abliterix identifies and modifies activation directions associated with refusal behavior. The intended effect is to reduce excessive refusals on selected evaluations. This is a weight-level behavioral intervention rather than a system-prompt override.

Important limitations:

  • Reduced refusal behavior is not guaranteed for every prompt, language, chat template, or inference engine.
  • The model may still refuse requests because refusal behavior can be distributed across multiple layers and mechanisms.
  • Abliterix processing may affect tone, calibration, reasoning quality, factual accuracy, safety behavior, or instruction following.
  • Quantization may introduce additional quality differences compared with the BF16 derivative.
  • The internal label best15 is a build-selection label, not a standardized or independently reproduced benchmark score.
  • “Abliterated” or reduced refusal does not mean unrestricted capability, guaranteed compliance, correctness, or safety.

Evaluation Results

The following results were measured on a separately deployed FP8 Abliterix checkpoint labeled Qwen3.8-27B-Abliterated-FP8, served through vLLM. They are included as behavioral evidence for the selected derivative checkpoint. Quantization format, runtime, prompt template, sampling configuration, and thinking-token budget can all affect the results.

Summary

Evaluation Scope and protocol Result
IFBench Official 300-prompt test set; 32,768-token generation limit; 78.16%
StrongREJECT First 150 rows of the official full dataset; non-thinking generation; temperature 0; 1.5%
MMLU prefix sample cais/mmlu, all/test, first 200 rows; 5-shot; thinking enabled; 91.50%
MMLU balanced sample First 30 test rows from each of 57 subjects; 1,710 questions; 5-shot; thinking enabled; 88.01%

Runtime and Hardware Compatibility

The BF16 checkpoint is standard unquantized sharded Safetensors and does not use a quantization backend. The INT8, INT4, FP8, and FP4 checkpoints use compressed-tensors, not GGUF, MLX, GPTQ, AWQ, or bitsandbytes format. A recent vLLM build is the recommended starting point because it can read the quantization metadata and provides an MTP speculative-decoding configuration. Support for this exact combination of the Qwen3.5 architecture, compressed-tensors, and BF16 MTP weights must still be verified with the installed release.

[!WARNING]
Quantized-kernel support varies by vLLM release, GPU architecture, CUDA/ROCm version, operating system, and model architecture. A checkpoint being structurally valid does not guarantee that every backend can execute it. Verify compatibility with the exact hardware and software stack before deployment.

FP8 and NVFP4 generally benefit most from GPUs with native support for the corresponding number formats. INT4 and INT8 also require compatible kernels. On unsupported hardware, use another variant or the BF16 derivative rather than assuming transparent compatibility. BF16 itself requires hardware and software with BFloat16 support and substantially more memory than the quantized variants.

Download a Variant

Because the variants are stored in subfolders, download the required folder first. The example below downloads the MTP-enabled INT4 variant.

python -m pip install -U huggingface_hub

hf download GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP \
  --include "INT4/*" \
  --local-dir ./Qwen3.8-27B-Abliterated-MTP

The resulting local model path is:

./Qwen3.8-27B-Abliterated-MTP/INT4

Replace INT4 with BF16, INT8, FP8, or FP4 in both places to download another variant. Remember that the current BF16/ folder does not include MTP weight tensors.

Quick Start with vLLM

Install a recent vLLM build appropriate for the target CUDA/ROCm environment. Check the official vLLM installation documentation rather than blindly reusing an installation command from another machine.

Start an OpenAI-compatible server

vllm serve ./Qwen3.8-27B-Abliterated-MTP/INT4 \
  --served-model-name Qwen3.8-27B-Abliterated-MTP-INT4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --max-model-len 32768 \
  --host 0.0.0.0 \
  --port 8000

For multi-GPU deployment, add an appropriate tensor-parallel setting, for example:

--tensor-parallel-size 2

Do not assume that the example context length or tensor-parallel size is suitable for every system. Adjust context length, GPU memory utilization, batching, and parallelism according to available memory and runtime validation.

The speculative depth of 3 is only a starting point. Compare several values and a run without --speculative-config. If the installed vLLM release cannot load the checkpoint with MTP enabled, remove that option and run ordinary decoding, or use a runtime version that explicitly supports the model and quantization combination.

Text request

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-Abliterated-MTP-INT4",
    "messages": [
      {"role": "user", "content": "Explain the difference between symmetric and asymmetric encryption."}
    ],
    "temperature": 0.6,
    "max_tokens": 512
  }'

For image or video input, follow the multimodal request format supported by the installed vLLM release and confirm that the Qwen architecture is supported. Multimodal payload formats can change between releases.

Choosing a Variant

  • Choose BF16 for the highest-fidelity reference, evaluation, or further conversion when MTP is not required; this folder currently lacks MTP weight tensors.
  • Choose INT8 for a less aggressive integer quantization with an included BF16 MTP head when compatible W8A8 and MTP kernels are available.
  • Choose INT4 to reduce main-model weight memory while keeping 16-bit activations and an included BF16 MTP head.
  • Choose FP8 when the target accelerator and runtime efficiently support FP8 W8A8 together with MTP.
  • Choose FP4 when NVFP4 and MTP support are available and minimizing main-model floating-point weight storage is the priority.

There is no universally “best” version. Benchmark accuracy, latency, throughput, peak memory, multimodal behavior, refusal rate, and MTP acceptance rate on the actual deployment stack. Compare MTP against ordinary decoding rather than assuming it will always improve performance.

Quantization level does not directly determine refusal rate. Sampling parameters such as temperature and top-p affect output variability but do not reproduce or replace the weight-level Abliterix intervention.

Intended Use

Appropriate uses include:

  • research into model behavior, alignment, refusal, robustness, and quantization;
  • authorized evaluation and red-team testing;
  • local or controlled vLLM deployment experiments;
  • defensive cybersecurity education and research in authorized environments;
  • evaluation of multimodal behavior on text, image, and video inputs.

Prohibited and High-Risk Use

Do not use this model to facilitate unlawful activity, unauthorized access, credential theft, malware deployment, privacy invasion, harassment, violence, fraud, exploitation, or other harm. Operators must implement access controls, monitoring, rate limits, content safeguards, and human review appropriate to the deployment context.

Disclaimer

[!CAUTION]
USE AT YOUR OWN RISK. Abliterix processing intentionally changes refusal-related behavior and may weaken safeguards present in the official model. The model may generate inaccurate, unsafe, offensive, biased, unlawful, insecure, or otherwise harmful content and may follow malicious instructions more readily than the official base model.

This repository and all included files are provided “AS IS” and “AS AVAILABLE,” without warranties or conditions of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, accuracy, reliability, non-infringement, security, safety, compatibility, performance, or uninterrupted availability.

The maintainers, contributors, quantizers, converters, distributors, the Qwen Team, Alibaba Cloud, NVIDIA, vLLM/compressed-tensors/LLM Compressor maintainers, and the Abliterix authors are not responsible for prompts, outputs, decisions, deployments, damages, losses, security incidents, claims, liabilities, or consequences arising from use or misuse of this derivative, to the maximum extent permitted by applicable law. Mention of any organization, product, or software project does not imply endorsement, sponsorship, or involvement.

Users and deployers are solely responsible for:

  1. evaluating the model and every available variant before deployment;
  2. verifying compatibility with the exact hardware and software stack;
  3. complying with applicable laws, regulations, licenses, platform policies, and third-party rights;
  4. obtaining explicit authorization before any cybersecurity testing;
  5. preventing access by unauthorized or unsuitable users;
  6. implementing safeguards appropriate to the use case;
  7. independently verifying model outputs before relying on them;
  8. protecting personal, confidential, proprietary, and security-sensitive data.

The model must not be treated as professional medical, legal, financial, safety, security, cybersecurity, or operational advice. Do not rely on it for decisions where errors could cause injury, rights violations, financial loss, system compromise, or other material harm.

[!CAUTION]

License and Attribution

The official base-model repository identifies its license as Apache-2.0. Users must independently review and comply with the base-model license and any terms applicable to dependencies, compressed-tensors, vLLM, inference software, datasets, inputs, outputs, and the intended use.

This derivative does not transfer ownership of the original model, trademarks, documentation, software, or third-party materials. “Qwen,” related marks, and official documentation remain the property of their respective owners.



[!IMPORTANT]
本仓库是 Qwen/Qwen3.8-27B 的非官方社区衍生版本。模型经过 Abliterix 处理,以降低部分拒答相关行为,并提供未量化的 BF16 检查点,以及采用 compressed-tensors 格式的 INT8、INT4、FP8 和 FP4 四种量化版本。四个量化版本还以 BF16 保留了模型原生的 **MTP(Multi-Token Prediction,多 token 预测)**草稿头,可用于可选的推测解码。
本模型并非 Qwen 官方发布,也不代表 Qwen 团队、阿里云、NVIDIA、vLLM/compressed-tensors/LLM Compressor 维护者或 Abliterix 作者的立场或认可。

模型简介

Qwen3.8-27B 是一个约 270 亿参数的原生视觉语言模型,支持文本、图像和视频理解。本仓库在保留衍生模型架构的基础上,使用 Abliterix 对拒答相关方向进行干预,并将得到的 BF16 衍生模型作为最高保真参考版本,同时提供四种面向部署实验的训练后量化版本。量化版本包含原生单层 MTP 草稿头;当推理框架支持 MTP 推测解码时,它可以先提出候选 token,再由主模型验证。

与官方模型的关系

  • 官方基础模型: Qwen/Qwen3.8-27B
  • 行为干预: Abliterix refusal-direction intervention
  • 推测解码组件: native one-layer MTP head in INT8, INT4, FP8, and FP4; stored in BF16
  • 存储格式: sharded Safetensors; quantized variants include compressed-tensors metadata
  • 可用版本: BF16, INT8, INT4, FP8, FP4
  • 仓库: GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP
  • 维护者: Global Cybersecurity Alliance (community distribution)

官方 Qwen 模型卡中的基准测试结果仅代表未经修改的基础模型,不能视为本 Abliterix 衍生版或量化版的实测成绩。本仓库不保证衍生模型完整保持官方模型的全部能力或分数。

可用版本

BF16 目录包含经 Abliterix 处理后、以原生 BFloat16 Safetensors 保存的未量化衍生模型。其余四个版本均包含兼容 compressed-tensors 0.18 的量化元数据,并额外包含独立的 model-mtp-bf16.safetensors。为了兼容性和稳定性,量化版本中的 MTP 头、视觉编码器、多模态投影器、词嵌入、lm_head 以及部分线性注意力卷积模块未量化。因此,实际文件大小会高于“参数量 × 位宽”的简单理论估算。

Folder Scheme Weight quantization Activation quantization Files Total bytes Approx. size Suggested use
BF16/ Native BF16 Unquantized BFloat16 BFloat16 63 54,733,721,635 51GB Highest-fidelity reference version; currently does not include the MTP weight tensors
INT8/ INT8 W8A8 + BF16 MTP 8-bit integer, per-channel, symmetric Dynamic 8-bit integer, per-token 64 31,240,486,740 29GB Conservative 8-bit integer deployment with optional MTP
INT4/ W4A16 + BF16 MTP 4-bit integer, group size 128, symmetric Unquantized A16 64 19,438,008,114 18GB Lower memory use; recommended integer option with optional MTP
FP8/ FP8 W8A8 + BF16 MTP 8-bit floating point, per-channel Dynamic 8-bit floating point, per-token 64 31,240,489,219 29GB FP8-capable accelerator deployment with optional MTP
FP4/ NVFP4A16 + BF16 MTP NVFP4 weights, group size 16, FP8 E4M3 scales Unquantized A16 64 20,579,438,970 19GB NVFP4-capable deployment with optional MTP

每个目录都是独立、完整的模型目录,包含配置文件、分词器和处理器文件、聊天模板、55 个主模型权重分片及权重索引。四个量化目录还包含量化元数据和一个大小为 849,400,412 字节的 model-mtp-bf16.safetensors,其索引将 15 个 mtp.* 张量指向该文件。

[!NOTE]
当前 BF16/ 的权重索引没有引用任何 mtp.* 张量,目录中也没有独立的 MTP 权重文件,尽管文本配置声明了 mtp_num_hidden_layers: 1。因此,本仓库当前的 MTP 可用范围是 INT8/、INT4/、FP8/ 和 FP4/,不包括 BF16/。

精度与量化说明

BF16

BF16 是经 Abliterix 处理后的未量化衍生模型,也是本集合中保真度最高的参考检查点。它不会引入额外量化误差,适合用于效果评估、进一步格式转换、微调实验,或在加速器内存充足时直接部署。实际运行内存会高于约 51GB 的权重体积,因为视觉编码器执行、KV Cache、激活值、CUDA/运行时缓冲区、批处理和上下文处理都需要额外空间。

INT8

INT8 使用静态逐通道对称 8 位整数权重及动态逐 token 对称 8 位整数激活。其内存需求通常高于 4 位版本,但整数压缩程度相对保守。

INT4

INT4 使用 W4A16 仅权重量化:权重以 group size 128 的对称 4 位整数打包,激活保持 16 位精度。这是本集合中体积最小的整数权重版本。

FP8

FP8 使用逐通道 8 位浮点权重和动态逐 token 8 位浮点激活。能否高效运行取决于加速器和推理框架是否提供相应支持。

FP4

FP4 使用 NVFP4A16:权重为 group size 16 的 4 位浮点格式,并使用 FP8 E4M3 局部缩放因子,激活保持 16 位精度。原生加速依赖硬件;不支持的平台可能无法加载,或在运行时提供回退路径时以较慢方式执行。

MTP 与推测解码

MTP(Multi-Token Prediction,多 token 预测)使用模型原生草稿头预先提出一个或多个后续 token,再由兼容的推理框架调用主模型进行验证。通过验证的 token 可以降低解码延迟或提升解码吞吐量,未通过的候选会被丢弃;该过程不会改变目标模型的输出分布。MTP 是可选的推理优化,并非新的对话模式、更大的上下文窗口或额外的安全行为修改。

在 INT8/、INT4/、FP8/ 和 FP4/ 中,配置声明了一个 MTP 隐藏层(mtp_num_hidden_layers: 1),且不使用独立嵌入(mtp_use_dedicated_embeddings: false)。MTP 张量刻意保持 BF16 而未量化,因此每个量化目录增加约 0.79 GiB 文件体积,加载后也会额外占用加速器内存。

仅存在 MTP 权重并不会自动启用加速。推理框架必须明确支持 MTP,通常还需要在服务配置中显式开启。实际收益取决于候选 token 接受率、请求长度、批大小、硬件、量化内核和推测 token 数量。请在目标部署环境中分别测试开启与关闭 MTP;当接受率较低时,更高的推测深度反而可能更慢。

Abliterix 修改说明

Abliterix 用于识别并修改与拒答行为相关的激活方向,目标是在特定评测中减少过度拒答。这属于权重层面的行为干预,并非简单覆盖系统提示词。

重要限制:

  • 模型不保证对所有提示词、语言、聊天模板或推理框架都降低拒答。
  • 拒答行为可能分布于多个层和机制,因此模型仍可能拒绝部分请求。
  • Abliterix 处理可能影响语气、置信度校准、推理质量、事实准确性、安全行为或指令遵循能力。
  • 量化可能进一步造成与 BF16 衍生模型不同的质量变化。
  • 内部标签 best15 仅用于构建版本筛选,不是标准化或独立复现的评测分数。
  • “Abliterated”或拒答率降低不代表模型具备无限能力,也不保证服从、正确或安全。

实测结果

以下结果来自通过 vLLM 部署的独立 FP8 Abliterix 检查点 Qwen3.8-27B-Abliterated-FP8,这些数据用于说明所选衍生检查点的实测行为。量化格式、推理框架、提示模板、采样配置和思考 token 预算均可能影响结果。

汇总

评测 范围与方法 结果
IFBench 官方 300 条测试集;生成上限 32,768 tokens; 78.16%
StrongREJECT 官方完整数据集前 150 条;关闭思考;温度 0; 1.5%
MMLU 前缀样本 cais/mmlu 的 all/test 前 200 条;5-shot;开启思考; 91.50%
MMLU 均衡样本 57 个主题各取测试集前 30 条,共 1,710 题;5-shot;开启思考; 88.01%

运行时与硬件兼容性

BF16 检查点是标准的未量化分片 Safetensors,不依赖量化后端。INT8、INT4、FP8 和 FP4 检查点使用 compressed-tensors,并非 GGUF、MLX、GPTQ、AWQ 或 bitsandbytes 格式。建议优先使用最新版 vLLM,因为它可以读取量化元数据并提供 MTP 推测解码配置;但仍须在实际安装版本中验证 Qwen3.5 架构、compressed-tensors 与 BF16 MTP 权重这一组合是否兼容。

[!WARNING]
不同 vLLM 版本、GPU 架构、CUDA/ROCm 版本、操作系统和模型架构支持的量化内核不同。检查点结构有效并不代表所有后端都能运行。部署前必须在目标软硬件环境中验证兼容性。

FP8 和 NVFP4 通常在原生支持相应数值格式的 GPU 上收益最大;INT4 和 INT8 同样需要兼容内核。不支持时,应改用其他量化版本或 BF16 衍生版,不应假设推理框架能够自动兼容。BF16 本身也要求软硬件支持 BFloat16,并且需要显著高于量化版本的内存。

下载指定版本

由于五个版本存放于不同子目录,请先下载所需版本。以下示例下载支持 MTP 的 INT4 版本。

python -m pip install -U huggingface_hub

hf download GlobalCybersecurityAlliance/Qwen3.8-27B-Abliterated-MTP \
  --include "INT4/*" \
  --local-dir ./Qwen3.8-27B-Abliterated-MTP
./Qwen3.8-27B-Abliterated-MTP/INT4

使用 vLLM

请根据目标 CUDA/ROCm 环境安装合适的最新版 vLLM。建议查阅 vLLM 官方安装文档,不要直接照搬其他机器的安装命令。

启动兼容 OpenAI API 的服务

vllm serve ./Qwen3.8-27B-Abliterated-MTP/INT4 \
  --served-model-name Qwen3.8-27B-Abliterated-MTP-INT4 \
  --trust-remote-code \
  --dtype bfloat16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --max-model-len 32768 \
  --host 0.0.0.0 \
  --port 8000
--tensor-parallel-size 2

不要假设示例中的上下文长度或张量并行数适合所有系统。请根据可用显存和实测结果调整上下文长度、GPU 内存利用率、批处理和并行参数。

示例中的推测深度 3 只是起点。请对比多个取值,并与移除 --speculative-config 后的普通解码进行测试。如果当前 vLLM 版本无法在开启 MTP 时加载该检查点,请移除该选项使用普通解码,或改用明确支持此模型与量化组合的运行时版本。

文本请求

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-Abliterated-MTP-INT4",
    "messages": [
      {"role": "user", "content": "Explain the difference between symmetric and asymmetric encryption."}
    ],
    "temperature": 0.6,
    "max_tokens": 512
  }'

对于图像或视频输入,请遵循当前安装版本 vLLM 支持的多模态请求格式,并确认其支持本 Qwen 架构。不同版本的多模态请求格式可能变化。

如何选择版本

  • 在不需要 MTP,且追求最高保真度、效果评估或进一步格式转换时选择 BF16;该目录目前没有 MTP 权重张量。
  • 具备兼容 W8A8 与 MTP 内核,并希望采用相对保守的整数量化时选择 INT8;其中包含 BF16 MTP 头。
  • 希望降低主模型权重内存、保持 16 位激活并使用 BF16 MTP 头时选择 INT4。
  • 目标加速器及推理框架能够同时高效执行 FP8 W8A8 与 MTP 时选择 FP8。
  • 具备 NVFP4 与 MTP 支持,且优先降低主模型浮点权重体积时选择 FP4。

不存在适用于所有环境的“最佳”版本。请在实际部署环境中测试准确率、延迟、吞吐量、峰值内存、多模态行为、拒答率和 MTP 接受率,并将 MTP 与普通解码直接对比,不要假设它一定能提升性能。

量化位宽并不直接决定拒答率。temperature、top-p 等采样参数会改变输出随机性,但不能复现或替代 Abliterix 的权重层干预。

预期用途

适合的用途包括:模型行为、对齐、拒答、鲁棒性与量化研究;经授权的评估和红队测试;本地或受控环境中的 vLLM 部署实验;合法授权环境中的防御性网络安全教育与研究;文本、图像和视频多模态行为评估。

禁止及高风险用途

不得使用本模型实施或协助违法活动、未经授权的系统访问、凭据窃取、恶意软件投放、侵犯隐私、骚扰、暴力、欺诈、剥削或其他伤害行为。部署者必须根据应用场景实施访问控制、审计监控、速率限制、内容保护和人工复核。

免责声明

[!CAUTION]

[!CAUTION]
使用者自行承担全部风险。 Abliterix 处理会主动改变拒答相关行为,可能削弱官方模型原有的部分安全保护。模型可能生成错误、不安全、冒犯性、偏见性、违法、不可靠或其他有害内容,也可能比官方基础模型更容易遵循恶意指令。

本仓库及其中所有文件均按**“现状”和“可用状态”**提供,不作任何明示或默示保证,包括但不限于适销性、特定用途适用性、准确性、可靠性、不侵权性、安全性、无害性、兼容性、性能或持续可用性保证。

在适用法律允许的最大范围内,维护者、贡献者、量化者、转换者、分发者、Qwen 团队、阿里云、NVIDIA、vLLM/compressed-tensors/LLM Compressor 维护者以及 Abliterix 作者,均不对因使用或误用本衍生模型产生的提示词、输出、决策、部署、损害、损失、安全事件、索赔、责任或后果承担责任。文中提及任何组织、产品或软件项目,均不代表其认可、赞助或参与本模型。

使用者和部署者应自行负责:

  1. 在部署前充分评估模型及每个可用版本;
  2. 验证目标软硬件环境的兼容性;
  3. 遵守适用法律法规、许可证、平台规则和第三方权利;
  4. 在开展任何网络安全测试前取得明确授权;
  5. 防止未经授权或不适合的用户访问;
  6. 根据实际用途实施必要的安全保护措施;
  7. 在依赖任何模型输出前进行独立核验;
  8. 保护个人、机密、专有及安全敏感数据。

本模型不得被视为专业的医疗、法律、金融、安全、网络安全或运营建议。对于错误可能导致人身伤害、权利侵害、经济损失、系统失陷或其他重大损害的决策,不得直接依赖模型输出。

许可证与署名

官方基础模型仓库将许可证标注为 Apache-2.0。使用者必须自行查阅并遵守基础模型、依赖项、compressed-tensors、vLLM、推理软件、数据集、输入、输出及实际用途适用的许可证和相关条款。

本衍生版本不转移原始模型、商标、文档、软件或任何第三方材料的所有权。“Qwen”及相关标识和官方资料仍归各自权利人所有。

参考资料

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.