base_model: ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:
- en
- fr
- es
- de
- zh
- ru
- ja
- ko
tags: - gguf
- llama.cpp
- quantized
- val-apex-i
- nanoplus-v3
- unsloth-studio
- qwen
- qwen38
- uncensored
- abliterated
- speculative-decoding
- mtp
- vision
- multimodal
- imatrix
- reasoning
- v3
pipeline_tag: text-generation
Swift-1.5-Qwen3.8-27B-Uncensored-MTP VAL-APEX-I NanoPlus V3 GGUF
Independently computed on NVIDIA RTX PRO 6000 Ada hardware clusters. If this handcrafted release saves you VRAM and runs faster on your GPU, consider fueling the community compute fund on Ko-fi.
Uncensored Frontier 27B Reasoning MoE with Intact Vision, Multilingual Syntax Preservation, and Native MTP Self-Speculation
Official VAL-APEX-I NanoPlus V3 quantization of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP. This edition quantizes the uncensored foundation model built on UkisAI's Swift-1.5 reasoning fine-tune of Qwen3.8-27B.
- Multilingual Stability Guaranteed: Unlike models whose refusal abliteration damages foreign-language syntax, this base preserves French, Spanish, German, and Chinese grammar and agreement markers.
- Native MTP Speculative Head Included: Features the built-in Multi-Token Prediction auxiliary head (
blk.64.*, 15 tensors), enabling fast speculative decoding directly in llama.cpp without an external draft model. - Multimodal Optical Vision: Retains the high-precision vision encoder (
mmproj-Q8_0.gguf) for visual and document reasoning. - ISTA-DASLab Calibrated: Quantized with the official ISTA-DASLab importance matrix (
imatrix-qwen3.8-27b.gguf) computed across diverse tokens to safeguard critical routing and projection weights.
VAL-APEX-I stands for:
Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization
[!NOTE]
EXPLORE THE QWEN3.8 27B VAL-APEX-I EDITIONS
Choose a model variant and quantization profile:
- Swift-1.5 Qwen3.8 27B Uncensored MTP — MiniPlus V3 — 14.28 GB (+5.70% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Swift-1.5 Qwen3.8 27B Uncensored MTP — NanoPlus V3 — 11.08 GB (+2.26% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B Base — MiniPlus V3 — 14.16 GB (+13.21% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B Base — NanoPlus V3 — 11.35 GB (+11.49% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 — 15.08 GB (+15.01% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B EfficientThink Uncensored — NanoPlus — 11.65 GB (+10.87% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 — 15.33 GB (+13.06% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
- Qwen3.8 27B Huihui Abliterated — NanoPlus — 11.90 GB (+9.14% improvement in WikiText-2 PPL vs GSQ-RCO IQ3_S).
VAL-APEX-I collection · MiniPlus collection · NanoPlus collection
[!IMPORTANT]
ULTRA-COMPACT 11.08GB BUILD
This V3 release uses the asymmetric VAL-APEX-I per-tensor allocation with an actual 11.08 GB file size and 2.85 whole-file BPW. Its empirical WikiText-2 PPL is 6.9102, providing a lightweight footprint that fits comfortably into smaller VRAM envelopes while beating standard baselines like GSQ-RCO IQ3_S.
Quick Navigation Index
- Quantization Comparison: Metrics & Tensor Map
- Model Files & Technical Specifications
- Native Context & Runtime Memory
- Recommended Configuration & Setup
- Recommended Generation Parameters
- Native Chat Template & Reasoning
- Community Compute Fund & Priority Model Requests
1. Quantization Comparison
Size & Quality Metrics
| Edition / quantization | File size | Whole-file BPW | WikiText-2 PPL | Improvement in PPL vs GSQ-RCO IQ3_S (7.07) |
|---|---|---|---|---|
| BF16 Base baseline (unmodified) | 54.65 GB (50.90 GiB) |
16-bit | 5.7942 | — |
| Swift-1.5 Qwen3.8 27B Uncensored MTP — NanoPlus V3 (this release) | 11.08 GB (10.32 GiB) |
2.85 | 6.9102 | 2.26% improvement (0.1598 PPL reduction) |
| Swift-1.5 Qwen3.8 27B Uncensored MTP — MiniPlus V3 | 14.28 GB (13.30 GiB) |
3.93 | 6.6669 | 5.70% improvement (0.4031 PPL reduction) |
| Qwen3.8 27B Base — MiniPlus V3 | 14.16 GB (13.19 GiB) |
4.15 | 6.1361 | 13.21% improvement (0.9339 PPL reduction) |
| Qwen3.8 27B Base — NanoPlus V3 | 11.35 GB (10.57 GiB) |
3.32 | 6.2577 | 11.49% improvement (0.8123 PPL reduction) |
| Qwen3.8 27B EfficientThink Uncensored — MiniPlus V2.1 | 15.08 GB (14.04 GiB) |
4.49 | 6.0091 | 15.01% improvement (1.0609 PPL reduction) |
| Qwen3.8 27B EfficientThink Uncensored — NanoPlus | 11.65 GB (10.85 GiB) |
3.46 | 6.3013 | 10.87% improvement (0.7687 PPL reduction) |
| Qwen3.8 27B Huihui Abliterated — MiniPlus V2.1 | 15.33 GB (14.28 GiB) |
4.49 | 6.1466 | 13.06% improvement (0.9234 PPL reduction) |
| Qwen3.8 27B Huihui Abliterated — NanoPlus | 11.90 GB (11.08 GiB) |
3.48 | 6.4238 | 9.14% improvement (0.6462 PPL reduction) |
| GSQ-RCO IQ3_S | 11.8 GB | 3.50 (published) | 7.07 (published) | 0.00% (comparison baseline) |
The NanoPlus V3 file is 11.08 GB, approximately 0.72 GB smaller than GSQ-RCO IQ3_S, with 2.26% lower reported WikiText-2 PPL (6.9102 vs 7.07). MiniPlus V3 reaches 5.70% lower reported PPL (6.6669 vs 7.07) at 14.28 GB.
[!NOTE]
Perplexity Dynamics & Swift Alignment Impact
WikiText-2 measures raw next-token cross-entropy against encyclopedic Wikipedia prose. The Swift-1.5 fine-tuning process by UkisAI introduces deliberate architectural and distribution shifts:
- Reasoning Density and Logit Compression: Swift-1.5 focuses probability mass on concise, dense chain-of-thought pathways (
<think>) and direct reasoning outputs, stripping verbose filler tokens. This purposeful shift away from general Wikipedia text distribution naturally yields a slightly higher raw perplexity score on WikiText-2 (e.g. 6.66 vs 6.13 on raw base) without representing any degradation in actual intelligence or reasoning capability.- Speculative Decoding Efficiency: The true efficiency dividend of Swift-1.5 lies in inference throughput, realized through the integrated Multi-Token Prediction (MTP) head (
blk.64.*, 844M parameters). In runtime, this yields substantial generation speedups via high speculative acceptance rates rather than static single-token prediction loss on historical corpora.
Tensor Precision Map
| Component | MiniPlus V3 | NanoPlus V3 | GSQ-RCO IQ3_S |
|---|---|---|---|
| Output head | Q6_K ×1 |
Q6_K ×1 |
Q4_K ×1 |
| Token embeddings | Q4_K ×1 |
Q3_K ×1 |
IQ2_S ×1 |
| Normalizations | F32 ×161 |
F32 ×161 |
F32 ×161 |
SSM A (ssm_a) |
F32 ×48 |
F32 ×48 |
F32 ×48 |
SSM convolution (ssm_conv1d) |
F32 ×48 |
F32 ×48 |
F32 ×48 |
SSM time-step bias (ssm_dt) |
F32 ×48 |
F32 ×48 |
F32 ×48 |
SSM norm (ssm_norm) |
F32 ×48 |
F32 ×48 |
F32 ×48 |
Attention gates (attn_gate) |
Q4_0 ×1Q4_K ×47 |
Q4_0 ×1Q4_K ×47 |
IQ2_S ×2IQ3_S ×18IQ3_XXS ×9IQ4_XS ×12Q2_K ×4Q4_K ×3 |
Linear QKV (attn_qkv) |
Q4_0 ×1Q4_K ×47 |
Q3_K ×47Q4_0 ×1 |
IQ2_XS ×1IQ2_XXS ×1IQ3_S ×22IQ3_XXS ×13IQ4_XS ×9Q2_K ×1Q4_K ×1 |
SSM alpha (ssm_alpha) |
F32 ×47Q4_0 ×1 |
F32 ×47Q4_0 ×1 |
BF16 ×48 |
SSM beta (ssm_beta) |
Q4_0 ×1Q8_0 ×47 |
Q4_0 ×1Q8_0 ×47 |
BF16 ×48 |
SSM output (ssm_out) |
Q4_0 ×1Q5_K ×47 |
Q4_0 ×1Q4_K ×47 |
IQ3_S ×22IQ3_XXS ×4IQ4_XS ×16Q4_K ×6 |
Full attention Q (attn_q) |
Q4_K ×16 |
Q3_K ×16 |
IQ2_XXS ×1IQ3_S ×3IQ3_XXS ×3IQ4_XS ×2Q2_K ×6Q4_K ×1 |
Full attention K (attn_k) |
Q5_K ×16 |
Q4_K ×16 |
IQ2_S ×1IQ3_S ×1IQ3_XXS ×1IQ4_XS ×8Q4_K ×5 |
Full attention V (attn_v) |
Q6_K ×16 |
Q5_K ×16 |
IQ3_S ×6IQ3_XXS ×1IQ4_XS ×1Q4_K ×8 |
Full attention output (attn_output) |
Q6_K ×16 |
Q6_K ×16 |
IQ3_S ×10IQ3_XXS ×1IQ4_XS ×2Q4_K ×3 |
MLP down (ffn_down) |
IQ4_XS ×64 |
IQ3_XXS ×48IQ4_XS ×16 |
IQ2_S ×4IQ2_XS ×3IQ3_S ×22IQ3_XXS ×7IQ4_XS ×21Q2_K ×1Q4_K ×6 |
MLP gate (ffn_gate) |
IQ3_S ×32IQ3_XXS ×32 |
IQ2_XXS ×48IQ3_XXS ×16 |
IQ1_M ×1IQ2_S ×4IQ2_XS ×4IQ2_XXS ×1IQ3_S ×15IQ3_XXS ×21IQ4_XS ×15Q2_K ×1Q4_K ×2 |
MLP up (ffn_up) |
IQ3_S ×32IQ3_XXS ×32 |
IQ2_XXS ×48IQ3_XXS ×16 |
IQ2_S ×5IQ2_XS ×1IQ2_XXS ×2IQ3_S ×25IQ3_XXS ×18IQ4_XS ×10Q4_K ×3 |
Additional MTP head (blk.64.*) |
F32 ×7Q4_0 ×6Q4_1 ×1Q6_K ×1 |
F32 ×7Q4_0 ×6Q4_1 ×1Q6_K ×1 |
— |
x is the number of tensors in each format. Counts were read from the released GGUF headers; the GSQ-RCO column uses its published tensor allocation.
Both V3 files contain 866 tensors: 851 main-model tensors plus 15 tensors in the additional MTP head. The 64 main layers are 48 Gated DeltaNet + 16 full-attention; blk.64.* is the MTP head, not a seventeenth main full-attention layer.
Fallback exceptions: blk.0.attn_gate.weight, blk.0.attn_qkv.weight, blk.0.ssm_alpha.weight, blk.0.ssm_beta.weight and blk.0.ssm_out.weight are Q4_0 in both editions. In the MTP head, Q/K/V, MLP gate/up and nextn.eh_proj are Q4_0; MLP down is Q4_1, attention output is Q6_K, and the remaining seven tensors are F32.
V3 layer allocation: MiniPlus uses IQ3_S for MLP gate/up in layers 0-15 and 48-63, and IQ3_XXS in layers 16-47; every main MLP down tensor is IQ4_XS. NanoPlus uses IQ3_XXS for gate/up in layers 0-7 and 56-63, and IQ2_XXS in layers 8-55; MLP down is IQ4_XS in the outer 16 layers and IQ3_XXS in the other 48.
Complete tensor-by-tensor audit: names, formats, shapes and parameter counts.
2. Model Files & Technical Specifications
| File Name | File Size | Whole-file BPW | Description |
|---|---|---|---|
Swift-1.5-Qwen3.8-27B-Uncensored-MTP-VAL-APEX-I-NanoPlus-V3.gguf |
11.08 GB |
2.85 BPW |
Primary NanoPlus V3 quantized model with integrated MTP head |
mmproj-Q8_0.gguf |
610 MB |
8.50 BPW | Dedicated Q8_0 multimodal vision projector |
- Source: ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP.
- Architecture:
qwen35; 64 main layers (48 Gated DeltaNet + 16 full-attention) plus one MTP block. - Stored parameters: 27,320,697,856, including MTP tensors.
- Tensor count: 866, including the 15-tensor MTP head.
- Native context: 131,072 tokens (128K context).
- Calibration: ISTA-DASLab Qwen3.8-27B Calibrated Matrix (
imatrix-qwen3.8-27b.gguf). - Scope: Complete language GGUF and companion multimodal vision projector (
mmproj-Q8_0.gguf).
3. Native Context & Runtime Memory
The model's native context is 131,072 tokens. File size describes the stored weights; runtime memory also includes the KV cache, recurrent state, compute buffers, batch settings and any optional components. Choose context and offload settings according to the allocations reported by your runtime.
4. Recommended Configuration & Setup
llama.cpp CLI
llama-cli \
-m Swift-1.5-Qwen3.8-27B-Uncensored-MTP-VAL-APEX-I-NanoPlus-V3.gguf \
--mmproj mmproj-Q8_0.gguf \
-p "<|im_start|>user\nHello! Explain your architecture.<|im_end|>\n<|im_start|>assistant\n" \
-ngl 99 \
-c 8192 \
--temp 0.6 \
--top-p 0.95
llama-server (OpenAI-Compatible API with Native MTP Self-Speculation)
llama-server \
-m Swift-1.5-Qwen3.8-27B-Uncensored-MTP-VAL-APEX-I-NanoPlus-V3.gguf \
--mmproj mmproj-Q8_0.gguf \
--port 8080 \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
-ngl 99 \
-c 16384
5. Recommended Generation Parameters
| Setting | Thinking | Non-thinking |
|---|---|---|
| Temperature | 1.0 | 0.7 |
| Top-P | 0.95 | 0.80 |
| Top-K | 20 | 20 |
| Min-P | 0.0 | 0.0 |
| Presence penalty | 0.0 | 1.5 |
| Repetition penalty | 1.0 | 1.0 |
6. Native Chat Template & Reasoning
The native Qwen Jinja chat template is embedded in this GGUF. Use --jinja where supported and preserve the native formatting for messages, tool calls and thinking blocks. Reasoning controls depend on the backend's implementation of the template arguments.
7. Community Compute Fund & Priority Model Requests
All IsValorum quantizations will always remain completely free and open to the public without paywalls.
However, cloud GPU compute is expensive. If you find these builds valuable and would like to support the project or request a specific model architecture to be prioritized for the next MiniPlus/NanoPlus release, you can sponsor GPU compute time through Ko-fi:
(When supporting on Ko-fi, feel free to leave a note with your Hugging Face handle and the specific model you would like prioritized).
