← back to catalog · registered 2026-08-25 23:02

jaromer/zerodigest-Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF

jaromer Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jaromer%2Fzerodigest-Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 11
  • hub_downloads_all_time 707
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
707
229 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-25
Downloads over time
Now809→from28↑2,789%
029659188728 on Aug 26809 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf text-generation quantizer autoround architecture-aware mamba ssm multi-token-prediction mtp uncensored 27b base_model:JonathanColetti/Qwen3.8-27B-Uncensored

Related

Total size
97.5 GB
Files
11
Quantizations
1
Registered
2026-08-25 23:02
Last updated on HF
2026-08-25 22:26

Files by quantization

Auxiliary files 11 files 97.5 GB
Qwen3.8-27B-Uncensored-YMQ-XL.gguf 18.4 GB ef933f56 download
Qwen3.8-27B-Uncensored-YMQ-L.gguf 16.0 GB d7c02bba download
Qwen3.8-27B-Uncensored-YMQ-M.gguf 13.6 GB a82068c0 download
Qwen3.8-27B-Uncensored-YMQ-S-Pro.gguf 11.7 GB 6ea94df0 download
Qwen3.8-27B-Uncensored-YMQ-XS-Pro.gguf 10.2 GB 13aaf616 download
Qwen3.8-27B-Uncensored-YMQ-XXS.gguf 9.29 GB 06e57946 download
Qwen3.8-27B-Uncensored-YMQ-XXS-Pro.gguf 9.21 GB c7f9fd7c download
Qwen3.8-27B-Uncensored-YMQ-XXS-NO-MTP.gguf 9.09 GB 78fff92f download
logo.jpg 167 KB 53d427f3 download
README.md 11.5 KB bb8db25a download
.gitattributes 2.54 KB 47c2e4f1 download

README current version from Hugging Face


license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
library_name: gguf
tags:

  • text-generation
  • gguf
  • quantizer
  • autoround
  • architecture-aware
  • mamba
  • ssm
  • multi-token-prediction
  • mtp
  • uncensored
  • 27b

Qwen3.8-27B-Uncensored-YMQ-MTP-GGUF

Source Model: JonathanColetti/Qwen3.8-27B-Uncensored

⚖️ An Architecture-Aware, AutoRound-Inspired Mixed Precision Layout

This repository features advanced, custom architecture-aware quantizations of Qwen3.8-27B (Uncensored) processed directly from official raw BF16 source files using the custom YMQ-Compiler (v2.0) log-space framework.

These builds natively support parallel multi-token prediction (MTP) speculation engines and utilize high-context optimization parameters tailored for demanding code development API execution environments (such as RooCode/Aider).

ZeroDigest YMQ Logo


📊 Quantization Preset Tier Details

Preset Tier Total Size Target Usage / Memory VRAM Profile Cognitive Real-World Coding Quality
XXS-Pro ~9.3 GB ⚠️ Experimental Low-VRAM Sandbox The Ultra-Compact Frontier. Perplexity = 8.2084. Packs the 27B dense matrix into a sub-10GB footprint. In intense multi-stage agent workflows, it can occasionally trigger context amnesia or formatting loops, but features heavily insulated upper routing tracks to protect basic logical structures.
XXS ~9.8 GB Absolute VRAM Squeeze / 12GB Card Lifeline Massive structural quantization noise. Best restricted to low-context, single-turn instructions. Fits 12GB cards with context cache breathing room.
XS-Pro ~11.0 GB ⚡ Dedicated 12GB VRAM Champion High compression economy baseline. Optimized to prevent API degradation during deep context tasks.
S-Pro ~12.5 GB 💎 Premium 16GB Workstation Driver The Efficiency Miracle. Holds an elite 7.06 PPL for pristine conversational fluidness while using full Q5_K attention armor.
M ~14.0 GB The Ultimate Coding Sweet Spot (Recommended) Elite logical stability. Complete logic clarity. It crushes standard industry 4-bit alternatives.
L ~17.0 GB Premium Single-GPU Processing / Heavy workloads Near-lossless instruction formatting. Pristine multi-turn architecture safety.
XL ~19.0 GB Maximum VRAM Fill / No Compromises Mathematical saturation ceiling. Full precision logic tracks for massive multi-file codebase operations.

📉 Perplexity Evaluation Metrics (WikiText-2)

The following metrics demonstrate the mathematical quality preservation of the YMQ-Compiler log-space cluster analysis compared to standard linear quantization layouts. Tested natively via llama-perplexity over a 4096 context window using the official WikiText-2 test corpus.

Model Preset Variant File Size Perplexity Score (Lower is Better) Cognitive Calibration Verdict
XXS-Pro ~9.3 GB 8.2084 Experimental VRAM economy boundary cliff.
XXS ~9.8 GB 7.6538 Extreme VRAM economy boundary cliff.
XS-Pro ~11.0 GB 7.1665 Isolated task profile fallback baseline.
S-Pro ~12.5 GB 7.0687 Full Q5_K reasoning armor.
M (Recommended) ~14.0 GB 6.8176 🎯 The Golden Architectural Sweet Spot
L ~17.0 GB 6.9832 Minor cumulative network drift from background bloat.
XL ~19.0 GB 6.8329 Full mathematical saturation ceiling.

🚨 CRITICAL ARCHITECTURAL UPDATE

  • Legacy S Preset Deprecated: The older, standard S configuration (~12.2 GB) has been officially removed from the repository.

  • Upgrade to S-Pro (~12.5 GB): We have replaced it with the newly engineered S-Pro preset.

  • Legacy XS Preset Deprecated: The older, standard XS configuration (~11.0 GB) has been permanently removed from the repository tree.

  • Upgrade to XS-Pro (~10.5 GB): We have officially replaced it with the newly engineered XS-Pro preset. Score 7.1665 vs old XS score 8.1516

💡 The Multi-Tier Grid Breakthrough: Standard S vs. S-Pro

During intensive local workspace validation passes, our architecture-aware YMQ-Compiler successfully mapped out a radical new bit-allocation matrix. By splitting the layer distribution, we created a premium, high-fidelity alternative to our standard budget tier:

  • Standard S Preset (~12.2 GB): Perplexity = 8.0351. Features a balanced log-space gradient. Highly capable of handling single-turn scripts and quick edits. Successfully ingested a clean 91kb codebase chunk to resolve deep priority-ordered dictionary bugs natively.
  • S-Pro Preset (~12.5 GB): Perplexity = 7.0687. By aggressively compressing auxiliary tensor lanes down to Q2_K but raising the background baseline floor to IQ3_XXS, S-Pro eliminates a massive wave of background quantization noise—dropping paper perplexity by a massive ~1.0 point while only adding a few megabytes of file weight.

The practical result is a premium, low-overhead everyday driver for 16GB GPU setups. Backed by full Q5_K reasoning armor.

💡 The Uncensored Performance Breakthrough

Notice that the Uncensored M preset achieves an elite score of 6.8176, outperforming even the original base model's score (6.8413). This occurs because removing the artificial refusal safety layers allows the model's underlying Attention and Mamba SSM weights to predict text paths with absolute, unrestricted mathematical clarity.

By pairing JonathanColetti's pristine abliteration weights with the YMQ-Compiler's log-space gate insulation, this preset matches the raw reasoning power of the massive 19GB XL file while clawing back a clean 5 Gigabytes of VRAM overhead cache space for local RooCode/Aider coding loops!

💡 Engineering Notes on the XXS-Pro Layout

The XXS-Pro preset is a highly aggressive exploration pass utilizing an optimized mixed-precision architecture template:
HIGH="IQ3_XXS" (3.0 BPW) -> MID="IQ3_XXS" -> LOW="IQ2_S" (2.5 BPW) -> FLOOR="IQ2_XXS" (2.06 BPW).

By adding custom FLOOR_TARGET parameters, we aggressively crushed the auxiliary and background matrix noise to stay beneath a hard 9.5 GB memory limit. While this compression level introduces enough quantization noise to challenge complex multi-file edit loops, our log-space steering gate protection allows the model to retain surprisingly strong English language capabilities and shorter script tracking entirely within low-VRAM graphics memory buffers!


⚖️ YMQ vs. Uniform Quantization (The AutoRound Philosophy)

Standard quantization pipelines apply a blunt, uniform bit-depth across every single layer in a model. This wastes valuable VRAM on silent background layers while starving critical logic anchors of necessary precision.

The YMQ-Compiler implements a philosophy similar to advanced weight-tuning frameworks like Intel's AutoRound:

  • Targeted Bit Allocation: It strips bits away from low-leverage background tensors and automatically re-allocates that saved VRAM budget straight into full high-fidelity shields for the model's highest cognitive spikes and boundary pathways.
  • Instant Optimization: Instead of running heavy, days-long optimization training loops, YMQ achieves a highly accurate mixed-precision layout instantly by analyzing layer importance metrics in log-space.

The result is a custom mixed-precision portfolio that matches the low perplexity and high context stability of premium optimized quants (like AutoRound), while maintaining an ultra-lightweight, high-speed single-GPU cache footprint.


🛠️ The YMQ Compilation Architecture

Standard quantization pipelines treat network tensors like a flat dataset, applying destructive blanket low-bit compression to delicate tracking networks. The YMQ-Compiler solves high-context logic decay by parsing model files dynamically via an automated, multi-tiered protection matrix:

  1. Log-Space Gap Detection Clustering: Instead of flat percentage thresholds, the engine computes statistical cluster variances in log-space, successfully isolating intermediate logical reasoning spikes and elevating them to stable non-linear 4-bit (IQ4_XS) formats, while compressing idle fact-storage layers to aggressive 2-bit baselines.
  2. Fading Boundary Tapering: Recognizes the extreme fragility of initial token entry data vectors, forcing an input wave cushion (L00=IQ4_NL → L01=IQ4_XS → L02=IQ3_XXS) that gradually stabilizes parameters before hitting the fallback pools.
  3. Dedicated Gate Insulation: Hard-shields volatile parallel Transformer Multi-Head Attention and Mamba Linear State Space Model (SSM) routing paths, keeping context tracking perfectly noise-free.
  4. Asymmetric Vocabulary Shielding: Fixes tied-weight boundary errors by mapping the final logit classification exit heads to robust configurations to completely eliminate formatting loops and API tag leakage under deep contexts.
  5. Native Next-N Speculative Stripping: Processed with advanced pre-tokenizer stripping to ensure zero index offset drift or layer-shifting risks across hybrid configurations.

🚀 Recommended Runtime Parameters (llama.cpp / llama-server)

$./llama-server -m models/Qwen3.8-27B-Uncensored-YMQ-M.gguf -ctk q8_0 -ctv q4_0 --ctx-size 245760 --mmproj models/Qwen3.8-27B-Uncensored-vision-Q6_K.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 --timeout 36000 --checkpoint-min-step 2048 --ctx-checkpoints 4 \
  --n-predict -1 --temp 0.6 --top-p 0.95 --top-k 20 --repeat-penalty 1.05 --jinja -fa

🖼️ Vision Projection (--mmproj)

For multimodal vision support, pair these builds with one of the following projection files:

Variant File Size Notes
Full Precision (F16) Qwen3.8-27B-Uncensored-vision-f16.gguf ~928 GB Full-precision vision tower, native to the Uncensored abliteration weights. Maximum fidelity for image reasoning tasks.
Q6_K Qwen3.8-27B-Uncensored-vision-Q6_K.gguf (this repo) ~587 MB High-fidelity quantized vision tower. Excellent quality-to-size balance with minimal perceptible degradation.
Q4_K_S Qwen3.8-27B-Uncensored-vision-Q4_K_S.gguf (this repo) ~478 MB Compact vision projection for VRAM-constrained setups. Retains strong image understanding at reduced footprint.

Pass via --mmproj <path-to-file> in your llama-server invocation (see example above).


☕ Support & Future R&D

If the YMQ-Compiler builds saved your context window from collapsing or optimized your active development cycle speeds, consider buying a coffee to fund further low-level optimization research. Your support keeps the server nodes baking future model scales!

👉 Support ZeroDigest Research on ko-fi


📦 Source Framework & Automation Code

The compiler pipeline automation engine, setup thresholds, and structural mapping rules are open-source. To view the implementation details or compile your own custom models natively using this profile layout, visit the official development hub:

👉 GitHub: ZeroDigest / YMQ-Compiler

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Duplicate from zerodigest/Qwen3.8-27B-Uncensored-YMQ-MTP-GGUFe7e5f0511.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration