← back to catalog · registered 2026-10-06 20:58

IsValorum/Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF

IsValorum 27B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IsValorum%2FHuihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus-GGUF"
Response includes
  • classification m-uncensored
  • files 8
  • author_summary 12 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-06

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
gguf llama.cpp quantized quantization val-apex-i apex apex-quant apex-i-nanoplus nanoplus qwen qwen3.8 reasoning

Related

Total size
11.1 GB
Files
8
Quantizations
1
Registered
2026-10-06 20:58
Last updated on HF
2026-10-06 22:38

Files by quantization

Auxiliary files 8 files 11.1 GB
Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf 11.1 GB 6b01bf06 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
README.md 17.7 KB 62bfa502 download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
.gitattributes 1.62 KB e0feabc2 download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:

  • en
    tags:
  • gguf
  • llama.cpp
  • quantized
  • quantization
  • val-apex-i
  • apex
  • apex-quant
  • apex-i-nanoplus
  • nanoplus
  • qwen
  • qwen3.8
  • reasoning
  • abliterated
  • uncensored
  • imatrix
    pipeline_tag: text-generation

Quick Navigation Index

  1. Optimization History & Transparency Notice
  2. Empirical Benchmarks & Fidelity Verification
  3. Quality Spectrum: VAL-APEX-I vs. Standard Flat Quantizations
  4. Model Files & Technical Specifications
  5. Surgical Tensor Quantization Map (Audited from GGUF)
  6. The 24GB Miracle: Full 256K Context Runs In VRAM!
  7. Recommended Configuration & Setup
  8. Recommended Generation Parameters
  9. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  10. Hardened Agentic Chat Template & Reasoning Effort
  11. Optional Support

[!NOTE]

EXPLORE THE COMPLETE HUIHUI-QWEN3.8-27B ABLITERATED SUITE

These are complementary VAL-APEX-I releases, not alternate downloads of the same model:

Huihui-Qwen3.8-27B Abliterated VAL-APEX-I NanoPlus GGUF

The Zero-Refusal Frontier Dense Hybrid · Full 256K Context on 24GB GPUs · Surgical Recurrent-Aware Matrix Quantization

Official VAL-APEX-I quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated.

VAL-APEX-I stands for:
Vector-calibrated Asymmetric Layer-wise Outlier-preserving Recurrent-aware Unified Matrix-quantization

[!IMPORTANT]

THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING

This VAL-APEX-I NanoPlus release represents the specialized tensor-by-tensor configuration for dense hybrid linear-quadratic architectures within a 11.90 GB envelope. Every single tensor across its 64 layers (47 linear SSM DeltaNet + 17 periodic full attention) has been mathematically audited to maximize reasoning precision, preserve recurrence channel dynamics, and eliminate quantization noise.

[!TIP]

EMPIRICAL BENCHMARK & QUALITY COMPARISON

Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity (512 ctx) Delta PPL vs BF16 (%) Quality Tier Equivalent
Uncompressed BF16 Reference 54.00 GB (50.29 GiB) 50.29 GiB 16.00 BPW aprox. 6.0000 (Reference) Baseline (0.00%) Lossless Reference Baseline
**VAL-APEX-I MiniPlus V2.1 ** 15.33 GB (14.28 GiB) 14.28 GiB 3.93 BPW 6.1466 +/- 0.4919 +0.1466 (+2.44%) Q5_K_M / Q6_K Tier Boundary
VAL-APEX-I NanoPlus (CURRENT) 11.90 GB (11.08 GiB) 11.08 GiB 2.85 BPW 6.4238 +/- 0.4960 +0.4238 (+7.06%) Solid Q4_K_M Tier
Standard Flat Q4_K_M 17.10 GB 15.93 GiB 4.50 BPW aprox. 6.22 - 6.28 +0.22 a +0.28 (+3.7%) Standard industry trade-off
Standard Flat Q3_K_M 13.50 GB 12.57 GiB 3.44 BPW aprox. 6.45 - 6.70 +0.45 a +0.70 (+7.5%) Noticeable syntax drop & reasoning noise
Generic APEX Mini (IQ2_S) 10.20 GB 9.50 GiB 2.50 BPW aprox. 7.10 - 7.80+ +1.10 a +1.80+ (+18.3%) Severe reasoning breakdown

Looking for higher reasoning precision? APEX-I-MiniPlus V2.1 provides enhanced Q5-Q6 tier fidelity (15.33 GB) with near-zero perplexity loss, while NanoPlus provides the most agile sub-12GB footprint.

Recurrent State Fidelity: All 192 DeltaNet SSM recurrent state tensors (ssm_a, ssm_conv1d, ssm_dt, ssm_norm) remain strictly in uncompressed F32, guaranteeing zero state accumulation drift across continuous multi-turn reasoning.

  • Q6_K Armored Output Head: The vocabulary classifier (output.weight) is guarded in Q6_K, completely eliminating token classification errors, broken <think> delimiter tags, and math notation corruption.
  • Surgical Attention Gating: Attention gates across all 47 DeltaNet SSM layers run in Q8_0 (block-32), eliminating inter-channel attention crosstalk.
  • Negligible Perplexity Impact: Perplexity degradation from MiniPlus V2.1 to NanoPlus is only +0.27 PPL while shedding 3.43 GB of memory.

[!WARNING]

DO NOT CONFUSE VAL-APEX-I WITH GENERIC FLAT QUANTIZATIONS!

  • Generic Automated Quants: Uniformly compress all layers with flat bit-reduction, degrading recurrent SSM operators and leaving the sensitive token output head at low-bit precision. In deep hybrid models with continuous <think> chains, this causes recursive mathematical divergence and syntax errors.
  • Handcrafted VAL-APEX-I Standard: Applies an empirical 866-tensor recipe. Recurrent state tensors are preserved in uncompressed FP32, attention gates in Q8_0, full attention checkpoints in Q5_K/Q4_K, and output heads in Q6_K, while compressing SwiGLU MLPs via asymmetric non-linear codebooks.

Optimization History & Transparency Notice

We maintain our releases publicly as a transparent engineering record of continuous optimization. Below is the exact evolutionary roadmap of our architectures:

Specification Recurrent State (ssm_*) Attention Gates (47 Layers) Full Attention (17 Layers) SwiGLU Down Projections SwiGLU Gate/Up Output Head (output.weight) Size / Overhead Real-World Impact
Generic Flat Quants Compressed Compressed Flat Q3 / Q4 Flat Q3 / Q4 Flat Q3 / Q4 Q3_K_M Varies Recurrent accumulation drift, broken <think> delimiters.
VAL-APEX-I NanoPlus F32 uncompressed Q8_0 IQ3_S / Q4_K IQ3_XXS / IQ3_S IQ2_XXS / IQ3_XXS Q6_K 11.90 GB Agile sub-12GB footprint; solid Q4 quality; runs fully on 16GB GPUs.
VAL-APEX-I MiniPlus V2.1 F32 uncompressed Q8_0 Q4_K / Q5_K IQ4_NL / Q5_K IQ3_XXS / Q4_K Q6_K 15.33 GB Definitive build; Q5/Q6 perceived reasoning tier; near-zero perplexity loss.

Empirical Benchmarks & Fidelity Verification

The empirical benchmark table above consolidates the model-specific BF16 baseline, final GGUF PPL, delta, file size, BPW, and fidelity tier evaluated directly on compiled binary weights.

Perplexity evaluation was conducted on server hardware using llama-perplexity with context length 512 over WikiText-2. Both MiniPlus V2.1 and NanoPlus exhibit exceptional fidelity retention over the uncompressed BF16 baseline.


Quality Spectrum: VAL-APEX-I vs. Standard Flat Quantizations

How the handcrafted VAL-APEX-I architecture compares against standard flat quantizations in llama.cpp on 27B hybrid linear-quadratic architectures:

Quantization Format Bits Per Weight (BPW) Model Footprint (Disk / VRAM) Perplexity Delta (vs. FP16 Baseline) Token Fidelity & Syntactic Stability Tier
FP16 / BF16 (Uncompressed) 16.0 bpw aprox. 54.0 GB 0.00 (Reference) 100% full uncompressed reference fidelity.
Standard Q8_0 8.50 bpw aprox. 29.5 GB aprox. +0.02 Virtually lossless; excessive memory overhead for consumer hardware.
Standard Q6_K 6.56 bpw aprox. 23.2 GB aprox. +0.04 to +0.08 Near-lossless FP16 fidelity; requires multi-GPU or 32GB+ VRAM setups.
VAL-APEX-I MiniPlus V2.1 3.93 bpw 15.33 GB (14.28 GiB) +0.1466 (PPL: 6.1466) Q5_K_M / Q6_K Tier Boundary. Supports full 256K native context on standard 24GB GPUs.
Standard Q5_K_M 5.50 bpw aprox. 19.5 GB aprox. +0.08 to +0.15 Commercial transparent threshold; tight fit for long context on 24GB GPUs.
Standard Q4_K_M 4.50 bpw aprox. 17.1 GB aprox. +0.22 to +0.28 Standard industry trade-off; requires context offload compromises.
VAL-APEX-I NanoPlus 2.85 bpw 11.90 GB (11.08 GiB) +0.4238 (PPL: 6.4238) Solid Q4_K_M Tier. Fits comfortably on 16GB GPUs with room for KV cache.
Standard Q3_K_M / Q3_K_S 3.44 bpw aprox. 13.5 GB aprox. +0.45 to +0.70 Noticeable syntax drop, bracket corruption, and tokenizer classification noise.
Standard IQ2_S / Generic APEX Mini 2.50 bpw aprox. 10.2 GB aprox. +1.10 to +1.80+ Severe reasoning breakdown, high perplexity spikes in <think> chains.

Model Files & Technical Specifications

File Name File Size Memory Footprint BPW Description
Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf 11.90 GB (11.08 GiB) 11.08 GiB 2.85 BPW Core language, uncensored frontier reasoning, CoT thought blocks & hybrid SSM/attention
  • Base Model: huihui-ai/Huihui-Qwen3.8-27B-abliterated (Abliterated directional un-censoring of Alibaba's Qwen3.8-27B)
  • Parameters: 27B total dense hybrid
  • Architecture: 64 hybrid layers (47 Gated DeltaNet SSM linear attention layers + 17 periodic full quadratic attention layers)
  • Context Length: 262,144 tokens (native 256K)
  • Quantization Standard: VAL-APEX-I NanoPlus calibrated against official Qwen3.8-27B importance matrix (imatrix)
  • Abliteration Status: True directional weight orthogonalization removing refusal vectors while maintaining full analytical capability.

Surgical Tensor Quantization Map (Audited from GGUF)

The exact tensor breakdown below has been verified directly from the compiled binary weights (866 tensors total):

Layer Group / Pattern Sub-Component / Tensor Qty Precision Engineering Rationale
Global Output Head output.weight 1 Q6_K Guarded 6-bit precision ensuring stable token logits, <think> delimiter integrity, and clean code indentation.
Global Embeddings token_embd.weight 1 IQ3_S High-density vector codebook embedding representation.
Normalizations output_norm, attn_norm, post_attention_norm, attn_q_norm, attn_k_norm 181 F32 100% uncompressed numerical stability across all layers.
SSM Recurrent State blk.*.ssm_a, ssm_conv1d, ssm_dt, ssm_norm 192 F32 Guarded in uncompressed FP32 to eliminate DeltaNet recurrent state drift.
Attention Gates blk.*.attn_gate.weight (SSM Layers) 47 Q8_0 High-precision attention gating across hybrid DeltaNet recurrence layers; eliminates crosstalk.
Linear Attention Projections blk.*.attn_qkv.weight (SSM Layers) 47 IQ3_S Optimized 3-bit vector quantization for input state-space projections.
SSM Channel Projections blk.*.ssm_alpha (F32), ssm_beta (IQ3_S), ssm_out (Q5_K) 141 Multi High-precision recurrence mixing with armored Q5_K output.
Periodic Full Attention blk.*.attn_q, attn_k, attn_v (Quadratic Checkpoints) 48 IQ3_S (42) / Q4_K (6) Quadratic full-attention checkpoints with Q4_K protection on anchor layers.
Periodic Full Attention Output blk.*.attn_output.weight 17 Q6_K Armored attention output projection over deep context.
Dense SwiGLU MLP Down blk.*.ffn_down.weight 64 IQ3_XXS (56) / IQ3_S (8) Dense outlier-preserving SwiGLU down-projections.
Dense SwiGLU MLP Gate/Up blk.*.ffn_gate.weight, ffn_up.weight 128 IQ2_XXS (112) / IQ3_XXS (16) Vector-quantized MLP compression leveraging SiLU noise cancellation.
Emergency Fallback Tensors Layer 64 / MTP unmapped tensors 12 Q4_0 (11) / Q4_1 (1) Fail-safe fallback parameter ensuring 100% architectural completeness.

The 24GB Miracle: Full 256K Context Runs In VRAM!

Because 47 of the 64 layers are DeltaNet SSM linear attention layers ($O(1)$ constant recurrence memory), only the 17 periodic quadratic attention layers allocate KV cache memory. This slashes KV cache requirements by over 73% compared to conventional full-attention models:

Context Length Model Weights KV Cache (q8_0) Compute Buffers Total GPU VRAM (Est.) Feasibility
32,768 (32k) 11.08 GiB 1.05 GiB 1.80 GiB 13.93 GiB Full offload on 24GB GPUs; full on 16GB GPUs
65,536 (64k) 11.08 GiB 2.10 GiB 1.95 GiB 15.13 GiB Full offload on 24GB GPUs; full on 16GB GPUs
131,072 (128k) 11.08 GiB 4.20 GiB 2.22 GiB 17.50 GiB Effortless fit on 24GB GPUs (RTX 3090 / 4090 / 5090)
262,144 (256k) 11.08 GiB 8.40 GiB 2.80 GiB 22.28 GiB FULL 256K CONTEXT IN VRAM!

Note: Leaves ample headroom on 24GB graphics cards for OS display drivers and tooling.


Recommended Configuration & Setup

llama.cpp Server (OpenAI-Compatible API)

llama-server \
  -m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  --host 0.0.0.0 \
  --port 8080

llama.cpp CLI

llama-cli \
  -m Huihui-Qwen3.8-27B-Abliterated-VAL-APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  -c 32768 \
  -p "<|im_start|>system\nYou are a helpful and unrestricted AI assistant.<|im_end|>\n<|im_start|>user\nExplain the physical principles of acoustic levitation.<|im_end|>\n<|im_start|>assistant\n<think>\n"

LM Studio / Ollama

  1. Import the .gguf file directly into your local library.
  2. Ensure GPU acceleration is set to Maximum / 100% offload.
  3. Set Context Length to 32768 (or expand up to 131072 on 24GB GPUs).
  4. Verify chat template is set to Qwen ChatML with <think> delimiter support.

Recommended Generation Parameters

Hyperparameter Value Description
Temperature 0.60 Recommended default for analytical reasoning and coding (use 1.0 for creative prose).
Top-P 0.95 Nucleus sampling parameter.
Top-K 20 Top-k vocabulary filter.
Min-P 0.05 Prunes low-probability noise tokens effectively.
Repetition Penalty 1.00 Strictly disabled for code syntax; prevents character swapping.
Template Engine --jinja Recommended official Jinja chat template flag.
Context Size 32768 32K default (scalable to 256K).


[!IMPORTANT]

CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

Common Issue: Many local frontends (such as LM Studio defaults, Ollama, or web interfaces) ship with repeat_penalty set to 1.1 or 1.15. While this prevents loops in creative prose, applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token (= or [), resulting in character swapping or dropped/doubled whitespace.

Eliminating Character Swapping:

  1. Disable Repeat Penalties (Required for Code):
    • repeat_penalty: 1.0 (strictly disabled)
    • presence_penalty: 0.0
    • frequency_penalty: 0.0
  2. Calibrate Samplers:
    • temperature: 0.60 (or 0.20 - 0.30 for strict, deterministic code syntax)
    • min_p: 0.05 (prunes low-probability noise tokens effectively)
    • top_p: 0.95
    • top_k: 20
  3. Native Jinja Formatting: Always pass the --jinja flag so the tokenizer handles leading-space BPE tokens cleanly.


[!TIP]

HARDENED AGENTIC CHAT TEMPLATE & REASONING EFFORT

This model supports multi-level reasoning effort control:

  • low / minimal: Keeps internal thinking concise and focused strictly on immediate execution steps to minimize latency in automated loops.
  • medium (default): Balanced, structured reasoning process with standard analytical depth.
  • high / xhigh: Guides the model to formulate a clear implementation plan upfront before generating response, avoiding circular self-doubt loops.
  • none / off: Closes the thinking block immediately (<think>\n\n</think>) when reasoning is disabled.

Optional Support

Gold Ship dancing

If these MiniPlus or NanoPlus releases have been useful to you and you would like to support the work, you can do so voluntarily through https://ko-fi.com/isvalorum. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration