← back to catalog · registered 2026-10-01 05:58

IsValorum/Xing4.0-29B-A4B-APEX-I-NanoPlus-Abliterated-GGUF

IsValorum 29B GGUF MoE second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IsValorum%2FXing4.0-29B-A4B-APEX-I-NanoPlus-Abliterated-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf llama.cpp quantized quantization apex apex-quant apex-i-nanoplus nanoplus custom-quantization imatrix abliterated uncensored

Related

Total size
10.9 GB
Files
3
Quantizations
1
Registered
2026-10-01 05:58
Last updated on HF
2026-10-01 06:23

Files by quantization

Auxiliary files 3 files 10.9 GB
Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf 10.9 GB ae562a11 download
README.md 15.0 KB c8435bf3 download
.gitattributes 1.57 KB 7dbd747c download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:

  • en
  • zh
    pipeline_tag: text-generation
    tags:
  • gguf
  • llama.cpp
  • quantized
  • quantization
  • apex
  • apex-quant
  • apex-i-nanoplus
  • nanoplus
  • abliterated
  • uncensored
  • moe
  • mla
  • hyper-connections
  • reasoning
  • text-generation
  • coding
  • xing4_0

Quick Navigation Index

  1. Architecture & Methodology Overview
  2. Empirical Benchmarks & Fidelity Verification
  3. Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
  4. Model Files & Technical Specifications
  5. Surgical Tensor Quantization Map (Audited from GGUF)
  6. Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
  7. Inference Quickstart
  8. Recommended Generation Parameters
  9. CRITICAL: Coding Syntax & Repeat Penalty Advisory
  10. Optional Support

Xing4.0-29B-A4B APEX-I-NanoPlus Abliterated GGUF

The Next-Generation Frontier MoE - Extreme 11.7 GB Footprint - Fast System RAM Streaming & Massive Context on 16GB VRAM

[!IMPORTANT]

THE DEFINITIVE SPECIFICATION IN THE 11 to 12 GB CEILING (STREAMLINED ARCHITECTURE)

This APEX-I-NanoPlus release represents the specialized ultra-compact configuration for sparse Mixture-of-Experts quantization within an 11 to 12 GB envelope. It compresses the 40-layer backbone into a compact 11.74 GB (10.94 GiB) file, enabling full offload with large context on consumer 16GB VRAM cards and fast memory bandwidth streaming from system RAM.

[!TIP]

EMPIRICAL BENCHMARK & QUALITY COMPARISON

Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 Base Quality Tier Equivalent
Unquantized BF16 Base 62.40 GB 58.11 GiB 16.00 BPW 7.3060 +/- 0.18633 Baseline (0.0000) Full precision reference
APEX-I-MiniPlus V2.1 13.76 GB 12.82 GiB 3.42 BPW 7.7238 +/- 0.19671 +0.4178 (+5.71%) Q5_K_M tier (bordering Q6_K)
APEX-I-NanoPlus (CURRENT) 11.74 GB 10.94 GiB 2.92 BPW 8.3197 +/- 0.21528 +1.0137 (+13.87%) Solid Q4_K_M / Q4_K_S Tier

Looking for higher precision? Xing4.0-29B-A4B APEX-I-MiniPlus-V2.1 offers the full 13.76 GB (12.82 GiB / 3.42 BPW) release, delivering Q5_K_M tier (bordering Q6_K) fidelity.

Routing: 100% of recipe-designated ffn_gate_inp router matrices remain in uncompressed F32, guaranteeing zero routing drift.

  • Solid Q4_K_M Tier in Language Modeling: WikiText-2 perplexity preserves 4-bit distributional fidelity (8.3197 +/- 0.21528) across standard generation in an ultra-lean footprint.
  • Solid Q4_K Tier in MoE Foundation Knowledge: Shared expert tensors (shexp) run in Q4_K, keeping core knowledge intact across 100% of tokens.
  • Protected MLA Geometry: Latent attention projections in Q4_K and the output head in Q6_K prevent attention collapse.

[!WARNING]

DO NOT CONFUSE APEX-I-NANOPLUS WITH UNCALIBRATED SUB-3-BIT QUANTS!

Never confuse handcrafted APEX-I-NanoPlus builds with uncalibrated homogeneous quantizations:

  • Uncalibrated Homogeneous Quants: Uniformly crush all core MoE experts down to aggressive 2-bit codebooks without importance calibration, leave the sensitive token output head unarmored, and compress latent attention projections down to sub-3-bit. In deep reasoning models with MLA, this triggers severe perplexity spikes, routing misdirection, and broken syntax.
  • Handcrafted APEX-I-NanoPlus: Applies a surgical tensor-by-tensor architecture that preserves 100% of expert routing matrices in uncompressed F32 (zero router drift), armors the token output head in high-precision Q6_K, safeguards attention projections in Q4_K, locks shared experts in Q4_K, fortifies down-projection experts in IQ3_XXS, and restricts 2-bit compression strictly to redundant gating/up projections guided by the official imatrix.

[!TIP]

SYSTEM RAM INFERENCE: FULL OR PARTIAL

This APEX-I-NanoPlus release is designed for full or partial system-RAM inference. Depending on the processor, memory bandwidth, and DDR4/DDR5 configuration, generation can range from 20 to 45 tok/s. With partial GPU offload, systems that cannot fit large contexts entirely in VRAM can place the remaining model and context load in system RAM, maintaining stable, responsive generation at longer context lengths.



Architecture & Methodology Overview

The APEX-I-NanoPlus release applies an ultra-compact tensor-by-tensor allocation engineered specifically for sparse Mixture-of-Experts architectures with Multi-head Latent Attention (MLA), prioritizing execution speed and extreme memory efficiency:

  • Zero Router Drift (F32): All routing matrices (ffn_gate_inp and ffn_gate_inp_shexp) are preserved in full 32-bit floating point, guaranteeing 100% routing fidelity matching the uncompressed base model.
  • Shared Foundation Backbone (Q4_K): The shared expert (shexp) active across 100% of tokens is armored in linear Q4_K across all 40 layers, securing foundational domain knowledge and smooth CPU/GPU streaming.
  • MLA Latent Projection Armor (Q4_K / Q6_K): Latent attention projections and output heads are preserved in Q4_K and Q6_K, eliminating attention collapse while keeping memory overhead minimal.
  • Fortified Residual Down-Projections (IQ3_XXS): All ffn_down_exps matrices are maintained at 3.06 BPW to protect the residual stream, while redundant SwiGLU gating and up-projections (ffn_gate_exps, ffn_up_exps) are compressed to 2.50 BPW (IQ2_S) using an authentic 820-chunk importance matrix.
  • Speculative Head Preservation (blk.40): The Multi-Token Prediction (MTP) draft head is quantized to Q4_K, enabling speculative acceleration in compatible runtimes.


Empirical Benchmarks & Fidelity Verification

Evaluated directly on the compiled GGUF binary using llama-perplexity over the standard WikiText-2 benchmark:

Metric Baseline (BF16 Unquantized) APEX-I-NanoPlus (Current) Evaluation Setup
WikiText-2 Perplexity 7.3060 +/- 0.18633 8.3197 +/- 0.21528 Measured directly on GGUF binary (2048 ctx, 512 batch, 10 chunks)
Delta vs BF16 Base 0.0000 (Reference) +1.0137 (+13.87%) Recovers full model coherence; operates at Solid Q4_K_M / Q4_K_S tier
Model Footprint 62.40 GB (58.11 GiB) 11.74 GB (10.94 GiB) approx. 81.2% reduction in storage and memory bandwidth
Active Memory Overhead >65 GB VRAM <14 GB VRAM Runs completely offloaded on a single RTX 4070 Ti / 4080 / 3090 GPU


Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations

How the handcrafted APEX-I-NanoPlus architecture compares against standard flat quantizations in llama.cpp on 29B Mixture-of-Experts architectures:

Quantization Format Bits Per Weight (BPW) Model Footprint (Disk / VRAM) Perplexity Delta (vs. BF16 Baseline) Token Fidelity & Syntactic Stability Tier
BF16 Base (Uncompressed) 16.00 bpw 62.40 GB 0.0000 (Reference) 100% full uncompressed reference fidelity.
Standard Q8_0 8.50 bpw approx. 33.2 GB approx. +0.0100 Virtually lossless; excessive memory overhead for consumer hardware.
Standard Q6_K 6.56 bpw approx. 26.0 GB approx. +0.0500 to +0.1000 Near-lossless BF16 fidelity; requires 32GB+ VRAM setups.
APEX-I-MiniPlus V2.1 3.42 bpw 13.76 GB (12.82 GiB) +0.4178 (PPL: 7.7238) Maximum fidelity Q5_K_M / Q6_K tier. Full native context on consumer GPUs.
APEX-I-NanoPlus (IsValorum) 2.92 bpw 11.74 GB (10.94 GiB) +1.0137 (PPL: 8.3197 +/- 0.21528) Solid Q4_K_M fidelity tier at only 11.74 GB (81.2% weight reduction). Enables full offload on 16GB GPUs with 32K context and zero AVX2 CPU stalls.
Standard Q4_K_M 4.50 bpw approx. 18.0 GB approx. +0.3500 to +0.5500 Standard industry trade-off; cannot fit in 16GB VRAM with context.
Standard Q3_K_M 3.44 bpw approx. 14.5 GB approx. +1.1000 to +1.8000 Noticeable syntax drop, bracket corruption, and tokenizer classification noise.
Standard IQ2_S / Generic Sub-3-Bit 2.50 bpw approx. 11.5 GB approx. +1.8000 to +4.0000+ Severe reasoning breakdown, high perplexity spikes in multi-step logic.

Model Files & Technical Specifications

File Name File Size Memory Footprint BPW Description
Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf 11.74 GB (10.94 GiB) 10.94 GiB 2.92 BPW Ultra-lean uncensored frontier MoE with MLA and Hyper-Connections
  • Base Model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
  • Parameters: 29.1B total (approx. 4.0B active per token)
  • Architecture: 40 layers, 64 routed MoE experts (4 active per token) + 1 shared expert + Multi-head Latent Attention (MLA)
  • Context Length: Native 131,072 tokens (128K)
  • Refusal Removal: Orthogonal activation steering via Heretic Directional Ablation


Surgical Tensor Quantization Map (Audited from GGUF)

The exact tensor breakdown below has been verified directly from the compiled binary weights:

Layer Group Sub-Component / Tensor Precision Engineering Rationale
Speculative Draft Head blk.40.* (All Tensors in MTP Block) Q4_K Surgical non-imatrix override bypassing imatrix lookup issues; enables speculative draft acceleration.
Global Output Head output.weight Q6_K Preserves near-FP16 token classification; eliminates syntax errors and bracket drops.
Global Embeddings token_embd.weight Q4_K High-fidelity vocabulary embedding representation.
All Normalizations output_norm, attn_*_norm, ffn_norm F32 100% uncompressed numerical stability across all layers.
Expert Routers blk.*.ffn_gate_inp, ffn_gate_inp_shexp F32 100% uncompressed routing fidelity across 64 experts; zero router drift.
Shared Foundation Experts blk.*.ffn_{gate,down,up}_shexp (All 40 Layers) Q4_K Foundation knowledge backbone active on 100% of tokens; protected in linear Q4_K for fast streaming.
Attention Output Heads blk.*.attn_output (All 40 Layers) Q6_K Armored attention output projection over deep context; prevents attention collapse.
Latent Attention Projections blk.*.attn_q_a, attn_q_b, attn_v_b, attn_kv_a_mqa Q4_K Preserves Multi-head Latent Attention (MLA) reconstruction geometry.
MoE Residual Down-Proj blk.*.ffn_down_exps (All 40 Layers) IQ3_XXS Fortified 3.06 bpw residual stream; preserves core mathematical and reasoning capacity.
MoE SwiGLU Gating & Up blk.*.ffn_gate_exps, blk.*.ffn_up_exps IQ2_S High-precision 2.50 bpw compression for redundant gating, guided by imatrix calibration.


Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)

Empirically verified in llama.cpp:

Hardware Target Offload Mode Generation Speed (Est.) Highlights
NVIDIA RTX 5080 / 5090 (Blackwell) Full GPU (-ngl 99) 130 to 175+ tok/s Extremely fast throughput with massive context buffers in VRAM
NVIDIA RTX 4090 / 3090 (24GB VRAM) Full GPU (-ngl 99) 85 to 120+ tok/s Full offload with headroom for maximum context
NVIDIA RTX 4080 / 4070 Ti Super (16GB VRAM) Full GPU (-ngl 99) 65 to 95+ tok/s Full 32K context completely inside 16GB VRAM
System RAM (DDR4 / DDR5) Hybrid / CPU Offload 25 to 50 tok/s Zero AVX2 CPU stalls; efficient streaming from system RAM


Inference Quickstart

1. llama-cli (Console Generation)

llama-cli \
  -m Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf \
  -ngl 99 \
  -c 16384 \
  --temp 0.7 \
  --top-p 0.90 \
  -p "<|im_start|>system\nYou are an expert uncensored assistant.<|im_end|>\n<|im_start|>user\nWrite a complete Rust implementation of a concurrent ring buffer.<|im_end|>\n<|im_start|>assistant\n"

2. llama-server (OpenAI-Compatible API)

llama-server \
  -m Xing4.0-29B-A4B.APEX-I-NanoPlus-Abliterated.gguf \
  --port 8080 \
  --host 0.0.0.0 \
  -ngl 99 \
  -c 32768 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0


Recommended Generation Parameters

Hyperparameter Value Description
Temperature 0.70 Balanced creativity and analytical precision
Top-P 0.90 Standard nucleus threshold
Top-K 40 Retains vocabulary diversity while trimming outliers
Repeat Penalty 1.00 to 1.05 Keep near 1.0 for coding workflows


[!IMPORTANT]

CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

Common Issue: Many local frontends ship with repeat_penalty set to 1.1 or 1.15. Applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token, resulting in character swapping or dropped whitespace.

Recommendation: When executing code generation or debugging tasks, always set repeat_penalty to 1.00 or maximum 1.02.



Optional Support

If you find this surgical quantization valuable for your production workflows, local research, or agentic coding environments, consider starring the repository and following the IsValorum Hugging Face profile for continuous updates and frontier releases.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-01Polish official APEX-I-NanoPlus release documentation4a19ea714.7 KB
    Loading...
  2. 2026-10-01Update definitive APEX-I-NanoPlus documentation and audited benchmarks60d502b14.3 KB
    Loading...
  3. 2026-10-01Upload README.md with huggingface_hubd6e09192.3 KB
    Loading...
  4. 2026-10-01Upload README.md with huggingface_hubb576aaa10.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.