← back to catalog · registered 2026-10-01 05:58

IsValorum/Xing4.0-29B-A4B-APEX-I-MiniPlus-V2.1-Abliterated-GGUF

IsValorum 29B GGUF MoE second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IsValorum%2FXing4.0-29B-A4B-APEX-I-MiniPlus-V2.1-Abliterated-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf llama.cpp quantized quantization apex apex-quant apex-i-miniplus v2.1 custom-quantization imatrix abliterated uncensored

Related

Total size
12.8 GB
Files
3
Quantizations
1
Registered
2026-10-01 05:58
Last updated on HF
2026-10-01 06:23

Files by quantization

Auxiliary files 3 files 12.8 GB
Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf 12.8 GB 95050149 download
README.md 16.1 KB 2ef80ad9 download
.gitattributes 1.57 KB 031b7630 download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
base_model_relation: quantized
quantized_by: IsValorum
library_name: gguf
license: apache-2.0
language:

  • en
  • zh
    pipeline_tag: text-generation
    tags:
  • gguf
  • llama.cpp
  • quantized
  • quantization
  • apex
  • apex-quant
  • apex-i-miniplus
  • v2.1
  • abliterated
  • uncensored
  • moe
  • mla
  • hyper-connections
  • reasoning
  • text-generation
  • coding
  • xing4_0

Quick Navigation Index

  1. Architecture & Methodology Overview
  2. Empirical Benchmarks & Fidelity Verification
  3. Quality Spectrum: APEX-I-MiniPlus V2.1 vs. Standard Flat Quantizations
  4. Model Files & Technical Specifications
  5. Surgical Tensor Quantization Map (Audited from GGUF)
  6. Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)
  7. The 16GB / 24GB Architecture: Full Context Runs In VRAM
  8. Recommended Configuration & Setup
  9. Recommended Generation Parameters
  10. CRITICAL: Coding Syntax & Repeat Penalty Advisory
  11. Optional Support

Xing4.0-29B-A4B APEX-I-MiniPlus-V2.1 Abliterated GGUF

The Next-Generation Frontier MoE - Uncensored & Unconstrained - MLA + Hyper-Connections - Efficient System RAM Offload

[!IMPORTANT]

THE DEFINITIVE SPECIFICATION IN THE 13 to 14 GB CEILING

This APEX-I-MiniPlus-V2.1 release represents the specialized tensor-by-tensor configuration for sparse Mixture-of-Experts quantization within a 13 to 14 GB envelope. Every tensor across its 40 layers, 64 routed experts (4 active per token), dedicated shared expert, and integrated MTP speculative block has been mathematically allocated to maximize reasoning precision, preserve routing behavior, and prevent avoidable CPU dequantization stalls.

[!TIP]

EMPIRICAL BENCHMARK & QUALITY COMPARISON

Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 Base Quality Tier Equivalent
Unquantized BF16 Base 62.40 GB 58.11 GiB 16.00 BPW 7.3060 +/- 0.18633 Baseline (0.0000) Full precision reference
APEX-I-MiniPlus V2.1 (CURRENT) 13.76 GB 12.82 GiB 3.42 BPW 7.7238 +/- 0.19671 +0.4178 (+5.71%) Q5_K_M tier (bordering Q6_K)
APEX-I-NanoPlus 11.74 GB 10.94 GiB 2.92 BPW 8.3197 +/- 0.21528 +1.0137 (+13.87%) Solid Q4_K_M / Q4_K_S tier

Prefer the ultra-compact Xing edition? APEX-I-NanoPlus compresses the same Xing family down to 11.74 GB (10.94 GiB / 2.92 BPW) for extreme 16GB VRAM and RAM streaming at Solid Q4_K_M tier quality, while MiniPlus V2.1 retains the higher 13.76 GB (12.82 GiB / 3.42 BPW) Q5_K_M tier fidelity.

Routing: 100% of recipe-designated ffn_gate_inp router matrices remain in uncompressed F32, guaranteeing zero routing drift.

  • Q6_K-Bordering Tier in Reasoning & Logic: WikiText-2 perplexity delta is exceptionally low (+0.4178 / +5.71% vs. BF16), placing overall semantic representation at the boundary of a 26 GB Q6_K build within an agile approx. 13.76 GB footprint.
  • Q5_K Tier in MoE Foundation Knowledge: All shared expert tensors (shexp) across all 40 layers run in uncompressed Q5_K, keeping core knowledge intact across 100% of tokens.
  • Solid MLA & Output Armor: Latent attention projections in Q4_K/Q5_K and the output head in Q6_K prevent token drift, attention collapse, and formatting errors.

[!WARNING]

DO NOT CONFUSE APEX-I-MINIPLUS WITH UNCALIBRATED SUB-3-BIT QUANTS!

Never confuse handcrafted APEX-I-MiniPlus builds with uncalibrated homogeneous quantizations:

  • Uncalibrated Homogeneous Quants: Uniformly compress all core MoE experts down to aggressive 2-bit codebooks without importance calibration, leave the sensitive token output head unarmored, and compress latent attention projections down to sub-3-bit. In deep reasoning models with MLA, this triggers severe perplexity spikes, routing misdirection, and broken syntax.
  • Handcrafted APEX-I-MiniPlus V2.1: Applies a custom tensor-by-tensor architecture that preserves specified router gates in uncompressed F32, armors the token output head in high-precision Q6_K, safeguards attention projections in Q4_K/Q5_K, locks shared experts in Q5_K, and calibrates core reasoning experts with an authentic 820-chunk importance matrix (imatrix).

[!TIP]

SYSTEM RAM INFERENCE: FULL OR PARTIAL

This APEX-I-MiniPlus release is designed for full or partial system-RAM inference. Depending on the processor, memory bandwidth, and DDR4/DDR5 configuration, generation can range from 20 to 45 tok/s. With partial GPU offload, systems that cannot fit large contexts entirely in VRAM can place the remaining model and context load in system RAM, maintaining stable, responsive generation at longer context lengths.



Architecture & Methodology Overview

The APEX-I-MiniPlus V2.1 release applies a mathematically audited, tensor-by-tensor allocation engineered specifically for sparse Mixture-of-Experts architectures with Multi-head Latent Attention (MLA):

  • Zero Router Drift (F32): All 64 routing matrices (ffn_gate_inp and ffn_gate_inp_shexp) are retained in full 32-bit floating point, ensuring exact expert selection identical to the uncompressed base model.
  • Shared Foundation Backbone (Q5_K): The shared expert (shexp) active across 100% of tokens is armored in high-precision linear Q5_K across all 40 layers, securing domain knowledge and syntactic consistency.
  • MLA Latent Projection Armor (Q4_K / Q5_K): Latent attention projections (attn_q_a, attn_q_b, attn_v_b, attn_kv_a_mqa) and output heads (attn_output in Q6_K) are safeguarded against the geometric distortion that breaks standard sub-4-bit quantizations.
  • Asymmetric MoE Allocation: Boundary layers (0 to 9 and 30 to 39) utilize Q3_K for representation stability, while deep core layers (10 to 29) apply IQ3_XXS calibrated against an authentic 820-chunk importance matrix (imatrix).
  • Speculative Head Preservation (blk.40): The Multi-Token Prediction (MTP) draft head is surgically isolated and quantized to Q4_K, ensuring operational speculative decoding without imatrix lookup issues.


Empirical Benchmarks & Fidelity Verification

Evaluated directly on the compiled GGUF binary using llama-perplexity over the standard WikiText-2 benchmark:

Metric Baseline (BF16 Unquantized) APEX-I-MiniPlus V2.1 (Current) Evaluation Setup
WikiText-2 Perplexity 7.3060 +/- 0.18633 7.7238 +/- 0.19671 Measured directly on GGUF binary (2048 ctx, 512 batch, 10 chunks)
Delta vs BF16 Base 0.0000 (Reference) +0.4178 (+5.71%) Near-zero degradation; sits at high-end Q4_K_M / Q5_K_M boundary
Model Footprint 62.40 GB (58.11 GiB) 13.76 GB (12.82 GiB) approx. 78.0% reduction in storage and memory bandwidth
Active Memory Overhead >65 GB VRAM <16 GB VRAM Runs completely offloaded on a single RTX 3090 / 4090 / 5080 GPU


Quality Spectrum: APEX-I-MiniPlus V2.1 vs. Standard Flat Quantizations

How the handcrafted APEX-I-MiniPlus V2.1 architecture compares against standard flat quantizations in llama.cpp on 29B Mixture-of-Experts architectures:

Quantization Format Bits Per Weight (BPW) Model Footprint (Disk / VRAM) Perplexity Delta (vs. BF16 Baseline) Token Fidelity & Syntactic Stability Tier
BF16 Base (Uncompressed) 16.00 bpw 62.40 GB 0.0000 (Reference) 100% full uncompressed reference fidelity.
Standard Q8_0 8.50 bpw approx. 33.2 GB approx. +0.0100 Virtually lossless; excessive memory overhead for consumer hardware.
Standard Q6_K 6.56 bpw approx. 26.0 GB approx. +0.0500 to +0.1000 Near-lossless BF16 fidelity; requires 32GB+ VRAM setups.
Standard Q5_K_M 5.50 bpw approx. 21.5 GB approx. +0.1500 to +0.2500 Commercial transparent threshold; exceeds standard single 16GB GPU limits.
Standard Q4_K_M 4.50 bpw approx. 18.0 GB approx. +0.3500 to +0.5500 Standard industry trade-off; requires partial offload on 16GB cards.
APEX-I-MiniPlus V2.1 (IsValorum) 3.42 bpw 13.76 GB (12.82 GiB) +0.4178 (PPL: 7.7238 +/- 0.19671) Delivers Q4_K_M / Q5_K_M class reasoning fidelity (+5.71% degradation) at a 78.0% weight-size reduction. Runs completely inside 16GB VRAM.
Standard Q3_K_M 3.44 bpw approx. 14.5 GB approx. +1.1000 to +1.8000 Noticeable syntax drop, bracket corruption, and tokenizer classification noise.
Standard IQ2_S / Generic Sub-3-Bit 2.50 bpw approx. 11.5 GB approx. +1.8000 to +4.0000+ Severe reasoning breakdown, high perplexity spikes in multi-step logic.

Model Files & Technical Specifications

File Name File Size Memory Footprint BPW Description
Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf 13.76 GB (12.82 GiB) 12.82 GiB 3.42 BPW Uncensored frontier MoE with MLA, Hyper-Connections, and MTP head
  • Base Model: huihui-ai/Huihui-Xing4.0-29B-A4B-abliterated
  • Parameters: 29.1B total (approx. 4.0B active per token)
  • Architecture: 40 layers, 64 routed MoE experts (4 active per token) + 1 shared expert + Multi-head Latent Attention (MLA)
  • Context Length: Native 131,072 tokens (128K)
  • Refusal Removal: Orthogonal activation steering via Heretic Directional Ablation


Surgical Tensor Quantization Map (Audited from GGUF)

The exact tensor breakdown below has been verified directly from the compiled binary weights:

Layer Group Sub-Component / Tensor Precision Engineering Rationale
Speculative Draft Head blk.40.* (All Tensors in MTP Block) Q4_K Surgical non-imatrix override bypassing imatrix lookup issues; enables speculative draft acceleration.
Global Output Head output.weight Q6_K Preserves near-FP16 token classification; eliminates syntax errors and bracket drops.
Global Embeddings token_embd.weight Q4_K High-fidelity vocabulary embedding representation.
All Normalizations output_norm, attn_*_norm, ffn_norm F32 100% uncompressed numerical stability across all layers.
Expert Routers blk.*.ffn_gate_inp, ffn_gate_inp_shexp F32 100% uncompressed routing fidelity across 64 experts; zero router drift.
Shared Foundation Experts blk.*.ffn_{gate,down,up}_shexp (All 40 Layers) Q5_K Foundation knowledge backbone active on 100% of tokens; protected in high-precision linear Q5_K.
Attention Output Heads blk.*.attn_output (All 40 Layers) Q6_K Armored attention output projection over deep context; prevents attention collapse.
Latent Attention Projections blk.*.attn_q_a, attn_q_b, attn_v_b, attn_kv_a_mqa Q4_K / Q5_K Preserves Multi-head Latent Attention (MLA) reconstruction geometry.
Boundary MoE Experts Layers 0 to 9 & 30 to 39 (ffn_*_exps) Q3_K Boundary layers preserved at Q3_K for input and output lexical stability.
Deep Core MoE Experts Layers 10 to 29 (ffn_*_exps) IQ3_XXS Calibrated with authentic 820-chunk importance matrix for maximum density in deep layers.


Hardware Throughput & Offload Benchmarks (RTX 30 / 40 / 50 & RAM Streaming)

Empirically verified in llama.cpp:

Hardware Target Offload Mode Generation Speed (Est.) Highlights
NVIDIA RTX 5080 / 5090 (Blackwell) Full GPU (-ngl 99) 120 to 160+ tok/s Full 128K context fits in VRAM with high throughput
NVIDIA RTX 4090 / 3090 (24GB VRAM) Full GPU (-ngl 99) 75 to 110+ tok/s Full offload with headroom for large context buffers
NVIDIA RTX 4080 / 4070 Ti Super (16GB VRAM) Full GPU (-ngl 99) 55 to 80+ tok/s 100% model fits in 16GB VRAM with native 32K context
System RAM (DDR4 / DDR5) Hybrid / CPU Offload 20 to 45 tok/s Zero AVX2 CPU stalls; efficient streaming from system RAM


The 16GB / 24GB Architecture: Full Context Runs In VRAM

Xing4.0-29B-A4B APEX-I-MiniPlus-V2.1 fits effortlessly within consumer VRAM limits:

Context Length Model Weights KV Cache (q8_0) Compute Buffers Total GPU VRAM (Est.) Feasibility
8,192 (8K) 12.82 GiB 0.35 GiB 0.80 GiB 13.97 GiB Full offload on 16GB GPUs
32,768 (32K) 12.82 GiB 0.75 GiB 1.10 GiB 14.67 GiB Full offload on 16GB GPUs
65,536 (64K) 12.82 GiB 1.40 GiB 1.45 GiB 15.67 GiB Tight fit on 16GB; effortless on 24GB
131,072 (128K) 12.82 GiB 2.70 GiB 1.90 GiB 17.42 GiB FULL 128K NATIVE IN 24GB VRAM!


Recommended Configuration & Setup

1. llama-cli (Console Generation)

llama-cli \
  -m Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf \
  -ngl 99 \
  -c 32768 \
  --temp 0.7 \
  --top-p 0.90 \
  -p "<|im_start|>system\nYou are an expert uncensored assistant.<|im_end|>\n<|im_start|>user\nExplain the architecture of multi-head latent attention.<|im_end|>\n<|im_start|>assistant\n"

2. llama-server (OpenAI-Compatible API)

llama-server \
  -m Xing4.0-29B-A4B.APEX-I-MiniPlus-V2.1-Abliterated.gguf \
  --port 8080 \
  --host 0.0.0.0 \
  -ngl 99 \
  -c 65536 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0


Recommended Generation Parameters

Hyperparameter Value Description
Temperature 0.70 Balanced creativity and analytical precision
Top-P 0.90 Standard nucleus threshold
Top-K 40 Retains vocabulary diversity while trimming outliers
Repeat Penalty 1.00 to 1.05 Keep near 1.0 for coding workflows


[!IMPORTANT]

CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

In programming code, brackets ({, }), assignment operators (=), and indentation whitespace repeat constantly across multi-line structures.

Common Issue: Many local frontends ship with repeat_penalty set to 1.1 or 1.15. Applying repeat penalties to code artificially penalizes necessary syntax tokens. When the logit of { drops, the model is forced to emit the next closest mathematical token, resulting in character swapping or dropped whitespace.

Recommendation: When executing code generation or debugging tasks, always set repeat_penalty to 1.00 or maximum 1.02.



Optional Support

If you find this surgical quantization valuable for your production workflows, local research, or agentic coding environments, consider starring the repository and following the IsValorum Hugging Face profile for continuous updates and frontier releases.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-01Polish official APEX-I-MiniPlus V2.1 release documentationf7829e115.8 KB
    Loading...
  2. 2026-10-01Update definitive APEX-I-MiniPlus V2.1 documentation and audited benchmarksc7180d615.4 KB
    Loading...
  3. 2026-10-01Upload README.md with huggingface_hubbe1f85f2.3 KB
    Loading...
  4. 2026-10-01Upload README.md with huggingface_hube01a81210.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.