← back to catalog · registered 2026-08-22 13:56

SC117/Agents-A1-Uncensored-MTP-APEX-GGUF

SC117 GGUF MoE multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FAgents-A1-Uncensored-MTP-APEX-GGUF"
Response includes
  • classification m-uncensored
  • files 9
  • hub_downloads_all_time 40,153
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
40K
2K last 30d - cooling
Likes
16
Model age
2mo ago
created 2026-07-18
Downloads over time
Now42K→from0↑0%
015.4K30.8K46.2K0 on Jul 1542K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 54 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
BF16
Tags
transformers gguf qwen3_5_moe qwen3_5 reasoning agentic uncensored mtp apex quantization multimodal text-generation

Related

Total size
142 GB
Files
9
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-10-03 10:25

Files by quantization

BF16 2 files 67.0 GB
Agents-A1-Uncensored-MTP-BF16.gguf 66.2 GB 375e3036 download
mmproj-Agents-A1-Uncensored-MTP-BF16.gguf 861 MB 46979c84 download
Auxiliary files 7 files 75.6 GB
Agents-A1-Uncensored-MTP-APEX-I-Balanced.gguf 24.3 GB d1b132ba download
Agents-A1-Uncensored-MTP-APEX-I-Quality.gguf 21.9 GB 397894c0 download
Agents-A1-Uncensored-MTP-APEX-I-Compact.gguf 16.1 GB d3815ae9 download
Agents-A1-Uncensored-MTP-APEX-I-Mini.gguf 13.3 GB e0c002d9 download
README.md 24.8 KB 22b8d227 download
README_zh.md 9.34 KB 4f6ed964 download
.gitattributes 1.94 KB f5a9d9eb download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/InternScience/Agents-A1/blob/main/LICENSE
pipeline_tag: text-generation
tags:

  • qwen3_5_moe
  • qwen3_5
  • reasoning
  • agentic
  • uncensored
  • mtp
  • apex
  • quantization
  • gguf
  • multimodal
    base_model:
  • InternScience/Agents-A1

APEX MTP Vision Apache-2.0

Agents-A1-Uncensored-MTP-APEX

English | 📖 中文文档

Uncensored 35B agentic MoE · APEX-quantized GGUFs + MTP + BF16 mmproj

🤖 About Agents-A1

Agents-A1 is a 35B-parameter Mixture-of-Experts agentic model from InternScience, post-trained on top of Qwen3.5-35B-A3B via a three-stage paradigm: full-domain SFT → domain-level teacher training → multi-teacher multi-domain on-policy distillation.

Despite operating in the ~35B model class, Agents-A1 delivers highly competitive performance against frontier-scale systems such as GPT-5.5, DeepSeek-V4-pro, and Kimi-K2.6 — achieving SOTA on Seal-0 (56.4), HiPhO (46.4), FrontierScience-Olympiad (79.0), IFBench (80.6), IFEval (94.8), and best-among-comparable on BrowseComp (75.5), XBench-DS-2510 (86.0), GAIA (96.0), SciCode (44.3), HLE (47.6), and MolBench-bind (56.8).

This GGUF package includes the mmproj-Agents-A1-Uncensored-MTP-BF16.gguf vision projector for multimodal (image + text) capabilities with llama.cpp. MTP layers are extracted from Qwen3.5-35B-A3B and injected into Agents-A1's safetensors (see MTP Extraction & Injection section). License: Apache-2.0.

⚠️ Uncensored Adaptation

This release merges the selected behavior-modification LoRA into the BF16 base before MTP injection and APEX quantization. The uncensored adaptation was produced with abliterix. It is intended to reduce refusal behavior and can produce responses that differ materially from the original model.

Evaluation result2 refusals / 100 prompts
KL divergence0.016058

Use responsibly: evaluate the model for your deployment context and apply appropriate safeguards.

🧠 Model Details
ArchitectureQwen3.5 MoE (Mixture of Experts)
Parameters35B total, 3B active per token
Experts256 routed experts, 8 active per token
Layers40 transformer layers + 1 MTP layer
Context262,144 tokens
MTP SourceQwen3.5-35B-A3B (1 layer, 785 tensors, injected)
Block Count41 (blk.0–39 + blk.40 MTP)
LicenseApache-2.0
🔧 MTP Extraction & Injection

The released InternScience/Agents-A1 checkpoint is a 40-layer Qwen3.5-35B-A3B MoE without MTP (Multi-Token Prediction) layers. To enable MTP acceleration in llama.cpp (which speeds up long-context generation by 10–30%), we extract the 1 MTP layer from Qwen3.5-35B-A3B and inject it into Agents-A1's safetensors before GGUF conversion.

The full pipeline has 4 steps: (1) Extract 785 MTP tensors from Qwen3.5-35B-A3B (filter keys containing "mtp"); (2) append as a new shard model-15-of-15.safetensors and update model.safetensors.index.json's metadata.total_size and weight_map (do not modify the original 14 shards); (3) convert to BF16 GGUF via the master-branch llama.cpp convert_hf_to_gguf.py — the master build auto-detects regular layers blk.0–39 + MTP layer blk.40.nextn.* (785 tensors); (4) quantize with APEX llama-quantize using qwen36_35b_mtp_*.txt configs which already include the blk.40 override (MTP kept at Q8_0 across all tiers — no manual patching needed). The imatrix Qwen3.5-35B-A3B.imatrix.gguf is reused directly (same architecture, compatible weights).

Reproduction command (example: I-Compact tier)

F:\llama.cpp\...\llama-quantize.exe ^
  --imatrix J:\Models\Qwen3.5-35B-A3B.imatrix.gguf ^
  --tensor-type-file E:\apex-quant\configs\qwen36_35b_mtp_compact.txt ^
  J:\Models\Agents-A1-Uncensored-MTP-GGUF\Agents-A1-Uncensored-MTP-BF16.gguf ^
  J:\Models\Agents-A1-Uncensored-MTP-APEX-GGUF\Agents-A1-Uncensored-MTP-APEX-I-Compact.gguf ^
  Q4_K_M
📊 BenchLocal Results (APEX-I-Compact, 16.14 GB)
ModeToolCall-15BugFind-15HermesAgent-20MaxEff.
Thinking100888791.271.2
No Thinking971008593.157.1

RTX 5070 Ti 16GB + 128GB RAM · No-thinking mode achieves higher ceiling (BugFind +12) but suffers more retries on complex agent scenarios.

🚀 Usage

llama.cpp (text only)

hf download SC117/Agents-A1-Uncensored-MTP-APEX-GGUF --include "*.gguf" --local-dir ./models
./llama-server -m ./models/Agents-A1-Uncensored-MTP-APEX-I-Compact.gguf -ngl 99 -c 131072

llama.cpp (vision + text)

./llama-server -m ./models/Agents-A1-Uncensored-MTP-APEX-I-Compact.gguf --mmproj ./models/mmproj-Agents-A1-Uncensored-MTP-BF16.gguf -ngl 99 -c 131072

vLLM

vllm serve SC117/Agents-A1-Uncensored-MTP-APEX-GGUF --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3
·
Tool-call variant
vllm serve SC117/Agents-A1-Uncensored-MTP-APEX-GGUF --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder

SGLang

python3 -m sglang.launch_server --model-path "SC117/Agents-A1-Uncensored-MTP-APEX-GGUF" --host 0.0.0.0 --port 30000
🎛️ Recommended Sampling Parameters

From the official Agents-A1 model card:

temperature0.85
top_p0.95
top_k20
min_p0.0
presence_penalty1.1
repetition_penalty1.0
💡 What is APEX?

These GGUF files are quantized using APEX, an MoE-aware mixed-precision quantization technique. APEX classifies every tensor by its role — routed expert, shared expert, SSM, or attention — and applies a layer-wise precision gradient, giving sensitive edge layers (including the MTP layer) higher precision and compressing redundant middle layers more aggressively.

APEX beats Q8_0 perplexity at half the size — and even beats F16 in some cases.

The qwen36_35b_mtp_*.txt configs include overrides for blk.40 (the MTP layer), preserving it at Q8_0 across all four I- tiers. The same Qwen3.5-35B-A3B.imatrix.gguf is reused (same architecture, compatible MoE expert layout).

📦 APEX Quantization Tiers
FileSizeProfileBest For
*-APEX-I-Quality.gguf21.75 GBI-QualityHigh quality (Q6_K + iq4_xs attention)
*-APEX-I-Balanced.gguf24.21 GBI-BalancedBest all-rounder (Q6_K + Q5_K experts)
*-APEX-I-Compact.gguf16.14 GBI-CompactBest quality/size ratio (Q4_K default)
*-APEX-I-Mini.gguf13.36 GBI-MiniMost compact, fits in 16GB VRAM (Q3_K + iq2_s)

BF16 source: Agents-A1-Uncensored-MTP-BF16.gguf (66.19 GB). imatrix: Qwen3.5-35B-A3B.imatrix.gguf (reused from base model).

Links

Citation

@misc{bai2026scalinghorizonparametersreaching,
      title={Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent},
      author={Lei Bai and Zongsheng Cao and Yang Chen and Zhiyao Cui and Shangheng Du and Yue Fan and Shiyang Feng and Zijie Guo and Haonan He and Liang He and Xiaohan He and Shuyue Hu and Yusong Hu and Songtao Huang and Yichen Jiang and Hao Li and Xin Li and Dahua Lin and Weihao Lin and Fenghua Ling and Dongrui Liu and Zhuo Liu and Runmin Ma and Chunjiang Mu and others},
      year={2026},
      eprint={2606.30616},
      archivePrefix={arXiv}
}

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-03Add optional Ko-fi support banner42f245e25 KB
    Loading...
  2. 2026-07-18Fix bilingual README links and add abliterix linkd36059424.6 KB
    Loading...
  3. 2026-07-18Add bilingual Uncensored model cardd6f795e24.5 KB
    Loading...
  4. 2026-07-18Add files using upload-large-folder tool40c824f24.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration