← back to catalog · registered 2026-08-22 13:56

Vegss/Step-3.7-Flash-AEON-Ultimate-Abliterated-GGUF

Vegss GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Vegss%2FStep-3.7-Flash-AEON-Ultimate-Abliterated-GGUF"
Response includes
  • classification m8
  • files 13
  • hub_downloads_all_time 669
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
669
66 last 30d - cooling
Likes
1
Model age
4mo ago
created 2026-06-08
Downloads over time
Now684→from233↑194%
210383556729233 on Jun 10684 on Oct 11684 on Oct 10JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q3_K
Tags
gguf abliterated uncensored moe vision-language thinking step3p7 llama.cpp imatrix experimental expert-granular-abliteration image-text-to-text

Related

Total size
335 GB
Files
13
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-06-08 13:49

Files by quantization

Q3_K 3 files 94.4 GB
Step-3.7-Flash-AEON-Ultimate-Abliterated-Q3_K_M-00001-of-00003.gguf 41.8 GB b2c275df download
Step-3.7-Flash-AEON-Ultimate-Abliterated-Q3_K_M-00002-of-00003.gguf 41.7 GB e74bdaeb download
Step-3.7-Flash-AEON-Ultimate-Abliterated-Q3_K_M-00003-of-00003.gguf 11.0 GB 14bd3899 download
F16 1 file 4.10 GB
mmproj-step37-flash-f16.gguf 4.10 GB efe6c322 download
Auxiliary files 9 files 240 GB
Step-3.7-Flash-AEON-Ultimate-Abliterated-IQ1_M-00001-of-00002.gguf 41.6 GB 6739fbc5 download
Step-3.7-Flash-AEON-Ultimate-Abliterated-q8_0-00002-of-00005.gguf 41.4 GB b35cc086 download
Step-3.7-Flash-AEON-Ultimate-Abliterated-q8_0-00004-of-00005.gguf 41.4 GB 424472ae download
Step-3.7-Flash-AEON-Ultimate-Abliterated-q8_0-00003-of-00005.gguf 41.4 GB 3b59284b download
Step-3.7-Flash-AEON-Ultimate-Abliterated-q8_0-00001-of-00005.gguf 40.5 GB 7a452d92 download
Step-3.7-Flash-AEON-Ultimate-Abliterated-q8_0-00005-of-00005.gguf 30.3 GB 82f6e183 download
Step-3.7-Flash-AEON-Ultimate-Abliterated-IQ1_M-00002-of-00002.gguf 3.45 GB 7c8b56b5 download
README.md 7.86 KB 52274692 download
.gitattributes 2.55 KB fe7d6b50 download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Step-3.7-Flash-AEON-Ultimate-Abliterated-BF16
base_model_relation: quantized
library_name: gguf
pipeline_tag: image-text-to-text
tags:

  • abliterated
  • uncensored
  • moe
  • vision-language
  • thinking
  • step3p7
  • gguf
  • llama.cpp
  • imatrix
  • experimental
  • expert-granular-abliteration
    language:
  • en

⚠️ KNOWN BROKEN — do not use for inference yet (fix in progress)

These GGUFs currently produce garbled output. The cause is NOT an upstream llama.cpp engine bug — any earlier note on this card claiming an "engine-blocked" / PR-#23845 dependency is outdated; please disregard it. The official Step-3.7 GGUF runs fine on a correctly-built llama.cpp.

Root cause: our Expert-Granular Abliteration interacts badly with low-bit quantization (same as our NVFP4) — the ablation zeroes a residual-stream subspace that is exact at BF16 but re-corrupted by quant noise at 3–4 bit → garbage. Coherent output requires BF16.

✅ Use instead: the BF16 release. A milder-ablation re-quant that survives low-bit is being validated; these files will be replaced or withdrawn once fixed.


Step-3.7-Flash-AEON-Ultimate-Abliterated-GGUF ⚠️ EXPERIMENTAL

GGUF quants of AEON-7/Step-3.7-Flash-AEON-Ultimate-Abliterated-BF16 (198B / ~11B-active sparse-MoE vision-language thinking model), built for single-DGX-Spark deployment.

⚠️ EXPERIMENTAL — NOT YET FUNCTIONAL (engine-blocked)

These GGUFs currently produce garbage output on every available GGUF runtime (llama.cpp, Ollama, LM Studio, KoboldCpp — all share the same engine). The cause is not these files — the quantized weights, abliteration, and tokenizer are all verified correct. The blocker is an open upstream bug in llama.cpp's Step-3.7 inference graph: Step-3.7 is routed through the step35 compute graph, which mis-runs its forward pass (garbage from the first token, independent of bit-width — even the near-lossless q8_0 is affected).

Dependency to use these properly: a corrected llama.cpp Step-3.7 inference implementation (tracking ggml-org/llama.cpp#23845 / a StepFun-fork fix). They are expected to work as-is once that lands — no re-quantization needed.

(Tokenizer note: it is correct. The right pre-tokenizer is deepseek-v3 — if a build defaults otherwise, pass --override-kv tokenizer.ggml.pre=str:deepseek-v3. This is a minor correctness item, not the blocker.)

Status: experimental until functionality is confirmed on a fixed engine. For working deployment today, use the BF16 or NVFP4 releases (table below).


Model family — formats, quality, validation

Release Format Size Target hardware Quality Refusals removed Validation state
…-BF16 BF16 safetensors 376 GB multi-GPU (≥2× Spark / Blackwell) reference (full) ✅ d≈10→0.35 ✅ working; weight-verified, prefill refusal-collapse confirmed
…-NVFP4 NVFP4 W4A4 (modelopt) 124 GB 2× DGX Spark (TP=2) near-full (RT err 0.095) ✅ ✅ working path; weight-verified (down 0.095; o_proj/up bit-exact)
…-GGUF / q8_0 GGUF (exp) 209 GB (near-lossless base) near-lossless ✅ (weights) ⚠️ experimental — engine-blocked
…-GGUF / Q3_K_M GGUF dynamic (exp) ~101 GB 1× DGX Spark high (3-bit dyn.) ✅ (weights) ⚠️ experimental — engine-blocked
…-GGUF / IQ1_M GGUF dynamic (exp) 48 GB (~1.95 bpw) 1× DGX Spark (max KV headroom) low (1.5-bit; below IQ2 cliff) ✅ (weights) ⚠️ experimental — engine-blocked

Legend: ✅ working today · ⚠️ experimental, awaiting the upstream engine fix.


Two independent things this build is

  1. Abliterated (behavior) — refusals removed via Expert-Granular Abliteration across all 288 experts (refusal subspace collapsed from Cohen's d≈10 → 0.35). Uncensored.
  2. Precisely quantized (fidelity) — a data-driven, per-component mixed-precision scheme + our own imatrix, not a uniform low-bit dump. Capable + still-uncensored after quantization (when the engine runs it).

Quantization methodology (data-driven selective allocation)

Per-component bits from our outlier study + refusal-subspace map (not stock Q3_K_M):

Component Q3_K_M tier IQ1_M tier Rationale (measured)
Expert gate/up_proj Q3_K IQ1_M cleanest family (FP4-g16 err 0.094) → bulk savings
Expert down_proj Q4_K IQ2_XXS most quant-sensitive expert block
self_attn.o_proj Q6_K Q5_K 13.1× outlier
q/k/v, attn-gate Q5_K Q4_K —
share_expert.* Q5/Q6_K Q4/Q5_K shared.down 18.7× outlier
dense MLP (L0–2) Q5_K Q4_K dense.down 24× outlier
router (ffn_gate_inp) FP32 FP32 routing fully preserved
embed / output Q6_K Q4/Q5_K —
vision (mmproj) F16 F16 kept

Plus a custom imatrix (diverse general/reasoning/code calibration).

Inference (once a fixed Step-3.7 engine is available)

# Build the StepFun step3.7 llama.cpp fork (or a future fixed mainline)
git clone -b step3.7 https://github.com/stepfun-ai/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=121 && cmake --build build -j --config Release

# Serve (shards auto-load from the first piece; mmproj for vision)
./build/bin/llama-server \
  -m  Step-3.7-Flash-…-Q3_K_M-00001-of-0000N.gguf \
  --mmproj mmproj-step3.7-flash-f16.gguf \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  -c 131072 --parallel 4 -ngl 999 --flash-attn \
  --override-kv tokenizer.ggml.pre=str:deepseek-v3 \
  --host 0.0.0.0 --port 8080

This will emit garbage until the upstream Step-3.7 graph bug (#23845) is fixed. Q3_K_M targets one Spark with moderate KV headroom; IQ1_M maximizes headroom (quality-tolerant, below the IQ2 cliff); q8_0 is the near-lossless base.


Quantized on NVIDIA B300 via the StepFun step3.7 llama.cpp fork + custom imatrix, from the AEON-Ultimate abliterated BF16. Base model © StepFun AI, Apache-2.0.


☕ Support the work

If this release has been useful, tips are deeply appreciated — they go directly toward more compute, more models, and more open releases.

₿ Bitcoin (BTC)
QR
bc1q09xmzn00q4z3c5raene0f3pzn9d9pvawfm0py4
Ξ Ethereum (ETH)
QR
0x1512667F6D61454ad531d2E45C0a5d1fd82D0500
◎ Solana (SOL)
QR
DgQsjHdAnT5PNLQTNpJdpLS3tYGpVcsHQCkpoiAKsw8t
ⓜ Monero (XMR)
QR
836XrSKw4R76vNi3QPJ5Fa9ugcyvE2cWmKSPv3AhpTNNKvqP8v5ba9JRL4Vh7UnFNjDz3E2GXZDVVenu3rkZaNdUFhjAvgd

Ethereum L2s (Base, Arbitrum, Optimism, Polygon, etc.) and EVM-compatible tokens can be sent to the same Ethereum address.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-08Duplicate from AEON-7/Step-3.7-Flash-AEON-Ultimate-Abliterated-GGUF559c8d47.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration