← back to catalog · registered 2026-08-22 13:56

h34v7/LING-3.0-FLASH-ABLITERATED-GGUF

h34v7 GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/h34v7%2FLING-3.0-FLASH-ABLITERATED-GGUF"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 160
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
160
Likes
1
Model age
7w ago
created 2026-08-19
Downloads over time
Now371→from105↑253%
92194296398105 on Aug 19371 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Quantizations
Q4_K
Tags
gguf moe abliterated bailingmoe3 en zh base_model:Blackfrost-AI/LING-3.0-FLASH-ABLITERATED base_model:quantized:Blackfrost-AI/LING-3.0-FLASH-ABLITERATED license:mit endpoints_compatible region:us conversational

Related

Total size
71.7 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 09:06

Files by quantization

Q4_K 1 file 71.7 GB
LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf 71.7 GB f47f38cf download
Auxiliary files 2 files 4.03 KB
README.md 2.47 KB a542b200 download
.gitattributes 1.56 KB dc10db22 download

README current version from Hugging Face


license: mit
base_model: Blackfrost-AI/LING-3.0-FLASH-ABLITERATED
language:

  • en
  • zh
    tags:
  • gguf
  • moe
  • abliterated
  • bailingmoe3

LING-3.0-FLASH-ABLITERATED — GGUF (Q4_K_M)

GGUF conversion of Blackfrost-AI/LING-3.0-FLASH-ABLITERATED — the abliterated (uncensored) variant of LING 3.0 Flash, converted with stock llama.cpp.

Property Value
Architecture bailingmoe3 (BailingMoeV3ForCausalLM)
Parameters 124B total / ~5.1B active (MoE, 512 experts, 8 active)
Layers 42 (layer group size 6)
Context 262,144 (hardware-dependent)
License MIT
Quantization Q4_K_M, 4.83 BPW — no imatrix
File size 77.0 GB (77,010,145,120 bytes)

Usage

Requires llama.cpp built with bailingmoe3 support (commit 6d0549831 or newer — upstream since Aug 2026).

llama-server

llama-server -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080 \
  -ngl 99          # offload all layers to GPU(s)
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 512
  }'

llama-cli

llama-cli -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf -ngl 99 \
  -p "Your prompt here" -n 512

Also works in LM Studio, Ollama (after ollama create), and any llama.cpp-compatible client.

Notes

  • Reasoning model: emits a hidden chain-of-thought before the final answer (surfaced as reasoning_content in the OpenAI-compatible API). Leave enough max_tokens headroom for thinking + answer.
  • No imatrix: plain Q4_K_M, not an i-quant. Quality is near-lossless relative to the f16 source (quantized via Q8_0 intermediate).
  • Fallback tensors: 8 of 938 tensors (blk.*.attn_k_b.weight, ncols=128 not divisible by 256) fell back to q5_0 due to the Q4_K_M block-size constraint — negligible impact.
  • MTP/NextN layer tensors are present in the GGUF; llama.cpp currently ignores them (harmless warning at load).

Verification

  • sha256: f47f38cfdac87837220fa34a3ba026b83498d9aa19996b18ba7f312b11be9fa6
  • Coherence-tested with llama.cpp 6d0549831 (fact/QA, math word problem, code generation, creative writing).

Original model

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Upload README.md with huggingface_hub9c974582.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration