license: mit
base_model: Blackfrost-AI/LING-3.0-FLASH-ABLITERATED
language:
- en
- zh
tags: - gguf
- moe
- abliterated
- bailingmoe3
LING-3.0-FLASH-ABLITERATED — GGUF (Q4_K_M)
GGUF conversion of Blackfrost-AI/LING-3.0-FLASH-ABLITERATED — the abliterated (uncensored) variant of LING 3.0 Flash, converted with stock llama.cpp.
| Property | Value |
|---|---|
| Architecture | bailingmoe3 (BailingMoeV3ForCausalLM) |
| Parameters | 124B total / ~5.1B active (MoE, 512 experts, 8 active) |
| Layers | 42 (layer group size 6) |
| Context | 262,144 (hardware-dependent) |
| License | MIT |
| Quantization | Q4_K_M, 4.83 BPW — no imatrix |
| File size | 77.0 GB (77,010,145,120 bytes) |
Usage
Requires llama.cpp built with bailingmoe3 support (commit 6d0549831 or newer — upstream since Aug 2026).
llama-server
llama-server -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf \
--host 0.0.0.0 --port 8080 \
-ngl 99 # offload all layers to GPU(s)
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 512
}'
llama-cli
llama-cli -m LING-3.0-FLASH-ABLITERATED-Q4_K_M.gguf -ngl 99 \
-p "Your prompt here" -n 512
Also works in LM Studio, Ollama (after ollama create), and any llama.cpp-compatible client.
Notes
- Reasoning model: emits a hidden chain-of-thought before the final answer (surfaced as
reasoning_contentin the OpenAI-compatible API). Leave enoughmax_tokensheadroom for thinking + answer. - No imatrix: plain Q4_K_M, not an i-quant. Quality is near-lossless relative to the f16 source (quantized via Q8_0 intermediate).
- Fallback tensors: 8 of 938 tensors (
blk.*.attn_k_b.weight,ncols=128not divisible by 256) fell back toq5_0due to the Q4_K_M block-size constraint — negligible impact. - MTP/NextN layer tensors are present in the GGUF; llama.cpp currently ignores them (harmless warning at load).
Verification
sha256: f47f38cfdac87837220fa34a3ba026b83498d9aa19996b18ba7f312b11be9fa6- Coherence-tested with llama.cpp
6d0549831(fact/QA, math word problem, code generation, creative writing).
Original model
- Repo: Blackfrost-AI/LING-3.0-FLASH-ABLITERATED
- Base: InclusionAI LING-3.0-Flash (open weights, MIT)