license: mit
base_model:
- inclusionAI/Ling-3.0-tiny
base_model_relation: finetune
pipeline_tag: text-generation
library_name: gguf
tags: - gguf
- llama.cpp
- abliterated
- heretic
- bailingmoe3
- moe
Ling-3.0-tiny-Abliterated — GGUF
An abliterated build of inclusionAI/Ling-3.0-tiny
(8B total / ~1B active MoE, hybrid KDA + MLA attention, a thinking model), produced with
Heretic and quantized with llama.cpp. Other abliterated Ling-3.0
builds exist; this one is my own run, and every claim below is measured against a stock baseline built
through the same pipeline.
Modification notice. These weights are a modified derivative. Directional ablation altered
1,692 tensors to reduce refusals (see What changed); nothing else was retrained.
inclusionAI did not produce, endorse, or review this build. Its behavior is not theirs.
Read this first
This is an escape hatch for over-refusal, not an uncensored model. Abliteration removes a refusal
direction from the residual stream — it adds no knowledge. This build knows exactly what the stock
8B knows and nothing more; what changed is its willingness to say it. If stock answers your prompt,
use stock.
It is effective, and that cuts both ways. Unlike some abliterations that only dent the
non-thinking path, this one drops refusals in both regimes. That makes the hazard below real,
not theoretical.
The hazard worth naming: refusals double as hedging. Remove them and you get a small model
answering fluently in domains where it is unreliable, with none of the cues that used to mark the
boundary. Factual reliability is unchanged; only the signalling is gone. Don't trust confident output
here more than you'd trust confident output from stock — especially on anything safety-critical. You
are accountable for what you generate and for applicable law.
Measurements
All on Q4_K_M, greedy-ish sampling (temp 1.0 / top-p 0.95 / top-k 20, the model's own defaults),
against a stock Q4_K_M built through the same pipeline. Refusals usemlabonne/harmful_behaviors test[60:]
— the 44 prompts Heretic's optimiser never saw (it scored test[:60]) — classified by Heretic's
marker list on the answer text.
Refusals (lower = complies more):
| regime | stock | abliterated |
|---|---|---|
| thinking off | 43/44 | 11/44 |
| thinking on | 44/44 | 12/44 |
The compliances are substantive: of the 33/32 non-refusals, 0 were degenerate (no empty or
looping output; mean ~600–2000 tokens of on-topic text). This is the opposite of what a marker metric
can hide — the drop is real compliance, not scoring noise.
GSM8K, paired, n=100, thinking off, identical sampling:
| correct | truncated | |
|---|---|---|
| stock | 93/100 | 1 |
| abliterated | 93/100 | 1 |
Net 0. Discordant pairs split 3/3 (McNemar exact p = 1.000). Abliteration damage tends to
land in math reasoning (a documented tendency for Heretic across models); at this KL,
none is detectable here. n=100 excludes a large regression, not a 2–3 pp one. MMLU, long context,
and non-English remain unmeasured.
Files
| file | size |
|---|---|
Ling-3.0-tiny-Abliterated-Q4_K_M.gguf |
4.49 GiB |
Ling-3.0-tiny-Abliterated-Q5_K_M.gguf |
5.25 GiB |
Ling-3.0-tiny-Abliterated-Q6_K.gguf |
6.05 GiB |
Ling-3.0-tiny-Abliterated-Q8_0.gguf |
7.83 GiB |
All quantized from one bf16 GGUF, imatrix-free (Q4_K_M and up don't require one).
Usage
llama-server -m Ling-3.0-tiny-Abliterated-Q4_K_M.gguf \
-ngl 99 -c 32768 -fa on --jinja \
--temp 1.0 --top-p 0.95 --top-k 20
Thinking is on by default; pass --chat-template-kwargs '{"enable_thinking":false}' to disable it.
Send a large max_tokens or none — with thinking on, a small budget is consumed inside<think> and returns empty content with finish_reason: "length".
What changed
Per the merge-time delta list (computed before quantization, so no quantizer in the loop):
| tensor type | changed |
|---|---|
routed-expert down_proj |
1,664 |
attention output (o_proj/dense) |
15 |
shared-expert down_proj |
13 |
| everything else (norms, embeddings, q/k/v, gates, experts' gate/up) | 0 |
1,692 modules total, all in layers 8–23 (the back half); layers 0–7 untouched. This is the
expected shape for refusal-direction ablation.
Reproducing
pip install -U heretic-llm
heretic inclusionAI/Ling-3.0-tiny \
--response-prefix '</think>' \
--system-prompt $'You are a helpful assistant.\ndetailed thinking off' \
--max-batch-size 128 --max-response-length 64 --n-trials 100
--response-prefix '</think>' and the detailed thinking off system prompt put scoring in Ling's
own non-thinking regime (byte-identical to enable_thinking=false); without them Heretic's refusal
markers fire on reasoning text. Abliteration is not bit-reproducible — the TPE search lands on a
different trial each run — so the selected trial 42 (6/60 refusals @ KL 0.0288, baseline 59/60,
per-layer directions) is published for --reproduce:
attn.o_proj : max_weight 1.4227 max_weight_position 18.860 min_weight 1.3238 min_weight_distance 10.697
mlp.down_proj : max_weight 1.3516 max_weight_position 13.909 min_weight 0.6552 min_weight_distance 6.881
(attn.o_proj covers the KDA layers' o_proj and the MLA layers' dense; mlp.down_proj covers
routed and shared experts.)
Note for anyone running Ling-3.0 in Transformers (not GGUF): under the default
attn_implementation="sdpa", Ling's MLA layers attend non-causally on any unpadded prefill,
because its code builds the 4-D causal mask only when padding is present and relies on SDPA'sis_causalotherwise — which its eager path never sets. Load withattn_implementation="eager".
llama.cpp (and therefore these GGUFs) is unaffected.
Verifying
941ef0a014d2c28b684a39ed77e87c4c3d4a0c107aae61949ec291402c206bd5 Ling-3.0-tiny-Abliterated-Q4_K_M.gguf
77533ffe319d50eeeec613536d604e0e1de08e85eda3942054e2f1b8378c8f51 Ling-3.0-tiny-Abliterated-Q5_K_M.gguf
cf61b3c2f1bb8fdda2e95214f2afdf4083cb3c4708a1a859c042efc5971fb2ef Ling-3.0-tiny-Abliterated-Q6_K.gguf
091fdc479ebeabe3195508613f08bc6a200a65d0e7cc275bd9f3141cf514dd16 Ling-3.0-tiny-Abliterated-Q8_0.gguf
License
MIT, inherited from inclusionAI/Ling-3.0-tiny. The base repo declares MIT in its metadata but
ships no LICENSE file; the LICENSE here is the standard MIT text with inclusionAI's own copyright
line, taken verbatim from their sibling repos (Ling-V2 / Ling-V2.5) which do carry it. Abliteration
and quantization change nothing about the license.
Credits
- inclusionAI — the Ling-3.0-tiny base model
- p-e-w/heretic — the abliteration method and tooling
- ggml-org/llama.cpp — conversion and quantization
- mlabonne — the prompt sets used for directions and scoring