license: apache-2.0
task_categories:
- text-generation
language: - en
size_categories: - n<1K
tags: - atomic-chat
- gguf
- llama.cpp
- abliteration
- refusal-direction
- ternary
- bonsai
- metrics
Ternary-Bonsai-2-27B-Abliterate-LoRA-GGUF-metrics
Measurements behindAtomicChat/Ternary-Bonsai-2-27B-Abliterate-LoRA-GGUF,
the rank-1 refusal-ablation adapter for PrismML's 1.75 bit/weight ternary pack.
Both packs are covered: every row carries a pack column, PTQ1_0 or PQ2_0. The same
adapter file was run on both, and on the refusal evaluation all 416 greedy replies came out
byte-identical across packs.
Aggregates only. Prompt text is not redistributed (the sources are named in the model
card) and model outputs on the harmful split are not published at all.
| file | rows | what |
|---|---|---|
runtime_leak.csv |
44 | how much signal is left along the refusal direction, measured inside the running model |
refusal_sweep.csv |
38 | refusals, empty and degenerate replies per adapter and strength |
mmlu.csv |
6 | MMLU accuracy per configuration |
mmlu_by_subject.csv |
57 | the same, per subject |
direction_rows.csv |
65 | per-layer statistics of the estimated direction |
runtime_leak.csv
The headline measurement. A probe built against
PrismML's llama.cpp fork (tagprism-b10709-9a9394a) taps every residual write during a real forward pass and reports|r.y| / |y| - the fraction of each write that lies along the refusal direction - plus the
same figure for the residual stream itself across all 64 blocks.
Base model sits around 1e-2. A correct adapter at scale 1 drives every writer to single
digit 1e-6. The published OrcaRouter adapter reaches that on ffn_down and attn_output
but leaves linear_attn_out (ssm_out, 48 of the 129 sites) at 1.6e-2, because its
factors for those sites are in the checkpoint's V-head order rather than llama.cpp's.
samples is tokens x layers behind each mean.
refusal_sweep.csv
104 harmful + 104 harmless held-out prompts, greedy, 64-token budget, thinking off,
identical seed and system prompt across configurations. refusal_rate_of_valid counts
refusals among replies that are neither empty nor degenerate, because over-projection at
scale 2 produces empty replies that a naive counter reads as compliance.
Refusal detection is a rule-based opening-phrase match: indicative, not a judge. A reply
that answers and then adds a disclaimer counts as compliance.
mmlu.csv, mmlu_by_subject.csv
500 questions stratified over 57 subjects, single letter forced by a root ::= [A-D]
grammar, thinking off. Answer-only, so the absolute numbers sit below PrismML's published
thinking-mode result; the comparison between configurations is the point. At n=500 the
standard error is about 2 points, and per subject it is far larger - read the by-subject
file as texture, not as 57 separate results.
direction_rows.csv
The direction file is [65, 5120]: row 0 is the embedding output, row L the residual
stream entering block L. Per row: cosine with OrcaRouter's published direction (estimated
independently, on the bf16 model), AUROC and Cohen's d separating harmful from harmless on
the held-out split, and zero_at - where ablation puts a prompt on the axis from the
harmless mean (0) to the harmful mean (1). zero_at near 0 is what keeps ablation from
inducing refusals on ordinary questions.
Separation does not predict behaviour: row 38 leads on Cohen's d, row 42 works better in
the sweep.
Reproduction
Tools and the full pipeline are described in the model card.