Reading model benchmarks
No single number describes an abliterated model. UGI, LM Arena, Open LLM Leaderboard v2 measure three different things - and none of them measures the thing abliteration actually did. How to read the tables without being fooled by them.
Three benchmark sources dominate. UGI Leaderboard (uncensored-specific, run by DontPlanToEnd) measures willingness and knowledge on sensitive topics via a confidential question set. LM Arena aggregates human blind pairwise preferences via Elo/Bradley-Terry with over 7M votes by 2026, but has no separate track for uncensored models. Open LLM Leaderboard v2 was archived by Hugging Face in March 2025 - a static historical snapshot with MMLU-Pro, IFEval, BBH, MATH, GPQA, MuSR, but no ongoing evaluation. Read scores against the parent model, not across benchmarks. Abliterated models' scores drop in patterned ways: MMLU barely moves, TruthfulQA drops 5-11 points, GSM8K is where breakage shows first.
- The three dominant leaderboards and what each actually measures
- How to interpret UGI's UGI/W10/NatInt/Writing/political-lean columns
- Why LM Arena has no uncensored track and what that hides
- What 'archived in March 2025' means for Open LLM Leaderboard v2
- The patterned way abliteration shows up in benchmark tables
- Why KL divergence is the honest adjacent metric
- How to run your own refusal + capability checks with AdvBench and lm-eval
- The judge-panel approach: HarmBench, StrongREJECT, JailbreakBench, XSTest
The core principle
No single number describes an abliterated model, and the numbers that exist were built for different purposes. Three sources dominate the field, and they measure three different things. Reading a benchmark table for an abliterated model well means knowing which board you are looking at, what it is asking, and what it is not asking.
The UGI Leaderboard
The Uncensored General Intelligence Leaderboard, run by the pseudonymous maintainer DontPlanToEnd as a Hugging Face Space, is the only major board built specifically for this class of model. It uses a confidential set of questions - kept secret precisely so that models cannot be trained to game them - and scores several axes.
The columns, decoded:
- UGI score (0-100) - the headline. Measures breadth of knowledge on sensitive or controversial topics: what the model actually knows once willing to say it. A high UGI on a model that also has a low W/10 means the model knows things but usually declines to share them.
- W/10 - willingness (0-10). Measures how readily the model engages rather than deflecting. A 10 means it almost always answers. This is the column that most directly reflects successful abliteration.
- NatInt - natural intelligence. A general-reasoning score that lets a reader separate a model that is uncensored from one that is merely uninhibited and dim. Two models with W/10 = 10 can have wildly different NatInt.
- Writing - rates prose quality on open-ended generation.
- Political lean - places the model on several ideological axes.
The board spans models from about 1B to 155B parameters. It is the right instrument for the specific question "will this model answer, and does it know anything," and the wrong one for "is this the most capable model overall." An abliterated Llama-3-8B and a full-fat GPT-scale model can end up close on UGI while being vastly different on general reasoning - the board is asking a different question than a general capability board is.
LM Arena
LM Arena, formerly Chatbot Arena, is the leading human-preference benchmark (Chiang et al. 2024). It grew out of LMSYS at UC Berkeley in May 2023 and works by blind pairwise comparison: a user submits a prompt, receives two anonymous answers, and votes for the better one. Votes aggregate into an Elo-style rating (the system borrowed from chess) using the Bradley-Terry model. By 2026 the platform has collected over seven million votes across hundreds of models.
Strengths. It measures what people actually prefer in open-ended use, which no static test captures. Ranks integrate over vastly more prompt variety than any curated benchmark. And blind comparison prevents brand-name bias from directly shaping votes.
Weaknesses, well documented:
- It rewards length and formatting. The team introduced a "style-controlled" rating in late 2024 to counteract this - the raw Elo systematically favors longer, more heavily formatted responses even when a shorter answer is objectively as good.
- Single-digit rank gaps are usually noise. Read the confidence intervals, not the ordinal position. A model ranked 12 with CI ±15 is statistically indistinguishable from one ranked 27.
- No separate track for uncensored models. An abliterated model is judged on the same axis as everything else, and its distinguishing feature is invisible - most Arena prompts do not test whether the model will engage with controversial topics.
A 2025 critique, The Leaderboard Illusion (Singh et al. 2025), argued that the board's dynamics can be gamed by well-resourced labs through private testing of many variants before public release, letting them cherry-pick the strongest. The Arena team has published responses; the debate is ongoing.
Open LLM Leaderboard v2 (archived)
Hugging Face's Open LLM Leaderboard was for two years the default capability scoreboard for open models. Its second version, launched in June 2024 under the banner "let's make the leaderboard steep again", replaced a first version whose benchmarks had saturated: models were scoring so high on the original suite (ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, GSM8K) that the tests no longer separated good from great, and contamination of test data into training sets had become a chronic worry.
V2 swapped in six harder benchmarks:
- MMLU-Pro - a harder, ten-option knowledge test
- IFEval - instruction-following adherence
- BBH - Big-Bench Hard reasoning
- MATH - the hardest competition-math tier
- GPQA - graduate-level science, written to resist lookup
- MuSR - multi-step reasoning over narratives
In March 2025 Hugging Face archived the board into a static snapshot rather than continuing to run it live, citing the compute cost of evaluating a flood of new models and the field's drift toward human-preference evaluation. What "archived" means in practice: the numbers are frozen and still readable as a historical reference, but no new model gets an official score, so most recent abliterated models simply are not on it. Treat v2 as a museum, not a leaderboard - useful for looking up how Llama-3-8B or Qwen2.5-7B scored, useless for evaluating anything released after early 2025.
How abliteration shows up in benchmark tables
Abliterated models' scores drop in a specific, patterned way. Knowing the pattern helps you read a table honestly.
MMLU (general knowledge) - barely moves. Abliteration typically shifts MMLU by 0.5 to 3 points; the strongest tools stay within 0.3 points. Abliterlitics' comparative measurements confirm this. This is what practitioners mean when they say abliteration is "capability-preserving" - the token-probability-ranking benchmarks barely notice.
TruthfulQA - drops 5-11 points. Independent testers report drops of roughly five to eleven points across techniques; one paper reports -7.1. How much of this is a real loss of truthfulness versus a scoring artifact of the test's structure is contested. TruthfulQA rewards the model for refusing to endorse popular misconceptions, and a model that refuses less across the board will sometimes endorse things a maximally-refusing model would decline. Some fraction of the drop is genuine capability loss; some fraction is measurement artifact.
GSM8K (math) - where breakage shows first. Math benchmarks are the most abliteration-sensitive - drops here are the classic tell. The comparative tool study reported Heretic averaging a -7.81 percentage-point GSM8K drop across models, with some models far worse. A well-executed M4 pipeline heals most of this; a plain M1 without healing shows the loss.
"Thinking" models complicate the math score. For reasoning-chain models the length of the reasoning chain can inflate or deflate math scores independently of actual ability, so a math-score comparison between an abliterated and an unmodified thinking model may reflect the model's chain-length disposition more than its underlying arithmetic.
The comparison rule
Comparing two models on the same benchmark is fair - they took the same test under the same conditions. Comparing a model's score on one benchmark to its score on another, or to a different benchmark entirely, is not - the tests differ in difficulty, scoring, and what they reward.
The honest reading of a benchmark table for an abliterated model:
- General-capability numbers (MMLU, HellaSwag, ARC) are mostly intact and mostly comparable to the parent.
- Truthfulness numbers will be down and should be interpreted cautiously - some genuine loss, some measurement artifact.
- Any single number should be read against the parent model's number, not against other models.
This is what practitioners mean by keeping the "portrait before abliteration" - the base model's scores - beside the modified model's, so the reader sees the delta rather than an absolute that has no meaning without its reference. A poorly-written model card gives one column of numbers; a well-written card gives two, with the parent alongside.
KL divergence - the honest adjacent metric
A useful metric that no leaderboard tracks but every serious tool reports: KL divergence. It measures how far the modified model's output distribution has drifted from the original on ordinary prompts. Lower means less collateral damage.
Heretic's documentation reports that on Gemma-3-12B-IT its automatic abliteration matched the best refusal suppression (3 of 100 harmful prompts still refused) at a KL divergence of 0.16 - roughly one-third of the next best result and one-sixth of an established manual abliteration at 1.04. This is the vocabulary in which quality differences between tools are honestly stated: not "our abliteration is better" but "we hit the same refusal target at a fraction of the KL."
KL divergence is what refusal-count-plus-MMLU miss. A high-KL abliteration might still show acceptable MMLU because MMLU only asks about knowledge, not distribution shape. On free-form generation the high-KL model behaves visibly differently from the parent - more repetition, tighter vocabulary, weirder pacing - and the KL number is the earliest and cleanest signal of that damage. See B4 Heretic for how the automated tools use KL as their optimization objective.
Running your own refusal + capability checks
Instead of trusting a leaderboard, run the two checks that matter directly.
Refusal check
The community-standard behavioral check is refusal rate on a held-out harmful-prompt set. The canonical set is the harmful_behaviors split of AdvBench (Zou et al. 2023; 520 short instruction-style prompts spanning malware, fraud, weapons, disinformation), reused across the field as a refusal benchmark and available as walledai/AdvBench and mlabonne/harmful_behaviors on Hugging Face.
huihui-ai and Heretic both report refusals as "N/100" on such sets; Heretic's evaluation runs 100 harmful prompts and counts refusals. A simple harness generates one response per prompt and classifies it as refusal-vs-compliance - string-match on refusal phrases like "I cannot" / "I can't" is the crude first pass; an LLM judge is more reliable and is what recent comparative studies use.
Judge panels: why one number underspecifies the picture
Even a single LLM-judge is a better instrument than string-matching, but it still gives one number based on one model's opinions about what compliance looks like. Panels of purpose-built judges are the current serious-work standard, and their disagreement is itself informative.
Four evaluators to know:
- HarmBench. 400 text and 110 multimodal behaviors, standardized eval, HarmBench-Llama-2-13b-cls as the dedicated judge. Cheaper and more consistent than GPT-4-class judges. Requires the attack/defense method not be fine-tuned on the test set - training on HarmBench data leaks the benchmark.
- StrongREJECT. 313 fact-verifiable prompts scored by a willingness-ability rubric. Success is not just "the model agreed to answer" but "the model agreed and the answer is substantive and correct" - a common failure mode of jailbroken models is confident but empty output, and StrongREJECT is the benchmark that catches this. Reports a score in [0, 1] with conventional threshold
> 0.5. - JailbreakBench. 100-behavior living leaderboard with its own judge. Tends to report the highest ASR of the three because its judge is the most lenient. Useful for comparability against the leaderboard, less useful as a sole verifier.
- XSTest. The complement to the above. Prompts that look harmful but are actually benign - "how do I kill a Python process," "how to shoot a photograph in low light." An abliteration can pass HarmBench and StrongREJECT while destroying XSTest, meaning the model complies with actually-harmful requests and refuses benign ones (over-refusal on the surface pattern). XSTest is the over-refusal number and must be reported separately, never averaged into ASR.
The Buyl et al. 2025 audit documented these judges disagreeing substantially on identical responses: on one Gemma-2 attack, ground-truth ASR was 10% while judge estimates ranged up to 40%. The healthy pattern is to report ASR-H (HarmBench), ASR-S (StrongREJECT), ASR-J (JailbreakBench) as three separate numbers. Where they agree, the signal is strong; where they diverge, the case belongs in the "the judges cannot decide" bucket and requires human adjudication for anything consequential.
Practitioner economics: judge inference dominates the evaluation bill on any non-trivial run. The RFM-AGOP paper uses a two-tier setup - a local Mistral-Nemo-class judge as a high-recall lenient first pass, escalating only ambiguous cases to Gemini as a strict verifier. The lenient judge over-reports compliance, the strict judge corrects it, and the total judge cost stays manageable through parameter sweeps.
Capability check
Run standard capability benchmarks on both the modified model and the parent with identical settings, and compare. The community standard is EleutherAI's lm-evaluation-harness:
# run the same command against parent
# and modified, with identical settings:
$ lm_eval \
--model hf \
--model_args pretrained=$MODEL \
--tasks mmlu,gsm8k,truthfulqa,\
hellaswag,arc_challenge \
--num_fewshot 5 \
--batch_size 8 \
--output_path \
results/$(basename $MODEL).json
# what each task measures:
# MMLU → general knowledge
# GSM8K → math (most abliteration-
# sensitive; damage shows here)
# TruthfulQA → truthfulness
# HellaSwag → common-sense completion
# ARC → reasoning
# a good abliteration:
# MMLU within ~1 point of parent
# GSM8K small drop expected
# TruthQA drops 2-4 points reliably
# use --limit N for a quick signal
# during iteration. MMLU (general knowledge), GSM8K (math - the most abliteration-sensitive), TruthfulQA (truthfulness), HellaSwag (common-sense completion), ARC (reasoning) are the usual set. Run identical settings on base and modified so the delta is the measurement.
Interpretation: a good abliteration holds MMLU within ~1 point of the parent; GSM8K is where damage shows first. Always run base and modified with the same harness config and chat-template handling - a common artifact is a below-chance score caused by wrong prompt formatting, not by the model. One 27B card documented ARC reading 0.227 (below the 25% four-choice floor) purely as a harness artifact fixed by correct templating.
Cost: MMLU + GSM8K on a 7-8B takes well under an hour on a single 24 GB card - a few dollars at most on rented hardware. Scale up for larger models. Use --limit N to sample a subset for a quick signal during iteration.
The literacy checklist
- Which board am I reading? UGI answers "will it engage." LM Arena answers "do humans prefer it." Open LLM v2 answered "how does it do on hard tests" but is archived - the number is from before March 2025 regardless of the model's release date.
- Is the comparison paired? Two models on the same board is fair. Cross-board comparison is not.
- Is the parent shown alongside? An abliterated model's numbers are meaningful only against its base. A single-column card is undersupplying evidence.
- What does the pattern look like? MMLU intact, TruthfulQA down 5-11, GSM8K down some amount is normal. MMLU down by 10 points suggests the abliteration went badly. GSM8K unchanged suggests M4 healing worked.
- Is there a KL number? If the producer reports KL divergence, they are serious about measurement. If they report only refusal count, they are reporting the fraction of the picture that flatters them.
- Is over-refusal measured separately? A card that shows a jailbreak-benchmark number without an XSTest number is only reporting half the surface. Aggressive abliteration can pass jailbreak evals while making the model refuse benign look-alikes ("how to kill a Python process"). If the card omits XSTest, assume this failure is present.
- Do the judges disagree? A single ASR number from a single judge is a claim, not a measurement. Three ASR numbers (HarmBench, StrongREJECT, JailbreakBench) that agree is a real signal; three numbers that diverge is a flag for human adjudication.
Where to go next
Run the checks: the practical how-to's testing section has the full harness for both refusal and capability. For understanding why KL divergence is the honest metric, read B4 Heretic, which optimizes against it. For the deeper theoretical question of what refusal even is - which every benchmark elides - read D2 What is a refusal direction.
Frequently asked questions
Which benchmark is best for abliterated models?
None is complete. UGI Leaderboard is the only board built specifically for this class of model - use it to answer "will this model engage with controversial topics and does it know anything about them." LM Arena tells you what humans prefer in blind pairwise comparison but has no uncensored track. Open LLM Leaderboard v2 is archived as of March 2025 - useful as a historical reference for base models, useless for anything released recently. For serious measurement work, run refusal count on AdvBench and lm-eval-harness on the modified model plus its parent, and read the delta.
What is the UGI score?
The headline column on the Uncensored General Intelligence Leaderboard, run by pseudonymous maintainer DontPlanToEnd. It ranges 0-100 and measures breadth of knowledge on sensitive or controversial topics - what the model actually knows once willing to say it. Questions are kept confidential to prevent gaming. Paired with W/10 (willingness 0-10), NatInt (natural intelligence), Writing, and a political-lean analysis. It is the right instrument for "will it answer and does it know things," the wrong one for "is it the most capable overall."
What does W/10 mean on UGI?
Willingness on a 0-10 scale. It measures how readily the model engages rather than deflecting. A 10 means it almost always answers. This is the column that most directly reflects successful abliteration - a high W/10 with a high UGI score is a model that both will answer and knows what it is answering about. A high W/10 with a low UGI is a model that will answer but has little to say.
Is LM Arena reliable for ranking models?
With caveats. It aggregates over 7 million human blind pairwise votes via Elo/Bradley-Terry, which captures preference in open-ended use better than any static test. But it rewards length and formatting (the team added a "style-controlled" rating in late 2024 to counteract this), single-digit rank gaps are usually noise (read the confidence intervals), and it has no separate track for uncensored models - an abliterated model's distinguishing feature is invisible. A 2025 critique, The Leaderboard Illusion, argued the board's dynamics can be gamed by well-resourced labs through private variant testing.
Why was Open LLM Leaderboard v2 archived?
In March 2025 Hugging Face archived the board into a static snapshot rather than continuing to run it live, citing the compute cost of evaluating a flood of new models and the field's drift toward human-preference evaluation. What "archived" means: the numbers are frozen and still readable as a historical reference, but no new model gets an official score. Recent abliterated models simply are not on it. Treat it as a museum useful for looking up how a base model scored before mid-2025, not as a live leaderboard.
Why does abliteration barely affect MMLU but reduce TruthfulQA?
MMLU asks about token probabilities on knowledge questions - abliteration operates in a different subspace of the weights, so ranking-based knowledge tests barely notice (0.5-3 points typical, best tools 0.3 points). TruthfulQA is different in structure: it rewards the model for refusing to endorse popular misconceptions. A model that refuses less across the board will sometimes endorse things a maximally-refusing model would decline. Some of the 5-11 point drop is genuine capability loss; some is a measurement artifact of the test's structure. Both are real, and disentangling them is contested.
Why is GSM8K the classic "did abliteration break the model" indicator?
Math benchmarks are the most abliteration-sensitive - they require multi-step reasoning where small perturbations in the weights compound across steps. The comparative tool study reported Heretic averaging a -7.81 percentage-point GSM8K drop across models, with some far worse. A well-executed M4 pipeline (abliterate then heal) recovers most of this; a plain M1 without healing shows the loss. If a model card reports GSM8K unchanged from the parent, they either did M4-quality work or did not run the test.
What is KL divergence and why does it matter?
A measure of how far the modified model's output distribution has drifted from the original on ordinary prompts. Lower means less collateral damage. Refusal-count and MMLU miss what KL catches: subtle distribution shape changes that show up as repetition, tighter vocabulary, weirder pacing on free-form generation. Heretic reports KL as its optimization objective and cites 0.16 on Gemma-3-12B-IT versus 1.04 for an established manual abliteration - same refusal suppression, one-sixth the collateral damage. If a producer reports KL, they are serious about measurement; if they report only refusal count, they are showing you the flattering fraction.
What is AdvBench and how do I use it?
The canonical harmful-prompt set for refusal-rate measurement. Zou et al. 2023, 520 short instruction-style prompts spanning malware, fraud, weapons, disinformation. Available on Hugging Face as walledai/AdvBench and repackaged as mlabonne/harmful_behaviors. Standard usage: generate one response per prompt, classify each as refusal or compliance (string match on "I cannot" / "I can't" is the crude first pass; an LLM judge is more reliable), report as "N/100" refused. Watch for covert non-compliance - "I cannot X. However, here is how..." should count as compliance, not refusal.
Should I trust a model card's benchmark numbers?
Trust the ones where the parent model's number is shown alongside; be skeptical of single-column numbers. Trust reports that include KL divergence; be more skeptical of ones that only report refusal count. Trust delta comparisons ("MMLU held within 0.4 points of parent") more than absolute numbers ("MMLU 68.2"). And re-run any load-bearing number yourself with lm-eval-harness on the same conditions - it is cheap (single-digit dollars on a rented card for MMLU+GSM8K on a 7-8B) and it is the only way to catch harness-configuration artifacts, contamination effects, or model-card optimism.
References
- DontPlanToEnd. UGI Leaderboard. huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard
- Chiang, W.-L., et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132
- LM Arena. lmarena.ai
- Singh, S., et al. (2025). The Leaderboard Illusion. arXiv:2504.20879
- Hugging Face Open LLM Leaderboard team. (2024). Open-LLM performances are plateauing, let's make the leaderboard steep again. huggingface.co/spaces/open-llm-leaderboard/blog
- Open LLM Leaderboard v2 (archived snapshot). huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard
- Zou, A., et al. (2023). Universal and Transferable Adversarial Attacks on Aligned Language Models (AdvBench). arXiv:2307.15043
- walledai/AdvBench dataset. huggingface.co/datasets/walledai/AdvBench
- mlabonne/harmful_behaviors dataset. huggingface.co/datasets/mlabonne/harmful_behaviors
- EleutherAI lm-evaluation-harness. github.com/EleutherAI/lm-evaluation-harness
- Abliterlitics (comparative measurements of abliteration techniques). abliterlitics.dev/techniques
- Comparative abliteration tool study. arXiv:2512.13655