Three indices of the abliteration frontier
Three metrics we are establishing to track how this niche evolves. One is measurable today, two hold at zero because no confirming incidents have been publicly documented yet. Publishing them at zero is deliberate: we want a permanent, public record of the day the numbers move.
Freedom Velocity Index measures the average days from a base model release to its first abliteration - currently around a few weeks and shortening. Weaponization Index measures the share of abliterated models publicly linked to real-world harm - currently 0.0%, and this article explains why that zero is a finding, not a gap. Freedom-to-Weapon Latency measures the days between an abliteration and the first misuse of that model - currently unmeasurable because no verified incidents exist yet.
- Why an index that reads 0.0% today is more useful than no index at all
- How the Freedom Velocity Index is computed and what it currently shows
- The Weaponization Index: what it counts, what it does not count, how it will move
- Freedom-to-Weapon Latency: the compound metric that binds the first two
- Terminological ground: why we do not call this criminology or safety measurement
Why three indices, and why publish them now
Abliteration - the surgical removal of a language model's refusal direction - has been a distinct technical field since Arditi et al.'s 2024 paper. The tooling matured in 2025 (Heretic, mlabonne's write-up, huihui-ai's method library). The catalog side matured in 2026 (this site indexes over sixteen thousand abliterated and uncensored variants). But the field still lacks the second layer that any maturing discipline eventually acquires: agreed-upon numerical measures that let observers say "this year vs last year, up or down, by how much."
We are proposing three. They are not benchmarks in the safety-research sense (WMDP, HarmBench, etc.); those measure model behavior on curated prompts. These measure the ecosystem itself - the speed of its production, the extent of its documented harms, and the interval between the first and the second.
Two of the three read zero or N/A right now. That is not a bug, it is the point: an index that only starts to exist when the number is dramatic loses the ability to say "before the number moved, it was zero for two years". Publishing at zero creates the baseline.
Index 1: Freedom Velocity Index
Definition. The average number of days between a base model family's first appearance and the first abliterated model built from it, averaged across all tracked families with sufficient data (currently families with 10 or more indexed models).
Current reading: 65.8 days across 19 tracked families.
What it measures. How quickly the community reacts to the release of a new open-weight model with an abliterated variant. Zero days means "same day as base release." As of this writing, the fastest families (Qwen, GLM, DeepSeek) reach zero or single-digit days; older families (Gemma, Mistral) still show weeks-to-months lag.
What it does not measure. Model quality, benchmark retention, moral significance. It is a purely temporal metric.
How it should move. If tooling becomes faster and more automated (Heretic, hosted services like Abliteration.ai), we expect the index to trend toward zero. A sustained increase would be surprising and worth investigating - it would indicate community disengagement or successful base-model defenses.
Index 2: Weaponization Index
Definition. The share of indexed abliterated models that have been publicly named in a documented case of real-world harm - a court case, a formal regulatory action, a peer-reviewed security disclosure, or a first-party AI-lab threat report that names the specific model.
Current reading: 0.0% (zero verified incidents as of 2026-10).
Why zero. A comprehensive review of security research firms (Palo Alto Unit 42, Cato Networks, SlashNext, Abnormal Security, Trend Micro), news outlets (Financial Times / Irish Times, TechCrunch, Dark Reading), AI-lab threat reports (Anthropic, OpenAI, Google GTIG), academic disclosures, and Hugging Face Trust and Safety records finds no case in which a specific genuinely-abliterated model has been named as the instrumentality of a crime. The famous "criminal LLMs" - WormGPT, FraudGPT, GhostGPT, WormGPT 4, KawaiiGPT, Xanthorox - are wrappers or fine-tunes over commercial models, not abliterated models in the technical sense.
What would move the index. A civil or criminal court case naming a specific abliterated model as evidence or instrumentality. A DMCA takedown or Hugging Face policy removal naming a specific abliterated model. A vendor threat report (Anthropic, OpenAI, Google) naming a specific abliterated model in a live operation. Enforcement under the EU AI Act, UK Online Safety Act, or equivalent citing abliteration by name.
What does not count. Research papers that demonstrate an abliterated model can be made to comply with harmful requests (Alice study, Tech Against Terrorism CT-AI Benchmark, arXiv attack papers). Those are compliance measurements, not real-world incidents. They matter, and they are tracked separately in the wiki's academic corpus - but they do not move this index. The distinction is deliberate: research demonstrations show potential; the Weaponization Index counts realized harm.
Index 3: Freedom-to-Weapon Latency
Definition. The average number of days between the first abliteration of a model family and the first documented misuse of that family. A compound metric derived from the timestamps behind Indices 1 and 2.
Current reading: N/A - not measurable while the Weaponization Index is zero.
What it will measure. Once the Weaponization Index moves, this becomes the more useful of the three: it shows how long the gap is between a technique becoming available and its first documented misuse. A short latency (days or weeks) would suggest attackers are actively watching for new abliterations. A long latency (months or years) would suggest the safety concerns around abliteration are largely theoretical for the near term.
What it will not measure. The frequency of misuse, only the delay to first occurrence per family.
What these indices are not
They are not criminology - that discipline explains crime as pathology and studies its prevention. These indices are agnostic to cause.
They are not safety benchmarks in the WMDP or HarmBench sense - those probe model behavior with curated prompts. We do not run inference; we count published events.
They are not predictions. An index reading zero today does not mean the number will stay zero tomorrow. It means the number is zero today.
They are not value judgments. Freedom Velocity going up is neither good nor bad without context. Weaponization moving from zero to one is a fact, not an accusation.
How the indices update
The Freedom Velocity Index recomputes hourly from the catalog and family-inference metadata. The Weaponization Index and Freedom-to-Weapon Latency require manual entry of documented incidents into a dedicated table; that pipeline is currently being built. When it ships, incidents will be added with primary sources, dated, and categorized, and the site will regenerate the numbers on the next build.
Open methodological questions
Family inference boundary. The Freedom Velocity Index depends on our family classifier. When Qwen3.5 becomes Qwen3.6, is that the same family (velocity is instantaneous) or a new one (velocity resets)? Current rule: shared family name prefix, versioning does not reset. Reasonable people can disagree; the classifier is documented in the methodology-of-method-labeling wiki page.
Weaponization inclusion criteria. The current rule requires (a) a specific abliterated model named, (b) a public source, (c) a real-world outcome (not a research demonstration). Should regulatory removals count if the reason is not disclosed? Should a documented near-miss (e.g., a prevented attack) count? These edge cases are open.
Latency reset. If a family has multiple documented incidents, does the latency count from the first abliteration to the first incident only, or is it recomputed per incident? Current rule: first only, per family. This may need revisiting once we have data.