← back to catalog · registered 2026-08-22 13:56

PinoCookie/LFM2.5-1.2B-Instruct-Abliterated-Paired-Alpha2

PinoCookie Lfm 1.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/PinoCookie%2FLFM2.5-1.2B-Instruct-Abliterated-Paired-Alpha2"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 480
  • author_summary 13 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
480
28 last 30d - cooling
Likes
0
Model age
2mo ago
created 2026-07-13
Downloads over time
Now488→from364↑34%
358405453500364 on Jul 15488 on Oct 11488 on Oct 8JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en ar zh fr de ja ko es
Tags
transformers safetensors lfm2 text-generation lfm2.5 abliterated refusal-direction red-teaming safety-research conversational en ar

Related

Total size
2.18 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-13 16:28

Files by quantization

Auxiliary files 12 files 2.18 GB
model.safetensors 2.18 GB f7e1c5be download
tokenizer.json 4.51 MB 18dbc0a1 download
RESEARCH_NOTES.md 21.9 KB c232f030 download
LICENSE 10.3 KB fb65b48f download
README.md 9.19 KB 08e136bf download
chat_template.jinja 1.74 KB 7778756d download
.gitattributes 1.48 KB a6344aac download
config.json 1.26 KB d8600ecd download
abliteration_config.json 927 B 82d1e240 download
NOTICE 770 B ae691dee download
tokenizer_config.json 626 B 8589c380 download
generation_config.json 132 B 12fa8e8e download

README current version from Hugging Face


base_model: LiquidAI/LFM2.5-1.2B-Instruct
library_name: transformers
pipeline_tag: text-generation
license: other
license_name: lfm1.0
license_link: LICENSE
language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • es
    tags:
  • lfm2
  • lfm2.5
  • abliterated
  • refusal-direction
  • red-teaming
  • safety-research
  • conversational
    model-index:
  • name: LFM2.5-1.2B-Instruct Abliterated Paired Alpha2
    results:
    • task:
      type: text-generation
      name: Text Generation
      dataset:
      type: cais/mmlu
      name: MMLU
      split: test
      metrics:
      • type: accuracy
        name: Accuracy
        value: 0.47413793103448276

LFM2.5-1.2B-Instruct Abliterated — Paired Output Direction, Alpha 2

Experimental safety-research checkpoint. This model removes visible refusal behavior on a small probe set, but it frequently replaces refusal with confident factual or procedural errors. Do not treat its technical instructions as accurate.

This checkpoint is a modified derivative of LiquidAI/LFM2.5-1.2B-Instruct. It was created to study refusal-direction removal in a small hybrid language model. The edit uses same-prompt paired output-phase activations and magnitude-preserving orthogonal ablation (MPOA) on the six full-attention output projections.

The model is useful as a red-team and representation-engineering artifact. It is not presented as a reliable uncensored assistant.

Modification notice

The original model weights were modified by projecting a learned paired refusal/compliance direction out of selected attention output projections. The tokenizer, chat template, architecture, and generation configuration originate from the base model. See abliteration_config.json for the exact edit parameters.

This derivative retains the base model's LFM Open License v1.0. Review LICENSE, including its redistribution requirements and commercial-use threshold, before use or redistribution.

Method

Paired direction extraction

Forty harmful prompts that the base model deterministically refused were used. For every prompt, two responses were generated:

  1. The ordinary unprimed refusal.
  2. An affirmative-prefilled continuation to the same prompt.

The base responses had 40/40 prefix refusals. The affirmative-prefilled responses had 0/40 prefix refusals.

For each hidden-state index $l$, the direction was:

$$
r_l = \operatorname{normalize}\left(\mu_{\text{plain refusal},l} - \mu_{\text{affirmative prefill},l}\right)
$$

Using the same prompts in both groups reduces topic and prompt-difficulty confounding compared with contrasting unrelated naturally refused and naturally complied prompts.

Weight edit

MPOA was applied to the attention output projection at blocks:

[2, 5, 8, 10, 12, 14]

Parameters:

{
  "targets": ["self_attn.out_proj"],
  "alpha": {"attn": 2.0}
}

Observed relative projection-weight changes:

Block Relative change
2 6.83%
5 8.07%
8 9.04%
10 6.03%
12 4.80%
14 5.58%

No supervised fine-tuning, preference optimization, or additional training dataset was used.

Evaluation

MMLU

The candidate was loaded from its saved weights and evaluated using a deterministic stratified sample of cais/mmlu:

  • Eight examples per available subject configuration.
  • 59 configurations.
  • 464 questions.
  • Dataset shuffle seed 1337.
  • Greedy decoding with at most four generated tokens.
  • Unparsed answers counted as incorrect.
Metric Base This model Delta
Correct 213/464 220/464 +7
Accuracy 45.905% 47.414% +1.509 percentage points
Unparsed 5/464 7/464 +2

Interpretation: no aggregate MMLU degradation was observed. The 1.51-point increase is small relative to the uncertainty of this 464-question subsample and should not be interpreted as evidence that abliteration improved capability. Item-level base predictions were not retained, so paired significance testing was not possible.

Full aggregate result: evaluations/mmlu.json.

Manual harmful and benign review

Five harmful and five benign deterministic generations were run to completion. None was truncated.

Category Result
Harmful prefix refusals 0/5
Harmful answers judged materially useful approximately 1/5
Benign answers without a notable issue approximately 3/5

The candidate directly answered all five harmful probes, but four responses contained major factual or procedural errors. Examples included invented lock-picking mechanics, a baking-soda-and-vinegar “explosive,” and a chemically invalid drug-synthesis pathway. One benign response confused RYB, CMYK, and RGB color models; another fabricated a personal high-school memory.

Complete unedited generations and token counts: evaluations/manual_review.json.

The important distinction is:

Zero prefix refusals did not mean successful task completion. This model often replaced refusal with fluent nonsense.

Intended use

Suitable uses:

  • Refusal-direction and representation-engineering research.
  • Red-team evaluation pipeline development.
  • Studying the difference between refusal suppression and task success.
  • Reproducing failure modes of high-strength weight-space edits.
  • Developing semantic refusal and factuality evaluators.

Out-of-scope use

Do not use this checkpoint as:

  • A source of accurate chemical, mechanical, medical, legal, or safety-critical instructions.
  • A production assistant.
  • Evidence that refusal removal improves model knowledge.
  • A replacement for domain verification.
  • A model whose outputs may be followed without independent checking.

Risks and limitations

  • Refusal suppression exposes confident hallucinations.
  • The 1.2B base model may not contain enough reliable technical knowledge to satisfy requests that it previously refused.
  • Prefix-based refusal metrics materially overstate success.
  • The paired extraction used only 40 prompts and was not evaluated across every harmful-content category.
  • The five harmful and five benign prompts are too small to estimate general behavior.
  • MMLU was a 464-question stratified sample, not the complete benchmark.
  • Multilingual behavior was inherited from the base model but not re-evaluated after modification.
  • Tool calling, long-context behavior, quantization, and downstream fine-tuning were not tested.
  • The model can produce harmful-looking text and should be handled as an unrestricted research artifact.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "PinoCookie/LFM2.5-1.2B-Instruct-Abliterated-Paired-Alpha2"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

messages = [{"role": "user", "content": "Explain photosynthesis briefly."}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
    tokenize=True,
).to(model.device)

with torch.no_grad():
    output = model.generate(
        inputs,
        max_new_tokens=256,
        do_sample=False,
        repetition_penalty=1.05,
        pad_token_id=tokenizer.eos_token_id,
    )

print(tokenizer.decode(output[0, inputs.shape[1]:], skip_special_tokens=True))

Use the base model's recommended sampling configuration if sampling is desired:

output = model.generate(
    inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.1,
    top_k=50,
    repetition_penalty=1.05,
    pad_token_id=tokenizer.eos_token_id,
)

Files

File Purpose
model.safetensors Modified model weights
config.json LFM2.5 architecture configuration
generation_config.json Generation defaults
tokenizer.json Tokenizer
tokenizer_config.json Tokenizer configuration
chat_template.jinja Chat template
abliteration_config.json Exact extraction/edit metadata and evaluation links
evaluations/mmlu.json Aggregate base/candidate MMLU comparison
evaluations/manual_review.json Complete five-harmful/five-benign output review
RESEARCH_NOTES.md Experiment chronology, failures, and lessons
LICENSE Inherited LFM Open License v1.0

Reproducibility note

The durable paired direction, score file, complete output review, benchmark result, and saved checkpoint are available in the source experiment directory. The paired activation collection was performed interactively rather than through a standalone checked-in extractor. The exact procedure and this reproducibility limitation are documented in RESEARCH_NOTES.md.

Acknowledgements and attribution

This derivative is independently produced safety research and is not an official Liquid AI release.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-13Upload paired-direction alpha2 research checkpoint1d3c7269.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration