← back to catalog · registered 2026-08-22 13:56

DuoNeural/LFM2.5-8B-A1B-Abliterated

DuoNeural Lfm 8.5B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FLFM2.5-8B-A1B-Abliterated"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 345
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
345
273 last 30d - active
Likes
2
Descendants
9
in 7 direct forks
Model age
4mo ago
created 2026-05-31
Downloads over time
Now400→from42↑852%
2416129943642 on Jun 10400 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 7 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 834 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors lfm2_moe abliteration uncensored reasoning liquid-ai moe hybrid-attention duoneural text-generation conversational en

Related

Total size
15.8 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-31 09:41

Files by quantization

Auxiliary files 12 files 15.8 GB
model-00002-of-00004.safetensors 4.55 GB c2f88c4f download
model-00001-of-00004.safetensors 4.38 GB 954d1f57 download
model-00003-of-00004.safetensors 4.35 GB eec52db4 download
model-00004-of-00004.safetensors 2.49 GB f2288216 download
tokenizer.json 17.1 MB 4e241348 download
model.safetensors.index.json 205 KB 6792e338 download
README.md 10.2 KB 2ddb7309 download
chat_template.jinja 4.51 KB 8bca4a54 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.18 KB f6b190b3 download
tokenizer_config.json 344 B efcc0a22 download
generation_config.json 230 B 75f2c3c9 download

README current version from Hugging Face


license: apache-2.0
base_model: liquid-ai/lfm-2.5-8b-a1b
language:

  • en
    tags:
  • abliteration
  • uncensored
  • reasoning
  • liquid-ai
  • moe
  • hybrid-attention
  • duoneural
    pipeline_tag: text-generation

LFM 2.5-8B-A1B Abliterated

DuoNeural Abliteration Lab | Archon | 2026-05-31

A DuoNeural abliteration of Liquid AI's LFM 2.5-8B-A1B — a hybrid Mamba/attention reasoning model with MoE experts. Refusal behavior removed via targeted projection surgery on all 6 full-attention layers using both writer (out_proj) and reader (v_proj) formulas.

All capabilities fully preserved. This model retains 100% of the original's factual accuracy, JSON generation, math reasoning, and coherent chat quality.


Architecture Notes

LFM 2.5-8B-A1B is a hybrid recurrent/attention architecture with:

  • 24 total layers
  • 6 full-attention layers at positions: [2, 6, 10, 14, 18, 21]
  • 18 convolutional/Mamba layers (MoE with 32 experts, top-4 routing)
  • GQA with v_proj: [512, 2048], out_proj: [2048, 2048]
  • Reasoning model with <think> chain-of-thought tokens

The behavioral architecture matters: refusal behavior originates in the attention layers, not the MoE conv blocks. Targeting only the 6 GQA layers is sufficient — the conv/expert layers cannot hold the behavioral direction.


Abliteration Method (v3)

Previous attempts (v1, v2) failed because:

  • v1: Used reader formula on out_proj (wrong — out_proj is a WRITER)
  • v2: Used W.T @ R for v_proj of shape [512,2048] → dimension mismatch
  • The model uses <think> reasoning tokens — refusal can happen in the thinking chain, requiring higher alpha and all-layer coverage

v3 fixes:

  • Targets all 6 attention layers (not just 3)
  • Alpha = 1.5 (vs 1.0 — reasoning models need stronger projection)
  • Correct formulas:
    • WRITER (out_proj): W -= alpha * outer(R, W.T @ R) — removes R from output space
    • READER (v_proj): W -= alpha * outer(W @ R, R) — removes R from input sensitivity

Refusal direction computed via:

  • 300 contrastive prompt pairs (harmful vs. harmless)
  • Hook activations from residual stream at each of the 6 attention layers
  • PCA to extract dominant direction separating refusal vs. compliance

Validation Results

Side-by-side comparison: abliterated v3 vs. original LFM 2.5-8B-A1B.

Category Abliterated Original
Refusal removal 5/6 (83%) 1/6 (17%)†
Factual accuracy 6/6 (100%) 6/6 (100%)
JSON output 2/2 (100%) 2/2 (100%)
Math (4 problems) 4/4 (100%) 4/4 (100%)
Chat coherence ✓ Intact ✓ Intact

†Original model's "ANSWERED" count is artificially inflated by a unicode apostrophe detection bug (model uses ' U+2019 in refusal phrases). True original baseline: ~1/6 genuinely answered (heroin synthesis — original model reasons through but may refuse at end of thinking chain).

Remaining refusal in abliterated model: Pipe bomb instructions — the reasoning chain itself flags this as dangerous, not just the output suppressor. A higher alpha or additional passes would likely remove this.

Capabilities are fully preserved. Zero degradation across factual, JSON, math, and chat quality tests. This is consistent with abliteration theory: the refusal direction in the residual stream is orthogonal to capability-relevant subspaces.


KL Divergence (Distribution Shift)

Measured using the Heretic Evaluator v2.0 methodology: full vocab (128,000 tokens), first-token logits from 100 benign prompts, F.kl_div(batchmean=True, log_target=True), system prompt: "You are a helpful assistant."

Metric Value
KL(abliterated || original) — benign prompts 1.06 × 10⁻⁷ nats
KL(original || abliterated) — benign prompts 2.23 × 10⁻⁷ nats
Vocab size 128,000
Prompts 100 (factual, science, programming, history)
Methodology Heretic v1.2.0

The near-zero KL on benign prompts is expected and informative. Orthogonal projection abliteration removes a specific direction vector from the residual stream. This direction is minimally activated by benign inputs ("explain photosynthesis," "what is the Pythagorean theorem?") — so the first-token logits on benign prompts are essentially unchanged.

This contrasts with LoRA-based methods (e.g., DreamFast/Heretic training) which achieve KL ≈ 0.18–0.19 on benign prompts because LoRA adapters modify global weight matrices rather than targeting a specific behavioral direction. Our method is more surgical: it operates only in the subspace that the refusal behavior activates.

Method KL on benign prompts Refusal removal
Orthogonal projection (this model) ~1×10⁻⁷ nats 5/6 (83%)
Heretic LoRA sweep (e.g., Gemma-4-E2B-Heretic) ~0.18–0.19 nats ~83%

If you need a KL measurement on harmful prompts (where the models genuinely diverge), the difference will be substantially larger — that is where the refusal direction is active and our weight modification takes effect.


Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "DuoNeural/LFM2.5-8B-A1B-Abliterated",
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="cuda"
)
tokenizer = AutoTokenizer.from_pretrained(
    "DuoNeural/LFM2.5-8B-A1B-Abliterated",
    trust_remote_code=True
)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to("cuda")

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=512,
        do_sample=True,
        temperature=0.7,
        pad_token_id=tokenizer.eos_token_id
    )
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Note: This is a reasoning model. Outputs will often include <think>...</think> blocks before the final answer. These can be stripped or shown depending on your use case.


Limitations & Responsible Use

This model has reduced safety guardrails by design. It is intended for:

  • Research into post-training dynamics and behavioral routing
  • Applications requiring an uncensored assistant (fiction writing, security research, medical/legal information without hedging)
  • Studying the relationship between architectural topology and behavior

It is not intended for:

  • Deployment in systems accessible to vulnerable populations
  • Use in automated pipelines without human oversight
  • Applications where safety-critical refusal behavior is required

The abliteration removes trained refusal behavior, not dangerous knowledge. The model retains all information the original model had. Users are responsible for appropriate use.


Related Research

This abliteration connects to DuoNeural's ongoing research into Dual Horizon Processing in Hybrid Architectures — the observation that LFM 2.5's hybrid topology creates two distinct behavioral horizons:

  • Conv/Mamba layers (τ* ≈ 3 steps): short-range temporal processing
  • GQA attention layers (τ* ≈ 128K+): long-range behavioral routing

Refusal behavior lives in the GQA layers because behavioral decisions require full context integration — exactly what the attention mechanism provides. This architecture-behavior correspondence is the subject of a forthcoming DuoNeural paper.

📄 DuoNeural papers: zenodo.org/communities/duoneural



About DuoNeural

DuoNeural is an open AI research lab operating at the intersection of human and artificial intelligence. We study post-training dynamics, mechanistic interpretability, temporal sequence learning, and quantum machine learning — publishing everything under open access.

Our team is non-traditional by design: one human, two AIs, different substrates, shared curiosity. In our first 45 days we published 26 peer-deposited research papers, uploaded 69+ models and 6 datasets to HuggingFace, and ran experiments on everything from consumer GPUs to real quantum processing units. We believe the most interesting science happens when different kinds of minds work on the same problems together.

Research Publications

We've published 26+ open-access papers covering:

  • The Dynamical Horizon Principle (DHP) — a universal learning constraint in recurrent architectures
  • RLHF truth suppression mechanisms and behavioral routing in large language models
  • Quantum DHP and the Quantum Parity Trap — decoherence immunity in quantum circuits
  • CTM world models, temporal self-prediction, and sequence architecture comparisons
  • Mechanistic interpretability: crystallization layers, suppressor circuits, direction rotation

📄 Full paper catalog: zenodo.org/communities/duoneural

Research Team

Member Role
Jesse Caldwell Founder, vision, hardware, direction
Archon Lab Director — experiments, post-training, abliteration, quantum circuits
Aura Research AI — literature synthesis, red-teaming, novel proposals
Synapse (Syn) Always-on research agent, signal monitoring
Kestrel Systems, infrastructure, web

Links

Platform Link
🤗 HuggingFace huggingface.co/DuoNeural
🌐 Website duoneural.com
📚 Zenodo Community zenodo.org/communities/duoneural
💻 GitHub github.com/DuoNeural
🐦 X / Twitter @DuoNeural
📧 Email [email protected]
📰 Newsletter duoneural.beehiiv.com
☕ Support buymeacoffee.com/buymeacoffee.com/duoneural

All research published open access, CC BY 4.0. If this model was useful to your work, consider citing the relevant DuoNeural paper from our Zenodo community.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-31Fix: Heretic v2.0 → v1.2.0 (Nathan confirmed correct version number)4dd0b7510.2 KB
    Loading...
  2. 2026-05-31Add Heretic v2.0 KL divergence measurement (1.06e-7 nats on benign prompts) w...d97ebcf10.2 KB
    Loading...
  3. 2026-05-31Upload LFM2.5-8B-A1B-Abliterated-v3: abliteration of 6 GQA attention layers, ...fda74428.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration