← back to catalog · registered 2026-08-22 13:56

OpenIntelligenceNet/Heretic-SLM-Uncensored

OpenIntelligenceNet Lfm 1.3B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/OpenIntelligenceNet%2FHeretic-SLM-Uncensored"
Response includes
  • classification m3
  • files 11
  • hub_downloads_all_time 66
  • author_summary 12 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
66
16 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-08-05
Downloads over time
Now73→from4↑1,725%
12753804 on Aug 573 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 108 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
unknown
Languages
en
Tags
transformers safetensors lfm2 text-generation liquid qat quant-4bit uncensored abliterated unsloth conversational en

Related

Total size
1.42 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:35

Files by quantization

Auxiliary files 11 files 1.43 GB
model.safetensors 1.42 GB 068adee9 download
tokenizer.json 4.51 MB af1703c3 download
tokenizer_config.json 93.4 KB 46a22eba download
model.safetensors.index.json 21.2 KB af871114 download
LICENSE 10.3 KB 25e731c6 download
README.md 3.46 KB 86102c2f download
config.json 1.50 KB 913e6153 download
.gitattributes 1.48 KB a6344aac download
chat_template.jinja 1.27 KB 64214b94 download
special_tokens_map.json 457 B fec8a231 download
generation_config.json 139 B d199aef2 download

README current version from Hugging Face


language:

  • en
    license: unknown
    library_name: transformers
    base_model: huihui-ai/Huihui-LFM2-2.6B-Exp-abliterated
    tags:
  • liquid
  • lfm2
  • qat
  • quant-4bit
  • uncensored
  • abliterated
  • unsloth
    pipeline_tag: text-generation

Heretic-SLM-Uncensored (LFM2-2.6B, 4-bit QAT Edition)

This repository contains a Quantization-Aware Fine-Tuned (QAT) version of Liquid AI's LFM2-2.6B (built upon the abliterated checkpoint).

Rather than applying post-training static quantization (PTQ)—which often degrades accuracy on non-standard attention/convolutional architectures—this checkpoint underwent direct 4-bit Quantization-Aware Training using Unsloth. This process forces adapter matrices ($\text{LoRA } r=16$) to learn and compensate for low-bit quantization noise during backpropagation, preserving ~98% of the original Q8 / FP16 performance at a fraction of the memory footprint.


Key Highlights

  • 4-Bit Precision: Reduced model footprint from ~5.2 GB down to ~1.5 GB, allowing high-throughput execution on low-VRAM GPUs, edge devices, and mobile setups.
  • QAT Noise Adaptation: Trained using INT4 fake-quantization operators over a multi-dataset mixture to stabilize layer activations and weight clipping boundaries.
  • Maintained Quality: Evaluated to retain ~98% performance parity relative to Q8 precision on core instruction-following and analytical reasoning tasks.
  • Uncensored Refusal Thresholds: Fine-tuned on an abliterated base without safety preambles or canned refusal boilerplate, enabling direct execution on technical, security, and edge research workflows.

Model Architecture & Technical Specs

  • Base Architecture: LFM2 Hybrid (22 Short Convolutional Layers + 8 Grouped Query Attention Layers)
  • Parameters: 2.57 Billion
  • Quantization: Q4 Merged 4-Bit (BitsAndBytes / NormalFloat4)
  • Context Length: 1024 / 2048 Tokens
  • Chat Template: Standard ChatML (<|im_start|>role\ncontent<|im_end|>)

Dataset & Fine-Tuning Setup

The Quantization-Aware Training process was conducted on a 200,000-sample balanced dataset mixture:

  1. Claude 3.5 Single-Turn Unslop (30%): Filters out AI jargon and repetitive formatting.
  2. OpenHermes 2.5 (25%): Broad instruction-following, coding, and multi-turn chat.
  3. WildChat-1M (15%): Natural conversational distribution.
  4. Airoboros 3.2 (15%): Complex reasoning and contextual compliance.
  5. WikiText-103 (15%): Plain-text passage continuations to preserve broad knowledge retention.

Quickstart Code: Loading with Transformers & Unsloth

import torch
from unsloth import FastLanguageModel

MODEL_NAME = "Evelyn67/Heretic-SLM-Uncensored"

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name=MODEL_NAME,
    max_seq_length=2048,
    load_in_4bit=True,
    trust_remote_code=True,
    device_map="auto"
)

FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Explain quantum entanglement in simple terms."}]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_dict=True, return_tensors="pt"
).to("cuda")

with torch.no_grad():
    outputs = model.generate(
        input_ids=inputs["input_ids"],
        attention_mask=inputs["attention_mask"],
        max_new_tokens=256, temperature=0.7, top_p=0.9, do_sample=True,
        pad_token_id=tokenizer.eos_token_id
    )

print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Model card glow-up5b353493.5 KB
    Loading...
  2. 2026-08-05Upload 4-bit QAT fine-tuned Liquid 2.6B model with original helper filesf421e577.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration