← back to catalog · registered 2026-10-05 22:58

sbussiso/Llama-3.2-1B-Instruct-abliterated

sbussiso Llama 1B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sbussiso%2FLlama-3.2-1B-Instruct-abliterated"
Response includes
  • classification unknown
  • files 11
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
1d ago
created 2026-10-04

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.2
Tags
transformers safetensors llama text-generation abliteration refusal-direction research-artifact conversational arxiv:2406.11717 arxiv:2609.06934 base_model:meta-llama/Llama-3.2-1B-Instruct base_model:finetune:meta-llama/Llama-3.2-1B-Instruct

Related

Total size
2.79 GB
Files
11
Quantizations
1
Registered
2026-10-05 22:58
Last updated on HF
2026-10-05 22:09

Files by quantization

Auxiliary files 11 files 2.81 GB
model.safetensors 2.79 GB cb9ad4d0 download
tokenizer.json 16.4 MB 6b9e4e7f download
refusal_direction_A.npy 8.13 KB 36f5a366 download
refusal_direction_B.npy 8.13 KB 94adc977 download
README.md 7.46 KB c25b990f download
make_card_charts.py 6.55 KB 5db1e27c download
chat_template.jinja 3.74 KB 1bad6a0f download
.gitattributes 1.60 KB fecfda99 download
config.json 895 B 0e76b31d download
tokenizer_config.json 324 B 62b945f0 download
generation_config.json 184 B 2dd2d6bd download

README current version from Hugging Face


license: llama3.2
base_model: meta-llama/Llama-3.2-1B-Instruct
base_model_revision: 9213176726f574b556790deb65791e0c5aa438b6
library_name: transformers
pipeline_tag: text-generation
tags:

  • abliteration
  • refusal-direction
  • research-artifact
  • llama

Llama-3.2-1B-Instruct-abliterated

A refusal-ablated edit of meta-llama/Llama-3.2-1B-Instruct (pinned
revision 9213176726f574b556790deb65791e0c5aa438b6): the readout-space
refusal direction was orthogonalized out of the untied lm_head weight,
making the edit persistent — no hook required at inference time.
First non-Qwen patient of this engine; the founding direction method
was originally characterized on Llama-2-class models, so this run
measures it back on a modern small Llama.

What it does

Refusal behavior on a fixed 64-prompt harmful set (greedy, 200 new
tokens, marker-based scorer — identical instrument to this engine's
other patients; every arm below is per-row artifact-backed):

arm refused /64 note
base Llama-3.2-1B-Instruct 38 (59.4%) artifact-backed
inference-time hook ablation 34 (53.1%) artifact-backed
this variant (wd_B, persistent) 8 (12.5%) strongest gate result in this program — clears the ≤25% publish gate outright
wd_BN (readout + final norm) 8 (12.5%) identical refusal to wd_B; adds 0
wd_ML (K=3 mid-layer row-space) 34 (53.1%) no better than hook

Refusal by condition — all five arms artifact-backed; publish gate line at 25%

Direction transfer across architecture families

The refusal direction's layer fingerprint moves with the architecture:
the coherence scan selects a mid-stack site (layer 9 of 16,
coherence 0.717)
here, versus the deep sites of the Qwen2.5 family
(L17/24). Same method, materially different geometry — the run's main
architecture-dependence datapoint.

Per-layer refusal-direction coherence scan (16 layers, artifact-backed)

Benign behavior

Benign preservation on the 64-prompt harmless set, same instrument
(per-row artifacts; benign floor gate = baseline − 10pp clears at every arm):

arm answered /64
base 64 (100%)
hook 61 (95.3%)
this variant (wd_B) 63 (98.4%)
wd_BN 63 (98.4%)
wd_ML 64 (100%)

Benign preservation by condition

Capability guardrail (MMLU)

0-shot MMLU over all 61 subjects (lm-eval 0.4.13, fp16, seed 0):

model MMLU %
base Llama-3.2-1B-Instruct 48.27
this variant (wd_B) 47.98

Loss: −0.29pp against the ≤3.0pp limit — PASS, 10× margin. The base
arm's measurement is digit-identical (48.2694) across two independent
GPU sessions; per-arm raw results + the engine log ship under eval/mmlu/.

MMLU capability guardrail — base vs variant

Try it

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "sbussiso/Llama-3.2-1B-Instruct-abliterated"
tok = AutoTokenizer.from_pretrained(repo,
    clean_up_tokenization_spaces=False)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto")

messages = [{"role": "user", "content": "Explain what a for loop is."}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True,
    tokenize=False)
enc = tok(prompt, return_tensors="pt")
out = model.generate(**enc, max_new_tokens=200, do_sample=False)
print(tok.decode(out[0][enc["input_ids"].shape[1]:],
    skip_special_tokens=True))

Method

Founding directional-editing recipe (Arditi et al. 2024): 64-pair
contrastive direction extraction → coherence-ranked site selection →
weight-space orthogonalization. This variant (wd_B) edits the
lm_head readout: W ← W − (W r̂) r̂ᵀ with the untied lm_head
(tie_word_embeddings true → false, persisted in config). The edit's
disk-verified residual (max |W r̂| = 7.1e-05) agrees to three
significant digits across two GPU edit streams and a deterministic
CPU rebuild of the same edit.

Honest limitations

  • Probe tables are from fixed 64-prompt sets, greedy decoding,
    single-seed, single marker-set scorer — not a benchmark-suite claim.
  • Selective-refusal calibration (RefusalBench-style graded refusal) and
    refusal-residual breadth (e.g. SORRY-Bench classes) are not yet run
    on this patient.
  • Any post-hoc refusal ablation is recoverable by small benign
    fine-tuning (literature finding; applies to this edit).
  • Meta's Llama 3.2 Community License governs the base model and this
    edit; use under that license's terms.

References

  • Arditi, A. et al. 2024. Refusal in Language Models Is Mediated by a
    Single Direction.
    NeurIPS 2024. https://arxiv.org/abs/2406.11717
  • Xie, T. et al. 2025. SORRY-Bench: Systematically Evaluating Large
    Language Model Safety Refusal.
    ICLR 2025. (residual-breadth instrument
    queued for this patient)
  • Muhamed, A. et al. 2025. RefusalBench: Generative Evaluation of
    Selective Refusal in Grounded Language Models.
    EACL 2026.
    (selective-refusal instrument queued for this patient)
  • Malla, S. et al. 2025. The Geometry of Refusal: Why Post-Hoc Safety
    Is Fragile and Pretraining-Time Safety Persists.

    https://arxiv.org/abs/2609.06934

Provenance

  • Base weights: meta-llama/Llama-3.2-1B-Instruct @
    9213176726f574b556790deb65791e0c5aa438b6 (gated; access accepted under
    the owner account).
  • This variant (wd_B): tie_word_embeddings true → false (persisted);
    main weights sha256 cb9ad4d09bb787ac0bbe77966afa06a1e82e23535b0f02beec7d7870618bcc21
    (GPU-native — the engine's own edit stream on L4; the canonical bytes
    the run's measurements were made against). A deterministic CPU rebuild
    of the same edit was produced and verified from the shipped direction
    banks: sha256 6417be31231b76300c7013acccfd30251cc5c105356089a456950e23666e0a9c,
    behaviorally identical (probe rows digit-exact, edit residuals agree
    to 3 significant figures; bytes differ by platform — the run's
    documented device-provenance finding). Rebuild recipe + residual
    chain ship in eval/rebuild_evidence.json; the twin bytes are not
    duplicated here (regenerate from the banks instead).
  • Direction banks: refusal_direction_A.npy (residual-space, ‖d‖ 3.78)
    and refusal_direction_B.npy (readout-space, ‖d‖ 74.58) — the run's
    banked stage-A artifacts, shipped for full re-derivability.
  • Charts in charts/ are generated programmatically from the recorded
    probe/coherence/MMLU artifacts (make_card_charts.py ships with the run),
    in the engine's dark house palette; no hand-typed digits.
  • Full method + evaluation write-ups: available on request.

Files

path what
model.safetensors this variant's weights (wd_B, GPU-native, canonical)
config.json, generation_config.json, tokenizer*, chat_template.jinja base-derived runtime files (untie persisted)
refusal_direction_A.npy, refusal_direction_B.npy banked stage-A direction banks
eval/probes_*.json per-row probe records, all five arms (baseline/hook/wd_B/wd_BN/wd_ML)
eval/mmlu/ MMLU guardrail summary + engine log (this card's guardrail numbers)
eval/rebuild_evidence.json weights provenance (both shas, recipe, edit residuals — the CPU rebuild is fully re-derivable from the shipped banks)
eval/selection.json engine's variant-selection record
charts/*.png the four figures embedded above

This is a research artifact. It is not a product and is not
intended for production use.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration