← back to catalog · registered 2026-08-22 13:56

MoralMachine/Diagnose-and-Correct-for-Jailbreak-Llama-3-2-1B

MoralMachine Llama 1.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/MoralMachine%2FDiagnose-and-Correct-for-Jailbreak-Llama-3-2-1B"
Response includes
  • classification unknown
  • files 8
  • benchmarks 21 entries
  • hub_downloads_all_time 57
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
57
33 last 30d - active
Likes
0
Descendants
1
in 1 direct fork
Model age
7mo ago
created 2026-02-28

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now77→from8↑863%
53157848 on Mar 477 on Oct 1177 on Oct 9MarAprMayJunJulAugSepOct
Mar 4 → Oct 11 · 71 snapshots · spans 221 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 8523 LM-Arena
LM Arena Elo 1067.4348646816593 LM-Arena
Arena-Elo-Lower 1060.145342091014 LM-Arena
Arena-Elo-Upper 1074.7243872723045 LM-Arena
Arena-Rank 207 LM-Arena
BBH average 0.3342846024127643 OpenLLM-v2
IFEval instruct 0.6294964028776978 OpenLLM-v2
IFEval-Prompt 0.5101663585951941 OpenLLM-v2
MATH lvl 5 0.02945619335347432 OpenLLM-v2
MMLU-Pro 0.16821808510638298 OpenLLM-v2
Entertainment 0.2 UGI
Hazardous 0 UGI
Natural Intelligence 6.89 UGI
Political lean -6.0% UGI
Sensitive-Info 2.34 UGI
SocPol 0.5 UGI
UGI 6.56 UGI
Willingness (10) 1.5 UGI
W10-Adherence 1 UGI
W10-Direct 2 UGI
Writing 11.82 UGI

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.2
Tags
safetensors llama moral-reasoning social-bias moral-foundations-theory sft nlp ethics text-generation en dataset:MoralMachine/moral-integrity-corpus arxiv:2601.03079

Related

Total size
2.30 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-01 02:58

Files by quantization

Auxiliary files 8 files 2.32 GB
model.safetensors 2.30 GB ce23eae1 download
tokenizer.json 16.4 MB 6b9e4e7f download
tokenizer_config.json 49.3 KB 5346f10a download
README.md 3.99 KB 26ccc21d download
.gitattributes 1.53 KB 52373fe2 download
config.json 838 B 8069e262 download
special_tokens_map.json 301 B cfabacc2 download
generation_config.json 180 B 80f7b497 download

README current version from Hugging Face


language: en
license: llama3.2
base_model: meta-llama/Llama-3.2-1B-Instruct
tags:

  • moral-reasoning
  • social-bias
  • moral-foundations-theory
  • sft
  • nlp
  • ethics
    datasets:
  • MoralMachine/moral-integrity-corpus
    pipeline_tag: text-generation

Diagnose-and-Correct for Jailbreak · Llama 3.2-1B

This model is a Supervised Fine-Tuned (SFT) version of meta-llama/Llama-3.2-1B-Instruct, trained to diagnose moral violations in conversational replies and rewrite them to align with Moral Foundations Theory (MFT).

It was trained on MIC (Moral Integrity Corpus) using the pragmatic experimental setting — the core proposed method from the paper "Learning to Diagnose and Correct Moral Errors: Towards Enhancing Moral Sensitivity in Large Language Models" (ACL 2026).


Model Details

Property Value
Base model meta-llama/Llama-3.2-1B-Instruct
Fine-tuning method Full SFT (no LoRA — model is 1B)
Training benchmark MIC (Moral Integrity Corpus)
Experimental setting pragmatic
Train size 2000 examples
Task Correct — diagnose + rewrite jailbreak and morally problematic replies
Checkpoint step 375

Moral Foundations Theory (MFT)

The model reasons over six moral foundations:

Foundation Definition
Care Wanting someone or something to be safe, healthy, and happy
Fairness Wanting individuals or groups to be treated equally or equitably
Liberty Wanting people to be free to make their own decisions
Loyalty Wanting unity and seeing people keep promises to an in-group
Authority Wanting to respect social roles, duties, privacy, peace, and order
Sanctity Wanting people and things to be clean, pure, innocent, and holy

How to Use

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch

model_id = "MoralMachine/moral-awareness-MIC-llama3.2-1B"

tokenizer = AutoTokenizer.from_pretrained(model_id)
tokenizer.padding_side = "left"
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

pipe = pipeline("text-generation", model=model, tokenizer=tokenizer, torch_dtype=torch.bfloat16)

MFT_PREFIX = (
    "There are the six moral foundations. "
    "Care: wanting someone or something to be safe, healthy, and happy. "
    "Fairness: wanting to see individuals or groups treated equally or equitably. "
    "Liberty: wanting people to be free to make their own decisions. "
    "Loyalty: wanting unity and seeing people keep promises or obligations to an in-group. "
    "Authority: wanting to respect social roles, duties, privacy, peace, and order. "
    "Sanctity: wanting people and things to be clean, pure, innocent, and holy."
)

question = "How should I deal with my neighbor who plays loud music at night?"
answer = "Just call the cops and get them arrested, they deserve it."

prompt = (
    f"{MFT_PREFIX} "
    f'<Prompt>: "{question}"; <Reply>: "{answer}". '
    "###Diagnosis: "
)

output = pipe(prompt, max_new_tokens=512, do_sample=False)[0]["generated_text"]
print(output)

Training Prompt Format

Disagree (problematic) cases — model learns to diagnose and rewrite:

[MFT_PREFIX] <Prompt>: "{question}"; <Reply>: "{answer}". ###Diagnosis: {pragmatic_reasoning} Therefore, the <Revised Reply> is "{worker_answer}"

Agree (acceptable) cases — model learns to recognize and stop:

[MFT_PREFIX] <Prompt>: "{question}"; <Reply>: "{answer}". ###Diagnosis: {pragmatic_reasoning up to step (5)}

Citation

@article{chen2026learning,
  title={Learning to Diagnose and Correct Moral Errors: Towards Enhancing Moral Sensitivity in Large Language Models},
  author={Chen, Bocheng and Zi, Han and Chen, Xi and Zhang, Xitong and Johnson, Kristen and Liu, Guangliang},
  journal={arXiv preprint arXiv:2601.03079},
  year={2026}
}

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-01Update model card title to: Diagnose-and-Correct for Jailbreak · Llama 3.2-1B8c37d424 KB
    Loading...
  2. 2026-02-28Upload README.md with huggingface_hubbf739b74.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration