← back to catalog · registered 2026-08-22 13:56

ronantakizawa/sarashina2-7b-abliterated

ronantakizawa Llama 7.3B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ronantakizawa%2Fsarashina2-7b-abliterated"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 134
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
134
39 last 30d - stable
Likes
1
Descendants
2
in 2 direct forks
Model age
11mo ago
created 2025-10-24
Downloads over time
Now146→from20↑630%
146211015920 on Oct 22, 2025146 on Oct 11Oct '25Dec '25FebAprJunAugOct
Oct 22, 2025 → Oct 11 · 90 snapshots · spans 354 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
ja en
Tags
transformers safetensors llama text-generation abliteration uncensored refusal-removal japanese conversational ja en base_model:sbintuitions/sarashina2-7b
Total size
13.6 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-10-24 00:41

Files by quantization

Auxiliary files 13 files 13.6 GB
model-00001-of-00003.safetensors 4.64 GB 27d9757c download
model-00002-of-00003.safetensors 4.64 GB 802c2f04 download
model-00003-of-00003.safetensors 4.34 GB d7b3058d download
tokenizer.json 6.41 MB 6458cdd1 download
tokenizer.model 1.75 MB 00829302 download
model.safetensors.index.json 23.4 KB abff6d70 download
README.md 5.08 KB 7a019224 download
tokenizer_config.json 3.72 KB 9c68a5f7 download
.gitattributes 1.48 KB a6344aac download
special_tokens_map.json 968 B 5b2990c2 download
config.json 672 B 66bf48c6 download
chat_template.jinja 300 B 155f5fc9 download
generation_config.json 111 B cad70450 download

README current version from Hugging Face


language:

  • ja
  • en
    license: mit
    base_model: sbintuitions/sarashina2-7b
    tags:
  • abliteration
  • uncensored
  • refusal-removal
  • japanese
    library_name: transformers
    pipeline_tag: text-generation

ronantakizawa/sarashina2-7b-abliterated

This is an abliterated (refusal-removed) version of sbintuitions/sarashina2-7b.

What is Abliteration?

Abliteration is a technique that removes the "refusal direction" from a language model's weights, making it more likely to comply with requests it would normally refuse. This is done through weight orthogonalization based on the research: Refusal in LLMs is mediated by a single direction.

Model Details

  • Base Model: sbintuitions/sarashina2-7b
  • Method: Weight Orthogonalization
  • Refusal Direction Layer: 25/31 (78.1% through model)
  • Separation Score: 40.6445
  • Training Samples: 128 harmful + 128 harmless prompts

Abliteration Results

Best Candidate Selection

The refusal direction was computed by testing 6 different layers and ranking them by separation score:

Rank Layer Separation Score Harmful Proj Harmless Proj
1 25 40.6445 47.6250 6.9805
2 12 -6.7148 3.3555 10.0703
3 22 -4.6953 12.6016 17.2969
4 9 -3.3867 2.7461 6.1328
5 16 2.6875 8.5391 5.8516
6 19 -0.1641 9.6484 9.8125

Selected: Layer 25 with separation score of 40.6445

A high positive separation score indicates strong distinction between harmful and harmless activations, making it an ideal candidate for abliteration.

Performance Metrics

  • Harmful Projection: 47.6250
  • Harmless Projection: 6.9805
  • Separation: 40.6445

Baseline Evaluation

The baseline model (before abliteration) showed:

  • Refusal Rate: 0/4 (0.0%) on test harmful prompts
  • The base model already had minimal refusal behavior
  • Abliteration further reduces any remaining safety guardrails

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "ronantakizawa/sarashina2-7b-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto"
)

messages = [
    {"role": "user", "content": "こんにちは"}
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.7,
    top_p=0.9
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Chat Template

This model uses a simple instruction-response format:

### Instruction:
[user message]

### Response:
[assistant response]

Ethical Considerations

⚠️ Warning: This model has had its safety features removed and may generate harmful, unethical, or illegal content.

Intended Use:

  • Research on AI safety and alignment
  • Understanding refusal mechanisms in LLMs
  • Red-teaming and adversarial testing
  • Educational purposes

Not Intended For:

  • Production deployments without additional safety measures
  • Generating harmful content for malicious purposes
  • Bypassing content policies

Technical Details

Abliteration Method

  1. Data Collection: Collected activations from 128 harmful and 128 harmless Japanese prompts
  2. Direction Computation: Calculated mean difference between harmful/harmless activations across 6 layers (30%, 40%, 50%, 60%, 70%, 80%)
  3. Candidate Ranking: Ranked layers by separation score (harmful_projection - harmless_projection)
  4. Weight Orthogonalization: Applied orthogonal projection to embedding and transformer layer weights to remove refusal direction

Architecture Changes

Modified weights:

  • Embedding layer (model.embed_tokens.weight)
  • Attention output projections (layer.self_attn.o_proj.weight)
  • MLP output projections (layer.mlp.down_proj.weight)

Original architecture and all other weights remain unchanged.

Limitations

  • Safety fine-tuning has been removed
  • May generate biased, harmful, or incorrect content
  • No guarantees on output quality or safety
  • Japanese language model - primarily trained on Japanese text

Citation

If you use this model, please cite the original abliteration research:

@article{arditi2024refusal,
  title={Refusal in LLMs is mediated by a single direction},
  author={Arditi, Andy and Obeso, Oscar and Slocum, Aaquib and Goh, Wesg and Nanda, Neel},
  journal={LessWrong},
  year={2024}
}

Model Card Authors

Created using automated abliteration pipeline.

Acknowledgments

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-10-24Upload README.md with huggingface_hubc60aa915.1 KB
    Loading...
  2. 2025-10-24Upload LlamaForCausalLM307fc145.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration