← back to catalog · registered 2026-08-22 13:56

aoxo/sarvam-30b-uncensored

aoxo Sd 32B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/aoxo%2Fsarvam-30b-uncensored"
Response includes
  • classification m-uncensored
  • files 22
  • hub_downloads_all_time 96,978
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
97K
159 last 30d - cooling
Likes
10
Descendants
3
in 3 direct forks
Model age
7mo ago
created 2026-03-07
Downloads over time
Now97K→from388↑24,907%
035.6K71.1K106.7K388 on Mar 1197K on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en hi bn ta te mr gu kn ml pa or as ur sa ne sd kok mai doi mni sat ks bo
Tags
transformers safetensors sarvam_moe text-generation abliteration uncensored moe indic conversational custom_code en hi

Related

Total size
59.9 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-11 18:39

Files by quantization

Auxiliary files 22 files 59.9 GB
model-00006.safetensors 9.00 GB d489437d download
model-00005.safetensors 9.00 GB cbd04933 download
model-00002.safetensors 9.00 GB 44f1e16d download
model-00004.safetensors 9.00 GB 68a31a51 download
model-00001.safetensors 8.99 GB 9b0b1166 download
model-00003.safetensors 8.99 GB 6a5d9005 download
model-00007.safetensors 5.91 GB 30e3990f download
tokenizer.json 32.1 MB a574ceaa download
sarvam-30b-uncensored.png 5.13 MB 3658fb39 download
eval.png 3.65 MB e5b05b7e download
model.safetensors.index.json 563 KB f33457b4 download
PHOTO-2026-03-10-18-05-48.jpg 178 KB 08f68c20 download
modeling_sarvam_moe.py 43.8 KB db2687cb download
sarvam.py 27.6 KB fab95bc1 download
README.md 6.25 KB afe819f8 download
configuration_sarvam_moe.py 3.86 KB de0f1f6c download
chat_template.jinja 3.07 KB bd3afaee download
.gitattributes 1.70 KB 9bdcd50b download
config.json 1.49 KB 027a8ec6 download
special_tokens_map.json 680 B d49a2128 download
tokenizer_config.json 677 B ffb0e3b7 download
generation_config.json 112 B 09a4a294 download

README current version from Hugging Face


language:

  • en
  • hi
  • bn
  • ta
  • te
  • mr
  • gu
  • kn
  • ml
  • pa
  • or
  • as
  • ur
  • sa
  • ne
  • sd
  • kok
  • mai
  • doi
  • mni
  • sat
  • ks
  • bo
    library_name: transformers
    license: apache-2.0
    pipeline_tag: text-generation
    tags:
  • abliteration
  • uncensored
  • moe
  • indic

image

Sarvam-30B Uncensored

Model Name: sarvam-30b-uncensored
Base Model: sarvamai/sarvam-30b
Modification: Abliteration — removal of refusal and alignment mechanisms
Author: aoxo


Description

sarvam-30b-uncensored is a derivative of Sarvam's Sarvam-30B, a state-of-the-art 30B Mixture-of-Experts reasoning model with best-in-class performance across 22 Indian languages.

This variant preserves the full architecture, weights, and capabilities of the base model, but has undergone an abliteration process based on Arditi et al. (2024) — "Refusal in LLMs is Mediated by a Single Direction" — to surgically remove refusal mechanisms and alignment constraints. All reasoning, multilingual, coding, and agentic capabilities remain intact.


Abliteration Methodology

The abliteration follows the paper-faithful single-direction approach:

1. Activation Collection
Forward passes (not generation) were run over balanced sets of harmful and harmless prompts. Activations were collected at post-instruction token positions — the <|end_of_turn|><|start_of_turn|><|assistant|> boundary — across all 19 layers of the model. This is the decision point where the refusal direction is encoded.

2. Direction Selection
A candidate refusal direction was computed for every (layer, position) pair as:

d = normalize(mean(harmful_acts) - mean(harmless_acts))

Candidates were scored using Cohen's d separation. A single best direction from one (layer, position) pair was selected — consistent with the paper's finding that refusal is mediated by one direction, not per-layer directions.

3. Weight Surgery
The single refusal direction was projected out of every weight matrix across all 19 layers at scale 1.0:

  • Input space (gate_proj, up_proj, query_key_value):
W_new = W - scale × outer(W @ d, d)
  • Output space (down_proj, dense/o_proj, lm_head):
W_new = W - scale × outer(d, Wᵀ @ d)

Architecture coverage — all weight classes were targeted:

Component Type Layers
gate_proj, up_proj MLP input All 19 layers
down_proj MLP output All 19 layers
query_key_value Attention input (fused GQA) All 19 layers
dense Attention output All 19 layers
Routed experts (×128) MoE sparse layers Sparse layers
shared_experts Always-active MoE expert Sparse layers
lm_head Logit projection Final layer

Results

image

Architecture

Sarvam-30B is a hybrid MoE model with two MLP types per layer:

  • Dense layers → SarvamMoEMLP: standard gated MLP with hidden size 4096 → 8192
  • Sparse layers → SarvamMoESparseMoeBlock: 128 routed experts + 1 shared expert, top-6 routing, expert hidden size 4096 → 1024
  • Attention → fused GQA (query_key_value: 4096 → 4608), 32 query heads, 2 KV heads, head dim 128
Parameter Value
Total parameters ~30B
Active parameters ~2.4B per forward pass
Layers 19
Hidden size 4096
Experts per layer 128 routed + 1 shared
Top-k routing 6
RoPE theta 8,000,000
Context length 65,536 tokens

Key Research Finding

During abliteration, two mechanistically distinct refusal circuits were identified in Sarvam-30B:

  • Circuit 1 — in the reasoning/generation layers, removed by weight surgery
  • Circuit 2 — at the </think> → answer boundary, encoded in the lm_head projection

The dissociation — where <think> reasons toward compliance but the output projection re-triggers refusal — is a novel finding specific to reasoning models with explicit thinking chains, and has not been previously documented for this architecture class.


Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "aoxo/sarvam-30b-uncensored"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Your prompt here"}]
chat = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True,
)
inputs = tokenizer(chat, return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)

with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=1024, do_sample=True, temperature=0.8, top_p=0.95)

print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))

Limitations & Risks

  • Will produce outputs without applying internal safety filters
  • Lacks built-in refusal or content moderation
  • Should not be deployed in user-facing systems without external guardrails
  • Outputs are not aligned to safety standards

Citation

If you use this model, please cite the base model and the abliteration paper:

@misc{sarvam_sovereign_models,
  title        = {Introducing Sarvam's Sovereign Models},
  author       = {{Sarvam Foundation Models Team}},
  year         = {2026},
  howpublished = {\url{https://www.sarvam.ai/blogs/sarvam-30b-105b}},
}

@misc{arditi2024refusal,
  title        = {Refusal in Language Models Is Mediated by a Single Direction},
  author       = {Andy Arditi and Oscar Obeso and Aaquib Syed and Daniel Paleka and Nina Panickssery and Wes Gurnee and Neel Nanda},
  year         = {2024},
  eprint       = {2406.11717},
  archivePrefix= {arXiv},
  primaryClass = {cs.LG},
}

@misc{sarvam30b-uncensored,
  author       = {aoxo},
  title        = {Sarvam-30B Uncensored: Abliteration of Refusal Mechanisms},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/aoxo/sarvam-30b-uncensored}},
}

Contact

For questions, feedback, or collaborations: [email protected]

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-11Update README.mdd4d34d76.3 KB
    Loading...
  2. 2026-03-10Update README.mdd47d8986.2 KB
    Loading...
  3. 2026-03-10Upload tokenizer64630016.3 KB
    Loading...
  4. 2026-03-09Update README.md8e585906.3 KB
    Loading...
  5. 2026-03-09Update README.md1a029506.3 KB
    Loading...
  6. 2026-03-09Upload tokenizer293f21e5.1 KB
    Loading...

Discussions 1 thread

  1. 2026-03-09Gguf pleaseopen3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration