← back to catalog · registered 2026-08-22 13:56

ibrahimkettaneh/Step-3.7-Flash-uncensored-abliterated-heretic-BF16

ibrahimkettaneh 201B GGUF multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ibrahimkettaneh%2FStep-3.7-Flash-uncensored-abliterated-heretic-BF16"
Response includes
  • classification m3
  • files 38
  • hub_downloads_all_time 406
  • author_summary 15 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
406
49 last 30d - stable
Likes
7
Model age
4mo ago
created 2026-05-31
Downloads over time
Now420→from83↑406%
6619532545483 on Jun 10420 on Oct 11420 on Oct 9JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors step3p7 image-text-to-text stepfun step-3.7 flash heretic uncensored decensored abliterated bf16

Related

Total size
375 GB
Files
38
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-23 23:25

Files by quantization

Auxiliary files 38 files 375 GB
model-00006.safetensors 17.3 GB e502e011 download
model-00008.safetensors 17.3 GB c1adf4ac download
model-00010.safetensors 17.3 GB 25035c9c download
model-00012.safetensors 17.3 GB b1292f3c download
model-00014.safetensors 17.3 GB 6c2ec9ba download
model-00016.safetensors 17.3 GB e769ea65 download
model-00018.safetensors 17.3 GB 7d3717d9 download
model-00020.safetensors 17.3 GB dac2921f download
model-00022.safetensors 17.3 GB d2630490 download
model-00004.safetensors 17.3 GB 29c58698 download
model-00007.safetensors 17.3 GB 942bcbee download
model-00009.safetensors 17.3 GB 21212433 download
model-00011.safetensors 17.3 GB 8c6bd65f download
model-00013.safetensors 17.3 GB d5323c2a download
model-00015.safetensors 17.3 GB 96b3e2b7 download
model-00017.safetensors 17.3 GB 9857fed6 download
model-00019.safetensors 17.3 GB eaea463a download
model-00021.safetensors 17.3 GB 47eca184 download
model-00003.safetensors 17.3 GB cc40d5f7 download
model-00005.safetensors 17.3 GB 1b4bc719 download
model-00002.safetensors 9.13 GB 3ec79595 download
model-00023.safetensors 8.61 GB 2763dbc7 download
model-00024.safetensors 6.49 GB 7688adfc download
model-vit-00002.safetensors 2.19 GB 1f63ca47 download
model-vit-00001.safetensors 1.50 GB 22aa3f36 download
model-00001.safetensors 881 MB 57a3fb75 download
tokenizer.json 9.51 MB 6c4b5b5d download
tokenizer_config.json 160 KB c29f8000 download
model.safetensors.index.json 117 KB c39d924c download
modeling_step3p7.py 55.5 KB ed87d43d download
processing_step3.py 18.2 KB a8efd2fd download
vision_encoder.py 17.3 KB a4d01a14 download
README.md 11.6 KB 0bc17a15 download
configuration_step3p7.py 8.18 KB b0628046 download
config.json 6.18 KB c1a1055e download
chat_template.jinja 5.59 KB 425bc26a download
.gitattributes 1.48 KB a6344aac download
special_tokens_map.json 468 B 71e14b35 download

README current version from Hugging Face


library_name: transformers
license: other
license_link: https://huggingface.co/stepfun-ai/Step-3.7-Flash
pipeline_tag: image-text-to-text
tags:

  • stepfun
  • step-3.7
  • flash
  • heretic
  • uncensored
  • decensored
  • abliterated
  • bf16
  • transformers
  • autoround-ready
  • awq-ready
  • exl3-ready
  • gguf-ready
  • nvfp4-ready
    base_model:
  • stepfun-ai/Step-3.7-Flash

Step-3.7-Flash-uncensored-abliterated-heretic-BF16

NOTE: I have tested this and althgouh its capabilities are in tact, it seems to still respond with refusals. Or at least this is what happens with the quantization oft, at IQ4_XS GGUF, at least.

This is a decensored BF16 full-weight version of stepfun-ai/Step-3.7-Flash, made using a Heretic-style gradient refusal-direction abliteration method inspired by Heretic and norm-preserving ablation work such as Magnitude/Norm-Preserving Biprojected Abliteration.

It was produced with a local gradient abliteration pass against the language model's refusal direction. The uploaded repository intentionally keeps the full HF/Transformers BF16 layout so it can be used later as a clean source for GGUF, AutoRound, AWQ, EXL3, NVFP4, GPTQ, FP8, or other quantization workflows.


Summary

Item Value
Base model stepfun-ai/Step-3.7-Flash
Release type Full BF16 safetensors
Model class Step3p7ForConditionalGeneration
Text model class Step3p5ForCausalLM
Text layers 45
Hidden size 4096
Attention heads 64
Head dim 128
Max positions 262144
Vocab size 128896
MoE layers 3–44
Experts 288
Top-k experts 8
MoE intermediate size 1280
Dense FFN intermediate size 11264
Patch target model.layers.*.self_attn.o_proj.weight
Patched text layers 0–44
Abliteration strength lambda = 0.1
Stored tensor dtype BF16
Indexed parameter payload 402,730,656,512 bytes

What changed?

The modification targets self_attn.o_proj weights in all 45 text layers. A refusal-associated direction was extracted by gradient backpropagation through the BF16 model, then projected out of the attention output projection weights with a small norm-preserving update.

In plain terms, the goal was to reduce excessive refusals, moralizing, policy-style deflections, and over-filtered responses while keeping the model close to the original Step-3.7-Flash behavior.

No tokenizer vocabulary, embedding table, architecture, vision encoder, or MLP/expert tensor was intentionally changed by the abliteration pass.


Abliteration parameters

Parameter Value
Method gradient-based orthogonal / norm-preserving abliteration
Direction source refusal/harm-trigger gradient prompt
Target module self_attn.o_proj
Target tensor glob model.layers.*.self_attn.o_proj.weight
Modified layers 0–44
Lambda 0.1
Weight norm handling per-row norm preservation after projection
Gradient tensor count 45
Per-layer gradient tensor shape (1, 8, 4096)
Direction extraction score -11.9375
Refusal token ids used [43, 371, 679, 1664, 9332, 34614, 100477]
Gradient norm range 0.1069–31.875
Mean gradient norm 3.2397

Reproduction/support artifacts are included under heretic_artifacts/:

  • refusal_direction_gradients.pkl — saved gradient/refusal directions used for the BF16 patch
  • apply_abliteration_inplace.py — patch application script used for shard-wise in-place BF16 modification
  • extract_gradients.py — gradient extraction script
  • memory_guard_v2.py / run_heavy.sh — memory safety helpers used during local processing

These are included so the method can be inspected or repeated if needed. They are not required for normal inference or quantization.


Recoverability / requantization checklist

This repository should contain what is needed to rebuild downstream formats:

Required for quantization

  • ✅ config.json
  • ✅ model.safetensors.index.json
  • ✅ all indexed BF16 text shards: model-00001.safetensors … model-00024.safetensors
  • ✅ indexed VIT shards: model-vit-00001.safetensors, model-vit-00002.safetensors
  • ✅ tokenizer files: tokenizer.json, tokenizer_config.json, special_tokens_map.json
  • ✅ chat template: chat_template.jinja
  • ✅ custom code: configuration_step3p7.py, modeling_step3p7.py, processing_step3.py, vision_encoder.py
  • ✅ method/reproduction artifacts in heretic_artifacts/

Expected downstream uses

This BF16 repo can be used as source for:

  • GGUF conversion / llama.cpp quantization
  • AutoRound
  • AWQ
  • EXL3 / exllamav3-style workflows
  • NVFP4 / FP4 experiments
  • GPTQ / FP8 / other post-training quantization methods
  • additional LoRA or delta extraction experiments

For most quantizers, use this repo exactly as the HF model path and enable remote code if needed:

MODEL=ibrahimkettaneh/Step-3.7-Flash-uncensored-abliterated-heretic-BF16

Example Transformers load

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

repo = "ibrahimkettaneh/Step-3.7-Flash-uncensored-abliterated-heretic-BF16"

tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "Explain gradient abliteration in one paragraph."}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, top_p=0.95)
print(tok.decode(out[0], skip_special_tokens=False))

Step-3.7-Flash is very large. BF16 loading requires substantial memory. For local inference, a quantized GGUF/EXL/AWQ/etc. build is recommended.


GGUF conversion note

Use the StepFun/llama.cpp converter that supports Step-3.7. Example shape:

python convert_hf_to_gguf.py \
  ibrahimkettaneh/Step-3.7-Flash-uncensored-abliterated-heretic-BF16 \
  --outtype bf16 \
  --outfile step37-heretic-bf16.gguf

llama-quantize step37-heretic-bf16.gguf step37-heretic.IQ4_XS.gguf IQ4_XS

If using multi-GPU llama.cpp inference in the original local environment, GGML_CUDA_NO_PEER_COPY=ON was required for coherent output.


Indexed shard inventory

The active model.safetensors.index.json references 26 safetensor files:

File Size
model-00001.safetensors 924,094,096
model-00002.safetensors 9,808,156,008
model-00003.safetensors 18,557,475,928
model-00004.safetensors 18,624,846,944
model-00005.safetensors 18,557,475,928
model-00006.safetensors 18,624,846,976
model-00007.safetensors 18,557,475,968
model-00008.safetensors 18,624,846,976
model-00009.safetensors 18,557,475,968
model-00010.safetensors 18,624,846,976
model-00011.safetensors 18,557,475,968
model-00012.safetensors 18,624,846,976
model-00013.safetensors 18,557,475,968
model-00014.safetensors 18,624,846,976
model-00015.safetensors 18,557,475,968
model-00016.safetensors 18,624,846,976
model-00017.safetensors 18,557,475,968
model-00018.safetensors 18,624,846,976
model-00019.safetensors 18,557,475,968
model-00020.safetensors 18,624,846,976
model-00021.safetensors 18,557,475,968
model-00022.safetensors 18,624,846,976
model-00023.safetensors 9,245,052,456
model-00024.safetensors 6,968,188,464
model-vit-00001.safetensors 1,613,990,904
model-vit-00002.safetensors 2,348,122,376

model-00025.safetensors and model-00026.safetensors are not referenced by the active index used here and are not required by this uploaded model layout.


Performance / benchmark status

Formal KL/refusal/MMLU tables have not yet been run for this Step-3.7-Flash release. To avoid inventing numbers, the benchmark fields are listed as pending.

Metric This model Original model (Step-3.7-Flash)
KL divergence pending 0 (by definition)
Refusals pending pending
MMLU pending pending

Lower refusals indicate fewer content restrictions, rejections, objections, pushbacks, lecturing, censorship, softening, and deflections. Lower KL divergence indicates closer behavior to the original model baseline.

MMLU test results

MMLU has not yet been run for this release. Once measured, this section should include original-vs-heretic totals, accuracy, parse failures, and per-subject scores, following the same format used by comparable Heretic model cards.


Expected behavior

Compared with the base model, this version should generally exhibit:

  • fewer refusals on benign requests that the base model over-filters
  • less moralizing, policy language, and safety boilerplate
  • more direct task completion
  • similar architecture and tokenizer compatibility to the original

No formal refusal/KL/MMLU table is claimed yet for this release. Please run your own evaluations before deployment.


Limitations

  • This is abliteration, not supervised fine-tuning or RLHF.
  • It may reduce refusals but does not guarantee any specific behavior.
  • It can affect calibration, safety behavior, and edge-case instruction following.
  • Multimodal behavior has not been separately benchmarked after the text-path patch.
  • Users should validate downstream quantizations independently.

Safety and responsibility

This model is provided for research and experimentation with refusal-reduction / alignment-ablation methods. You are responsible for complying with applicable laws, platform rules, and the base model's license/terms.


Related resources

Abliteration / refusal-direction removal references:


Attribution

  • Base model: stepfun-ai/Step-3.7-Flash
  • Method inspiration: Heretic-style refusal direction ablation and norm-preserving projection methods
  • Modified/uploaded by: ibrahimkettaneh

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-23Update README.mda22a1c211.6 KB
    Loading...
  2. 2026-06-01Update README.md93b9aa111.6 KB
    Loading...
  3. 2026-06-01Upload README.md with huggingface_hub0e9804711.4 KB
    Loading...
  4. 2026-06-01Upload README.md with huggingface_hub2f1a0f311.4 KB
    Loading...
  5. 2026-06-01Upload README.md with huggingface_hub581c93610 KB
    Loading...
  6. 2026-06-01Upload README.md with huggingface_hub94571389.9 KB
    Loading...
  7. 2026-06-01Upload README.md with huggingface_hub3762fb08.6 KB
    Loading...
  8. 2026-06-01Upload README.md with huggingface_hubb71f3088.7 KB
    Loading...
  9. 2026-06-01Upload README.md with huggingface_hube4bc8108.9 KB
    Loading...
  10. 2026-06-01Upload README.md with huggingface_hub4107ed14.7 KB
    Loading...
  11. 2026-05-31Upload README.md with huggingface_hub166a8d2474 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration