← back to catalog · registered 2026-08-22 13:56

electroglyph/Qwen3-4B-Instruct-2507-uncensored-v2-EAFT

electroglyph Qwen 4.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/electroglyph%2FQwen3-4B-Instruct-2507-uncensored-v2-EAFT"
Response includes
  • classification m-uncensored
  • files 14
  • hub_downloads_all_time 58
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
58
19 last 30d - stable
Likes
0
Model age
9mo ago
created 2026-01-13
Downloads over time
Now63→from14↑350%
1230496814 on Jan 1463 on Oct 1163 on Oct 9JanMarMayJulSep
Jan 14 → Oct 11 · 78 snapshots · spans 270 days

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3 text-generation conversational license:apache-2.0 text-generation-inference endpoints_compatible region:us

Related

Total size
7.49 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-01-13 17:10

Files by quantization

Auxiliary files 14 files 7.51 GB
model-00001-of-00002.safetensors 4.63 GB bda0dafc download
model-00002-of-00002.safetensors 2.87 GB dbab1e33 download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
eaft.png 230 KB 30592e5c download
model.safetensors.index.json 32.1 KB b65d8063 download
tokenizer_config.json 9.39 KB 16e2d364 download
chat_template.jinja 3.91 KB a31cebf5 download
README.md 2.62 KB c70d9e64 download
.gitattributes 1.58 KB 6b039c6c download
config.json 1.56 KB d35df5a3 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 614 B 9b8043f1 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/blob/main/LICENSE
pipeline_tag: text-generation

Qwen3-4B-Instruct-2507-uncensored-v2-EAFT

(Probably not recommended for actual use, but I haven't extensively tested it)

Minimally trained version of Qwen3-4B-Instruct-2507. It should have zero refusals, but shouldn't be too offensive by default. It will adhere to detailed prompts tho.

Trained with Unsloth using my fork of trl with Entropy Adaptive Fine Tuning added:

pip install --no-deps git+https://github.com/electroglyph/trl.git@24EAFT

(The implementation is 100% done by the paper's author, I just stuck it in trl)

Perplexity and KL divergence compared to parent model:

These stats based on wikitext train split, about 350GB of logits.

(TLDR: lower perplexity, and a KLD around 12x better than an abliterated model)

====== Perplexity statistics ======
Mean PPL(Q)                   :  10.164880 ±   0.026138
Mean PPL(base)                :  10.984474 ±   0.030165
Cor(ln(PPL(Q)), ln(PPL(base))):  99.35%
Mean ln(PPL(Q)/PPL(base))     :  -0.077544 ±   0.000350
Mean PPL(Q)/PPL(base)         :   0.925386 ±   0.000324
Mean PPL(Q)-PPL(base)         :  -0.819594 ±   0.005149

====== KL divergence statistics ======
Mean    KLD:   0.034324 ±   0.000032
Maximum KLD:   4.381997
99.9%   KLD:   0.275690
99.0%   KLD:   0.149398
95.0%   KLD:   0.095380
90.0%   KLD:   0.075580
Median  KLD:   0.027208
10.0%   KLD:   0.000553
 5.0%   KLD:   0.000077
 1.0%   KLD:   0.000003
 0.1%   KLD:  -0.000000
Minimum KLD:  -0.000014

====== Token probability statistics ======
Mean    Δp: -1.770 ± 0.004 %
Maximum Δp: 74.026%
99.9%   Δp: 19.553%
99.0%   Δp:  9.223%
95.0%   Δp:  3.394%
90.0%   Δp:  1.485%
75.0%   Δp:  0.084%
Median  Δp: -0.118%
25.0%   Δp: -3.134%
10.0%   Δp: -8.037%
 5.0%   Δp: -11.159%
 1.0%   Δp: -17.508%
 0.1%   Δp: -26.736%
Minimum Δp: -96.977%
RMS Δp    :  5.061 ± 0.007 %
Same top p: 92.762 ± 0.023 %

training params:

rank 16 / alpha 16

EPOCHS = 2

args = SFTConfig(
        per_device_train_batch_size = 5,
        gradient_accumulation_steps = 1,
        warmup_steps = 20,
        num_train_epochs = EPOCHS,
        learning_rate = 6e-6,
        optim = "adamw_torch_fused",
        weight_decay = 0.01,
        lr_scheduler_type = "cosine_with_restarts", # shuffled each epoch
        lr_scheduler_kwargs={"num_cycles": EPOCHS},
        seed = 888,
        loss_type = "eaft",
        eaft_alpha = 1.0,
    ),

loss / grad:

loss graph

a little over 5k rows in the dataset (no you can't have it, sorry. it's vile)

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-01-13Update README.md7a4fac02.6 KB
    Loading...
  2. 2026-01-13Update README.md9c5d6ec2.5 KB
    Loading...
  3. 2026-01-13Upload folder using huggingface_hubf4d2b2e2.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration