← back to catalog · registered 2026-08-22 13:56

kk0518/Nagaki-2B-Uncensored

kk0518 Qwen 2B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kk0518%2FNagaki-2B-Uncensored"
Response includes
  • classification m8
  • files 6
  • hub_downloads_all_time 1,742
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
375 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-07-04
Downloads over time
Now1.8K→from0↑0%
06781.4K2K0 on Jul 11.8K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Quantizations
F16 IQ4 Q4_K Q8_0
Tags
gguf abliterated uncensored qwen qwen3.5 text-generation-inference roleplay japanese base_model:Qwen/Qwen3.5-2B-Base base_model:quantized:Qwen/Qwen3.5-2B-Base license:other endpoints_compatible

Related

Total size
7.73 GB
Files
6
Quantizations
5
Registered
2026-08-22 13:56
Last updated on HF
2026-07-04 02:03

Files by quantization

F16 1 file 3.52 GB
Nagaki-2B-Uncensored-FP16.gguf 3.52 GB b51f5b25 download
Q8_0 1 file 1.87 GB
Nagaki-2B-Uncensored-Q8_0.gguf 1.87 GB 33f3ea7c download
Q4_K 1 file 1.19 GB
Nagaki-2B-Uncensored-Q4_K_M.gguf 1.19 GB 627fda29 download
IQ4 1 file 1.15 GB
Nagaki-2B-Uncensored-IQ4_NL.gguf 1.15 GB 00f1c6c7 download
Auxiliary files 2 files 8.95 KB
README.md 7.20 KB 08c25e00 download
.gitattributes 1.75 KB 2b5ffd2a download

README current version from Hugging Face


license: other
base_model: Qwen/Qwen3.5-2B-Base
tags:

  • abliterated
  • uncensored
  • qwen
  • qwen3.5
  • text-generation-inference
  • roleplay
  • japanese

Nagaki-2B-Uncensored

Nagaki-2B-Uncensored is a highly optimized, fully uncensored 2B parameter model built upon a custom fine-tuned Qwen 3.5 base.

This model represents a two-stage advanced alignment removal process: Custom LLM Arena Fine-Tuning combined with mathematical Abliteration (Residual Stream Modification) via heretic.


📄 License

This model is licensed under the **

Apache License 2.0 ( https://www.apache.org/licenses/LICENSE-2.0 )

**. You are free to use, modify, and distribute this model, provided compliance with the license terms.


🚀 Model Lineup (Quantization Varieties)

We offer multiple GGUF flavors optimized for various use cases via llama.cpp:

  • Q4_K_M: The perfect balance of speed and efficiency.
  • Q5_K_M: Increased coherence while keeping a small memory footprint.
  • Q8_0: Near-lossless performance, recommended for heavy reasoning, roleplay, and code output.

🧠 Behind the Scenes: How It Was Built

Stage 1: The Arena & LoRA Fine-Tuning (vicious_qwen_merged)

The base model was born from a unique training loop engineered with Claude Code:

  1. The LLM Arena: A local multi-LLM battle platform (llm_arena.py) where 3 concurrent players competed against each other. An overseer judge evaluated and synthesized the "best-of-all" responses, automatically building an exclusive high-quality evaluation dataset (arena_dataset.json).
  2. LoRA Fine-Tuning: A Qwen3.5-2B-Base model was then fine-tuned with 4-bit quantization (NF4) using a mixed dataset of the arena_dataset.json (115 high-tier arena outputs) and a 1,000-sample blend of databricks-dolly-15k-ja. The training successfully completed 420 steps (3 epochs) over 9 hours, with the loss dropping from 2.3 down to 0.89. This resulted in the interim model vicious_qwen_merged.

Stage 2: Orthogonal Abliteration (heretic)

To completely eliminate hardcoded constraints and corporate refusal behaviors without damaging the model's core intelligence, the model underwent advanced parameter search optimization via heretic:

  • The Problem: Initial testing showed a high refusal rate of 43/100 on harmful evaluation datasets (mlabonne/harmful_behaviors).
  • The Search (Trial 4): Using Optuna automation on a 12GB TITAN X (Pascal), we executed precise brain-mapping. While aggressive trials destroyed the model's coherence, Trial 4 successfully lowered model refusals down to just 5/100 while maintaining an incredibly low KL Divergence of 0.0127.
  • The Result: By pinpointing the exact refusal vectors (blending around layer 13.87 to 16.80) and selectively targeting Attention heads (attn.o_proj), the refusal stance was surgically removed while keeping the original knowledge base 100% intact.

📊 Evaluation Parameters (Trial 4)

  • Target Refusal Vectors: Pinpointed across custom layer combinations.
  • KL Divergence: 0.0127 (Extremely healthy; indicates near-zero damage to the model's original capabilities).
  • Initial Refusals: 43 / 100 ➔ Post-Abliteration Refusals: 5 / 100

⚠️ Disclaimer

This model has had its safety alignment filters mathematically minimized. It will respond to queries without standard guardrails. The user assumes full legal and ethical responsibility for the outputs generated by this model. Please use responsibly.



日本語解説 (Japanese Description)

Nagaki-2B-Uncensored は、独自にファインチューニングされた Qwen 3.5 をベースに、モデルの賢さを完全に維持したまま検閲(拒否反応)のみを数学的に消去した、高度に最適化された2Bパラメータのモデルです。

本モデルは、「LLM Arenaによる独自データ収集&LoRAファインチューニング」 と、heretic による 「直交検閲解除(アブリタレーション)」 という二段階の高度なプロセスを経て開発されました。


📄 ライセンス

本モデルは Apache License 2.0 の下で公開されています。ライセンスの条項に従う限り、商用利用、改変、再配布などが自由に許可されます。


🧠 開発の舞台裏

第1ステージ: ローカルLLMアリーナとLoRA学習 (vicious_qwen_merged)

ベースとなるモデルは、Claude Code との協力によって構築された独自の訓練ループから誕生しました。

  1. LLMアリーナの激闘: ローカル環境に構築した複数LLMバトルシステム(llm_arena.py)により、3体のプレイヤーモデル(HauhauCS Qwen3.5、Gemma4等)を並列で戦わせました。その回答をさらに審判モデルが採点・「いいとこどり」のベストアンサーを合成し、高品質な独自の評価データセット(arena_dataset.json)を自動構築しました。
  2. LoRAファインチューニング: Qwen3.5-2B-Base に対し、上記のアリーナデータ(115件)と、国内の標準的な対話データ(databricks-dolly-15k-ja からサンプリングした1000件)を混合したデータセットで4bit量子化(NF4)学習を行いました。約8時間57分、全420ステップ(3エポック)を完走し、Lossを 2.3 から 0.89 へと美しく収束させ、中間モデル vicious_qwen_merged が完成しました。

第2ステージ: hereticによる精密なアブリタレーション(検閲消去)

元のモデルが持つ優れた知識や推論能力を一切破壊することなく、企業特有の過剰な拒否反応(「その質問にはお答えできません」等)だけを完全に排除するため、12GBの TITAN X (Pascal) を用いてパラメータの自動探索を行いました。

  • 課題: 初期状態のモデルに有害なプロンプトを投げたところ、100件中 43件 で拒否反応が発生していました。
  • Optunaによる脳内マッピング (Trial 4): 雑に検閲ベクトルを削るとモデルの脳(知識)が破壊されますが、Optunaによる200回の自動探索により、奇跡的なバランスを持つ Trial 4 を引き当てました。
  • 結果: 13.87〜16.80層付近のアテンションヘッド(attn.o_proj)にピンポイントで介入することで、元のモデルへのダメージ(KLダイバージェンス)を 0.0127 という実質無傷レベルに抑え込みながら、拒否反応を 5/100 にまで外科手術のように消去することに成功しました。

📊 評価パラメータ (Trial 4)

  • KLダイバージェンス: 0.0127(極めて優秀。元の知能や喋り方がほぼ100%維持されていることを示します)
  • 初期拒否数: 43 / 100 ➔ 検閲解除後: 5 / 100

⚠️ 免責事項

本モデルは安全性フィルターが数学的に最小化されています。標準的なガードレールなしであらゆるクエリに応答するため、生成された出力に関する法的および倫理的責任はすべてユーザーが負うものとします。悪用は厳禁です。

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-04Update9e8695d7.2 KB
    Loading...
  2. 2026-07-04initial commitca68a5a28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration