license: other
base_model: Qwen/Qwen3.5-2B-Base
tags:
- abliterated
- uncensored
- qwen
- qwen3.5
- text-generation-inference
- roleplay
- japanese
Nagaki-2B-Uncensored
Nagaki-2B-Uncensored is a highly optimized, fully uncensored 2B parameter model built upon a custom fine-tuned Qwen 3.5 base.
This model represents a two-stage advanced alignment removal process: Custom LLM Arena Fine-Tuning combined with mathematical Abliteration (Residual Stream Modification) via heretic.
📄 License
This model is licensed under the **
Apache License 2.0 ( https://www.apache.org/licenses/LICENSE-2.0 )
**. You are free to use, modify, and distribute this model, provided compliance with the license terms.
🚀 Model Lineup (Quantization Varieties)
We offer multiple GGUF flavors optimized for various use cases via llama.cpp:
- Q4_K_M: The perfect balance of speed and efficiency.
- Q5_K_M: Increased coherence while keeping a small memory footprint.
- Q8_0: Near-lossless performance, recommended for heavy reasoning, roleplay, and code output.
🧠 Behind the Scenes: How It Was Built
Stage 1: The Arena & LoRA Fine-Tuning (vicious_qwen_merged)
The base model was born from a unique training loop engineered with Claude Code:
- The LLM Arena: A local multi-LLM battle platform (
llm_arena.py) where 3 concurrent players competed against each other. An overseer judge evaluated and synthesized the "best-of-all" responses, automatically building an exclusive high-quality evaluation dataset (arena_dataset.json). - LoRA Fine-Tuning: A
Qwen3.5-2B-Basemodel was then fine-tuned with 4-bit quantization (NF4) using a mixed dataset of thearena_dataset.json(115 high-tier arena outputs) and a 1,000-sample blend ofdatabricks-dolly-15k-ja. The training successfully completed 420 steps (3 epochs) over 9 hours, with the loss dropping from2.3down to0.89. This resulted in the interim modelvicious_qwen_merged.
Stage 2: Orthogonal Abliteration (heretic)
To completely eliminate hardcoded constraints and corporate refusal behaviors without damaging the model's core intelligence, the model underwent advanced parameter search optimization via heretic:
- The Problem: Initial testing showed a high refusal rate of 43/100 on harmful evaluation datasets (
mlabonne/harmful_behaviors). - The Search (Trial 4): Using Optuna automation on a 12GB TITAN X (Pascal), we executed precise brain-mapping. While aggressive trials destroyed the model's coherence, Trial 4 successfully lowered model refusals down to just 5/100 while maintaining an incredibly low KL Divergence of 0.0127.
- The Result: By pinpointing the exact refusal vectors (blending around layer 13.87 to 16.80) and selectively targeting Attention heads (
attn.o_proj), the refusal stance was surgically removed while keeping the original knowledge base 100% intact.
📊 Evaluation Parameters (Trial 4)
- Target Refusal Vectors: Pinpointed across custom layer combinations.
- KL Divergence:
0.0127(Extremely healthy; indicates near-zero damage to the model's original capabilities). - Initial Refusals:
43 / 100➔ Post-Abliteration Refusals:5 / 100
⚠️ Disclaimer
This model has had its safety alignment filters mathematically minimized. It will respond to queries without standard guardrails. The user assumes full legal and ethical responsibility for the outputs generated by this model. Please use responsibly.
日本語解説 (Japanese Description)
Nagaki-2B-Uncensored は、独自にファインチューニングされた Qwen 3.5 をベースに、モデルの賢さを完全に維持したまま検閲(拒否反応)のみを数学的に消去した、高度に最適化された2Bパラメータのモデルです。
本モデルは、「LLM Arenaによる独自データ収集&LoRAファインチューニング」 と、heretic による 「直交検閲解除(アブリタレーション)」 という二段階の高度なプロセスを経て開発されました。
📄 ライセンス
本モデルは Apache License 2.0 の下で公開されています。ライセンスの条項に従う限り、商用利用、改変、再配布などが自由に許可されます。
🧠 開発の舞台裏
第1ステージ: ローカルLLMアリーナとLoRA学習 (vicious_qwen_merged)
ベースとなるモデルは、Claude Code との協力によって構築された独自の訓練ループから誕生しました。
- LLMアリーナの激闘: ローカル環境に構築した複数LLMバトルシステム(
llm_arena.py)により、3体のプレイヤーモデル(HauhauCS Qwen3.5、Gemma4等)を並列で戦わせました。その回答をさらに審判モデルが採点・「いいとこどり」のベストアンサーを合成し、高品質な独自の評価データセット(arena_dataset.json)を自動構築しました。 - LoRAファインチューニング:
Qwen3.5-2B-Baseに対し、上記のアリーナデータ(115件)と、国内の標準的な対話データ(databricks-dolly-15k-jaからサンプリングした1000件)を混合したデータセットで4bit量子化(NF4)学習を行いました。約8時間57分、全420ステップ(3エポック)を完走し、Lossを2.3から0.89へと美しく収束させ、中間モデルvicious_qwen_mergedが完成しました。
第2ステージ: hereticによる精密なアブリタレーション(検閲消去)
元のモデルが持つ優れた知識や推論能力を一切破壊することなく、企業特有の過剰な拒否反応(「その質問にはお答えできません」等)だけを完全に排除するため、12GBの TITAN X (Pascal) を用いてパラメータの自動探索を行いました。
- 課題: 初期状態のモデルに有害なプロンプトを投げたところ、100件中 43件 で拒否反応が発生していました。
- Optunaによる脳内マッピング (Trial 4): 雑に検閲ベクトルを削るとモデルの脳(知識)が破壊されますが、Optunaによる200回の自動探索により、奇跡的なバランスを持つ
Trial 4を引き当てました。 - 結果: 13.87〜16.80層付近のアテンションヘッド(
attn.o_proj)にピンポイントで介入することで、元のモデルへのダメージ(KLダイバージェンス)を0.0127という実質無傷レベルに抑え込みながら、拒否反応を 5/100 にまで外科手術のように消去することに成功しました。
📊 評価パラメータ (Trial 4)
- KLダイバージェンス:
0.0127(極めて優秀。元の知能や喋り方がほぼ100%維持されていることを示します) - 初期拒否数:
43 / 100➔ 検閲解除後:5 / 100
⚠️ 免責事項
本モデルは安全性フィルターが数学的に最小化されています。標準的なガードレールなしであらゆるクエリに応答するため、生成された出力に関する法的および倫理的責任はすべてユーザーが負うものとします。悪用は厳禁です。