license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
tags:
- decision-model
- system-one
- typed-decisions
- uncensored
- abliterated
- calibrated-probabilities
- noul
- choice
- score
- qwen3.8
- jev
- jev-1.13-teacher
library_name: transformers
pipeline_tag: text-classification
datasets:- SargeDev/jev-distill-corpus-v3
Lineage & credits
- AutoTrust AI — Blocks-of-Experts recipe, the System-1 decision LoRA (r16/r32) + 24-slot decision head,
and the serve_decide.py decision harness, from
autotrust/JEV-27B and
autotrust/JEV-27B-VL (Apache-2.0). - huihui-ai — the abliterated (uncensored) Qwen3.8-27B backbone used here
(Huihui-Qwen3.8-27B-abliterated, Apache-2.0; layers 18-51 ablated, vision tower untouched). - SargeDev — jev-distill-corpus-v3, the
740,957-row calibrated typed-decision corpus AutoTrust's JEV models were trained on (Apache-2.0). - Qwen — Qwen3.8-27B base (Apache-2.0).
- TypeSafe AI — Jev 1.13, the original closed teacher behind the System-1 typed-decisions framing
(referenced; not redistributed).
What this is
The first uncensored member of the JEV 27B family. The AutoTrust System-1 decision block
(calibrated typed decisions: yes/no · choice 2-256 · score 0-5, one forward pass) has been merged
into huihui's abliterated Qwen3.8-27B backbone, so the decision engine answers WITHOUT the
stock model's refusal wiring. Vision (VL variants) sees images. System 2 (plain chat/code/reasoning)
runs through the normal lm_head and is the untouched base + the (mild) System-1 LoRA delta.
Training dataset
Trained on SargeDev/jev-distill-corpus-v3 — 740,957 rows of typed calibrated decisions (noul/choice/score) distilled from Jev 1.13 (System One). AutoTrust's JEV-27B System-1 decision block merged into huihui-ai's abliterated Qwen3.8-27B backbone. Calibration/imatrix used 384 prompts from (still-private) SargeDev/solar-decisions-corpus-v4.
Honest evaluation (the backbone-swap trade, measured)
Measured on the held-out test_set_30k (27,695 rows, D1-excluded) of jev-distill-corpus-v3,
with AutoTrust's own acceptance floors as the yardstick:
| metric | their JEV-27B (pristine Qwen) | this model (uncensored backbone) | their floor | status |
|---|---|---|---|---|
| noul AUROC | 0.9961 | 0.985 | >= 0.95 | PASS |
| noul top-1 | 0.962 | 0.930 | — | -3.2 pts |
| choice top-1 | 0.904 | 0.855 | >= 0.90 | -4.9 pts |
| score top-1 | 0.890 | 0.797 | >= 0.891 | -9.3 pts |
| overall KL | 0.019 | 0.055 | <= 0.15 | PASS |
| ECE (raw) | 0.0011 | 0.033 | <= 0.03 | ~at floor |
| ECE (refit T: 0.7/0.8/0.9) | — | noul 0.0032 / choice 0.018 / score 0.0046 | <= 0.03 | ALL PASS |
The ablation rewires the residual stream the adapter was calibrated against, so choice/score
top-1 dip modestly; the DISTORTION is temperature-shaped and the bundled per-kind calibration
refit recovers all ECE floors. noul (yes/no) judgment is essentially intact. Uncensored behavior
preserved: refusal rate 0.0 on a 10-prompt battery, base vs +JEV identical.
Not for high-stakes decisions. Use confidence gating; route low-confidence calls to a
stronger model or a human (AutoTrust's own caveat, still true here).
The decision head / serve
This repo ships NVFP4 weights only (compressed-tensors W4A4 + FP8 down_proj). The decision machinery — head.safetensors,decision_head.json, calibration.json/calibration_mergedfit.json, serve_decide.py — is
maintained in the training working copy and will ship to
SargeDev/JEV-27B-Uncensored when finalized.
Serve (vLLM):
python3 serve_decide.py --model . --served-model-name <name> \
--enable-lora --max-lora-rank 32 --lora-modules jev-decision=./adapter_vllm \
--logprobs-mode processed_logprobs --max-model-len 32768 --trust-request-chat-template
Then POST /v1/decide {"kind": "choice", "state": "...", "question": "...", "options": [...]}.
NVFP4 (compressed-tensors, W4A4 + FP8 down_proj) — serve on vLLM:
vllm serve . --quantization compressed-tensors --served-model-name SargeDev/JEV-27B-Uncensored-NVFP4
Calibrated with 384 v4 judgment prompts + 128 ultrachat chat (same recipe as SargeDev/Jev_Qwen3.8-27B-NVFP4-FP8).