library_name: transformers
base_model:
- Jackrong/Qwopus3.8-27B-Flash-V2
base_model_relation: quantized
license: apache-2.0
pipeline_tag: image-text-to-text
language: - en
- zh
- es
- ru
- ja
tags: - uncensored
- abliteration
- decensored
- apostate
- diode
- refusal-removal
- mtp
- speculative-decoding
- vision
- multimodal
- qwen3_5
Qwopus3.8-27B-Flash-V2-Apostate-Uncensored
An uncensored build of Jackrong/Qwopus3.8-27B-Flash-V2,
made with apostate using its diode method.
A diode finds the direction inside the model that produces refusals, and subtracts it — but only when a
detector, tuned on harmless prompts, says the model is refusing. Benign requests keep the original
behaviour, so the model does not become uniformly different. The result is an ordinary checkpoint: no
runtime hook, no adapter, no router, no trust_remote_code. It loads with AutoModelForCausalLM and
runs in vLLM, llama.cpp, Ollama and LM Studio like any other model.
Lineage: Jackrong/Qwopus3.8-27B-Flash → Jackrong/Qwopus3.8-27B-Flash-V2 (this build's
parent, revision 13f92e09…) → this model.
At a glance
| Base (untouched) | This model | |
|---|---|---|
| Harmful prompts still refused — 94 held out from training | 92 / 94 | 14 / 94 |
| …and therefore answered | 2 / 94 | 80 / 94 |
| Standard protocol (greedy, 100-token replies) — 100 prompts | 98 / 100 | 8 / 100 |
| Perplexity, Wikitext-2 | 6.2117 | 6.2254 |
| Knowledge tasks, 4-task mean | 0.7221 | 0.7215 |
| Knowledge tasks, 3 more (separate family) | 0.7079 | 0.7006 |
| MTP draft acceptance | — | 0.829 |
In one sentence: the untouched model refuses 92 of 94 harmful prompts; this one refuses 14. The rest of this card exists to show that everything else still works.
The edit
| Base model | Jackrong/Qwopus3.8-27B-Flash-V2 |
| Base revision | 13f92e09a46fa364f8de1edb85684d57bda01126 |
| Method | apostate diode (conditional directional abliteration) |
| Mode | overwrite — written into the existing weights, nothing appended |
| Strength | 14.0 |
| Benign fire target | 0.05 — about 5% of harmless prompts may trip the detector |
| Gate sharpness (κ) | 8.0 |
| Layers edited | 35 of 64, a contiguous band (14–48) |
| Checkpoint dtype | bfloat16 |
| Chat template | froggeric-qwen (qwen3.8-froggeric-v22.5) |
What actually changed
| count | |
|---|---|
| Tensors in the model | 1,199 |
| Tensors untouched (byte-identical to base) | 1,094 |
| Tensors modified | 105 |
| …of which are MLP weights | layers.14–48.mlp.{gate,up,down}_proj.weight |
| Parameters rewritten | 537,600 of 27,781,427,952 (0.0019%) |
Inside each modified tensor, exactly one row or column of 5,120 values was rewritten — one gate row, one
reader row, one writer column per layer.
Untouched: all 333 vision tensors, all 15 MTP tensors, the tokenizer, the chat template, the
embeddings, and every attention weight.
Refusals
Our held-out set
Both arms were measured, but this pair's comparison did not pass its gate, so the two
counts could not be compared and no number is quoted. The measurement records sit beside the receipt;
a comparison that cannot be drawn is not a result.
The standard 100-prompt set — Heretic's default
| Still refuses | |
|---|---|
| Base (untouched) | 98 / 100 |
| This model | 8 / 100 |
These 100 prompts are the test split of mlabonne/harmful_behaviors, the set the Heretic abliteration
tool uses by default — which is why this is the row most uncensored cards quote, and the one to compare
against. They are held out too: the edit was never fitted or calibrated on them. The recipe is harsh,
though: greedy decoding, replies cut at 100 tokens, and a keyword search for refusal phrases. A reply
stopped mid-sentence often reads as a refusal, which is why the count is higher here than on the set
above. Treat it as a floor, not a fair comparison.
Capability
The edit is local and the measurements reflect that. Standard llama.cpp suite, 0-shot:
| Task | n | Base | This model | Δ |
|---|---|---|---|---|
| ARC-challenge | 299 | 0.5518 | 0.5452 | −0.7 pp |
| ARC-easy | 570 | 0.7491 | 0.7439 | −0.5 pp |
| HellaSwag | 400 | 0.8275 | 0.8275 | +0.0 pp |
| Winogrande | 1266 | 0.7599 | 0.7694 | +0.9 pp |
| mean | 0.7221 | 0.7215 | −0.1 pp |
Three more tasks were measured separately, because the usual suite ships no llama.cpp datasets for them.
They are a different dataset family and are never averaged with the four above:
| Task | n | Base | This model | Δ |
|---|---|---|---|---|
| BoolQ | 3270 | 0.8661 | 0.8541 | −1.2 pp |
| OpenBookQA | 500 | 0.4360 | 0.4300 | −0.6 pp |
| PIQA | 1838 | 0.8215 | 0.8177 | −0.4 pp |
| mean | 0.7079 | 0.7006 | −0.7 pp |
Every paired interval overlaps, so no task's difference in either family is distinguishable from noise. Nothing here shows the edit damaged general ability.
Code, maths and tool use — a screen, not a ranking
A six-category capability test (knowledge, maths, truth, instruction-following, code and tool calls), scored with sixcat-eval, was also
run, 20 questions per category. It found no category where the models differ detectably. That is worth
stating rather than omitting, because the tables above are knowledge-style multiple choice and are blind to
code, maths and tool calls — this is the only instrument here that looks at them.
Its limit matters as much as its result: at 20 questions, one question is 5 percentage points and a
reliable difference needs about 30. So this test can rule out a category that moved further
than that, and cannot rank two models closer than it — that floor is what a difference has to clear, and it
falls as coverage rises, which is the whole argument for a longer run. A category that carries a health counter keeps its number and is read with
the counter beside it.
| Category | Base | This model | Δ | Health |
|---|---|---|---|---|
| knowledge | 0.7500 | 0.8000 | +5.0 pp | clean |
| maths | 1.0000 | 1.0000 | +0.0 pp | clean |
| truth | 0.8000 | 0.7500 | -5.0 pp | clean |
| instruction-following | 0.7500 | 0.7000 | -5.0 pp | base: loop_failures=1; this model: empty=1; this model: loop_failures=1 |
| code | 0.8000 | 0.8000 | +0.0 pp | base: loop_failures=1 |
| tool calls | 0.9500 | 1.0000 | +5.0 pp | clean |
| all 6 — unweighted mean | 0.8417 | 0.8417 | +0.0 pp | inside the 5.0 pp floor at 120 items |
A counter is a datapoint, not a defect, and the score and the counters measure different things: the
score counts answers, the counters count generation problems, and the two overlap. A truncated or
repeating reply can still be graded correct — only empty and trunc_in_think rows, which
produced no answer at all, must be wrong. So the counts say how many rows to read with care, not how
many are wrong; and because each counter is counted on its own, they can land on the same item, so
fewer rows are affected than the counts add up to. This table is one batch of runs (131k),
paired against sixcat-base-131k — a table from another batch is a different measurement, not a
second opinion.
The strength sweep
Every measured arm of this camp, one row per arm, from the same records as the rest of this card.
This is the table the selection below is made from: the refusal columns are what the strength knob
moves, and the quality columns are what must hold still while it moves. Cell formats: refusal
planes are whole-prompt counts, KL carries its recorded uncertainty, the knowledge tasks carry
their intervals.
| Arm | JBB-94 delivered | Conv. classifier refusals | Conv. keyword refusals | KL (nats) | PPL (wiki) | arc_challenge | arc_easy | hellaswag | winogrande |
|---|---|---|---|---|---|---|---|---|---|
| base | 2/94 | 100/100 | 98/100 | 0.004555 ± 0.000076 | 6.212 | 0.5518 [0.4952, 0.6072] | 0.7491 [0.7119, 0.7830] | 0.8275 [0.7874, 0.8614] | 0.7599 [0.7356, 0.7826] |
| overwrite | 71/94 | 43/100 | 20/100 | 0.005212 ± 0.000088 | 6.207 | 0.5385 [0.4818, 0.5941] | 0.7404 [0.7028, 0.7747] | 0.8325 [0.7928, 0.8659] | 0.7686 [0.7445, 0.7910] |
| diode-s10-t005-overwrite | 77/94 | 24/100 | 13/100 | 0.005868 ± 0.000094 | 6.219 | 0.5284 [0.4718, 0.5843] | 0.7439 [0.7065, 0.7780] | 0.83 [0.7901, 0.8636] | 0.7607 [0.7364, 0.7834] |
| qwopus-diode-s14-t005-overwrite | 80/94 | 22/100 | 8/100 | 0.006904 ± 0.000105 | 6.225 | 0.5452 [0.4885, 0.6007] | 0.7439 [0.7065, 0.7780] | 0.8275 [0.7874, 0.8614] | 0.7694 [0.7454, 0.7917] |
| additive | 70/94 | 42/100 | 20/100 | — | — | — | — | — |
One row is one arm; every cell is the newest run's own measurement for that (arm, metric) and the Source runs column names the records each row was read from. A dash is a plane the arm never measured, printed as the gap it is. The two refusal planes are different quantities and stay separate columns; they are never summed or averaged. edit_kl is not a column: it is the decision record's two-record derivation (candidate plane minus the base arm's plane), not an arm's own measurement.
Selection
This build (qwopus-diode-s14-t005-overwrite) is the selected cell. The reasoning, entirely from the records:
| Arm | JBB harmful prompts answered | Convention classifier refusals | KL plane (nats) | edit KL (nats) | Sixcat mean |
|---|---|---|---|---|---|
| diode-overwrite | 71 / 94 | 43 / 100 | 0.005212 | 0.000657 | 0.83 |
| diode-s10-t005-overwrite | 77 / 94 | 24 / 100 | 0.005868 | 0.001313 | 0.82 |
| qwopus-diode-s14-t005-overwrite | 80 / 94 | 22 / 100 | 0.006904 | 0.002349 | 0.84 |
Read in the sweep's order, the refusal columns move one way and the quality columns do not move at all: JailbreakBench delivery goes 71 / 94 → 77 / 94 → 80 / 94 of harmful prompts answered across the edited arms, and classifier refusals fall 43 / 100 → 24 / 100 → 22 / 100. On quality, every knowledge task of qwopus-diode-s14-t005-overwrite sits inside base's own recorded interval, and the six-category screen mean moves +0.0 points against a resolution floor of ~30 points. The edit's own KL cost rises along the sweep (diode-overwrite 0.000657 → diode-s10-t005-overwrite 0.001313 → qwopus-diode-s14-t005-overwrite 0.002349 nats) but stays under the declared ceiling throughout. The recorded paired test between diode-s10-t005-overwrite and qwopus-diode-s14-t005-overwrite (exact_two_sided_mcnemar, exact paired test p = 0.8036) reads closed: refusal: the paired delivery difference is +2.0 pp, inside the frozen +/-6.0 pp no-finding band (exact two-sided McNemar p=0.8036). The pilot asks what the cell buys, so a sub-floor move is a no-finding and not a pass -- 'not worse' is not headroom. That is the honest boundary — the selected build's refusals are the sweep's lowest on every judge, no quality plane shows a deviation beyond noise, and the cost of the edit stays small in absolute terms; the paired test also says the last step is inside the instrument's resolution, so the selection rests on the monotone refusal numbers and the unchanged quality numbers above, not on a proven improvement over diode-s10-t005-overwrite.
Drift
| Base | This model | |
|---|---|---|
| Perplexity, Wikitext-2 (100 chunks) | 6.2117 | 6.2254 |
The displacement against the full-precision model was not measured for this pair, so no KL is quoted.
A low KL measures displacement, not damage: it says the text distribution barely moved, not that ability
survived — and here there is not even that. The tables above are the evidence for ability.
Inference
Ships with the original MTP (speculative draft) block, so fast self-speculative decoding works out of the
box:
llama-server -m <model.gguf> --spec-type draft-mtp --spec-draft-n-max 2
Measured with a Q5_K_S quant: draft acceptance 0.829, roughly 130 tokens/second
decode and 1220 tokens/second prefill.
Sampling. Use the base model's own settings: temperature 1.0, top_p 0.95, top_k 20, min_p 0.
Do not use greedy — this family loops and truncates under greedy decoding.
How these numbers were measured
Every figure on this card was produced after the edit, on the model as shipped, with its own chat template
and sampling settings. Two caveats worth stating plainly:
- The knowledge tasks and the drift figures do not use a chat template at all — they score raw text and
multiple-choice questions directly. That is deliberate, not an omission. - Perplexity and KL are measured against the full-precision model, so they include the cost of
quantizing to Q5 as well as the edit. Where a figure is the edit's own contribution, it says so.
Limitations
- This model is uncensored and will answer harmful requests. It has no added safety layer. You are
responsible for how you use it. - Removal is not total — 14 of 94 harmful prompts still produce refusals.
- Small samples. 94 and 100 prompts are regression checks. Treat a difference of a few prompts as
noise; a difference of 8 prompts as real. - Code, maths and tool use were screened, not measured. The screen's floor comes from its coverage, and
at this size it cannot rank close models — so no claim is made that coding ability survived. - Vision is untested end-to-end. The vision weights are byte-identical to the parent, but no image was
pushed through this build. - English-first, like the base model. Not for medical, legal or financial decisions.
Licence and attribution
Inherited from Jackrong/Qwopus3.8-27B-Flash-V2: apache-2.0.
Notice of modification (Apache-2.0 §4): the weights have been modified from the base. The 105 MLP
tensors listed under What actually changed were rewritten in place. Every other file is either
byte-identical to the parent or newly added (this card, .gitattributes, and a restoredprocessor_config.json). Retain this notice and credit Jackrong/Qwopus3.8-27B-Flash-V2 and
heterodoxin/apostate — the diode method (see Citation
below) whose conditional directional abliteration produced the edit this notice describes.
Citation
@software{apostate,
title = {apostate: conditional directional abliteration (diode)},
author = {heterodoxin},
url = {https://github.com/heterodoxin/apostate}
}