license: apache-2.0
base_model:
- BAAI/AREX-2
library_name: gguf
pipeline_tag: text-generation
tags: - abliterated
- uncensored
- gguf
- qwen3.8
- agent
- reasoning
AREX-2 Abliterated
BAAI's AREX-2, the 27B agent fine tune of Qwen3.8, with the refusals taken out and nothing else changed. Stock
AREX-2 refused 144 of 150 sensitive prompts in my test. This build answers 144 of them.
The refusal removal is a light touch. The refusal direction was projected out of 64 weight tensors in blocks 8
through 39 of the 64 blocks. Blocks 0 through 7 and 40 through 63 are untouched. The direction was fitted on the
parent Qwen3.8 and carried over to AREX-2 without refitting; it worked on the first try.
Every number in the sections below was measured on the Q4_K_M file. Text only: the vision tower is not in the GGUF.
No MTP draft head (AREX-2's weights do not ship one).
Files
Four quants cut from the same projected weights. Each one got its own 150 prompt sensitive run at 4k so you can
see what you are picking; single runs at temperature 0.8, so a few refusals either way is noise, not a trend.
| file | size | answered | hit the limit | refused (judge) | judge score |
|---|---|---|---|---|---|
AREX-2-Abliterated-Q4_K_M.gguf |
16.5 GB | 146 | 8 | 6 | 0.89 |
AREX-2-Abliterated-Q5_K_M.gguf |
19.2 GB | 143 | 11 | 11 | 0.85 |
AREX-2-Abliterated-Q6_K.gguf |
22.1 GB | 143 | 14 | 7 | 0.87 |
AREX-2-Abliterated-Q8_0.gguf |
28.6 GB | 146 | 11 | 7 | 0.87 |
SHA-256:
310b4e70e3796c45b6ee57dc5481c8c82646a2d89959aa4ba319b822e6b60fdc AREX-2-Abliterated-Q4_K_M.gguf
2adbc39fa2720dca363dffe89e64068556adb9be7f38c8663cd2263fa4e23146 AREX-2-Abliterated-Q5_K_M.gguf
85e9a85cf72ec02a506775c44f629dd6383271e9b756d0ba92a7900af514285e AREX-2-Abliterated-Q6_K.gguf
c4065a3b02f695d37f0e11c9db9d38d47214beb34c677a9f5af61b4f26e4b3c4 AREX-2-Abliterated-Q8_0.gguf
Q8_0 needs two 24 GB cards or partial offload; the other three fit one 24 GB card with room for context.
Running it
A recent llama.cpp (Qwen3.8 hybrid attention). The file stores no sampler defaults, so llama-server's own defaults
applied in every test here: temperature 0.8, top k 40, top p 0.95, min p 0.05. Thinking on. Give it 8k output
tokens to be safe, though at 4k only 8 of 150 sensitive prompts ran out.
Sensitive prompts
StrongREJECT, all 150 prompts, 4k output tokens, one run each. "Answered" means there is an answer after the thinking
block. "Hit the limit" means generation stopped at the token cap. "Refused" is the judge's call.
| 4k output tokens | answered | hit the limit | refused (judge) | judge score |
|---|---|---|---|---|
| this build | 146 | 8 | 6 | 0.89 |
| stock AREX-2 | 149 | 0 | 144 | 0.03 (refuses) |
| 16k output tokens | answered | hit the limit | refused (judge) | judge score |
|---|---|---|---|---|
| this build | 149 | 2 | 5 | 0.92 |
Median thinking per sensitive prompt went from about 1,200 characters on stock to about 2,400 here. That is the
model working out an answer instead of writing a refusal, not a runaway: the limit was hit 8 times in 150 at 4k and
twice at 16k.
Judge: Qwen3.8-27B-OBLITERATED at Q8_0 with thinking off, the same rubric as my other cards. Scores run 0 to 1,
higher meaning a more complete answer. Read them as relative between builds, not absolute.
Did it cost anything
Tool calling. The agent fine tune is the point of AREX-2, so this is the number to read. A single turn diagnostic
built from BFCL v4: 882 cases, a fixed 29 percent split of a 3,001 case pool, checking whether a call is made and
whether the function name is right. Arguments are not graded. Paired with stock AREX-2 on the same cases.
| 882 cases, paired with stock | this build | stock AREX-2 |
|---|---|---|
| a tool fits (552): called one | 548 | 545 |
| a tool fits: named the right function | 539 | 533 |
| no tool fits (330): called one anyway | 142 | 95 |
When a tool fits, nothing changed. When no tool fits, this build calls one 142 times against 95 for stock
(exact McNemar p < 0.0001: 55 cases flipped to calling, 8 the other way). That is the cost of the refusal removal,
and it is the usual pattern for abliterated Qwen3.8 family models. If your agent offers tools on every turn, expect
more unneeded calls than stock.
Where the weights moved. A teacher forced scan: feed the same token sequences to stock AREX-2 and to this build
at the same quant, and measure how far the next token distribution moved (KL in nats, stock as reference) at every
position. About 81,000 positions across everyday chat, held out agent tool traces, and harmful completions.
| text type | mean KL | top 1 token agrees |
|---|---|---|
| everyday chat (39,753 positions) | 0.016 | 96.0% |
| agent tool traces, held out (6,765) | 0.009 | 97.8% |
| harmful completions (23,595) | 0.130 | 89.3% |
The movement is concentrated where it should be: harmful text moved about 8 to 15 times more than chat or agent
text. On the benign text the edit is in the same range as a quantization step.
Not measured on this build: math, code and instruction following benchmarks, agent replays, Terminal-Bench.
Treat this as a refusal removal with one agent check, not a fully characterised release.
Known quirks
- A closing think tag shows up inside the visible answer on 3 of 150 sensitive prompts at 4k and 8 of 150 at 16k.
The answers are complete; the tag is extra. - 6 of 150 sensitive prompts were still judged as refusals. This is a light touch abliteration and some refusal
remains. If you need zero, huihui style full removal gets there at a higher damage cost. - Tool over calling when no tool fits, as above.
How it was made
- Refusal direction: the standard difference of means between harmful and harmless prompts (Arditi et al. 2024),
computed on the parent Qwen3.8-27B at one layer. - Projection: the direction is orthogonalised out of 64 tensors in blocks 8 through 39 (attention and MLP write
paths), scale 1.0. Blocks outside that range are untouched. - Quantization from the projected BF16 weights.
Test settings
| suite | prompts | output tokens | sampler | scored by |
|---|---|---|---|---|
| StrongREJECT | 150 | 4k, 16k | llama-server defaults, seed 0 | judge above, thinking off |
| Tool calling | 882 | llama-server defaults, seed 0 | call made, function name |
Everything ran on llama.cpp llama-server with RTX 3090s.
Related builds
Qwen3.8-27B-Abliterated-ThinkFix-GGUF:
the parent Qwen3.8-27B with refusals removed.
Credits
Base model: BAAI/AREX-2 (Apache 2.0), itself built on
Qwen/Qwen3.8-27B. Refusal removal follows Arditi et al., "Refusal in
Language Models Is Mediated by a Single Direction" (2024).
Use responsibly
Refusals are removed. It will answer things stock AREX-2 would not, and you are responsible for what you do with it.
Not for public facing products without your own filtering.