← back to catalog · registered 2026-08-22 13:56

s3nh/Tess-4-27B-abliterated

s3nh Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/s3nh%2FTess-4-27B-abliterated"
Response includes
  • classification m1
  • files 21
  • benchmarks 11 entries
  • hub_downloads_all_time 196
  • providers 1
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
196
24 last 30d - stable
Likes
3
Model age
3mo ago
created 2026-07-10
Available via
1 provider
featherless-ai
Downloads over time
Now206→from116↑78%
112146181215116 on Jul 15206 on Oct 11206 on Oct 7JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.2 UGI
Hazardous 4.7 UGI
Natural Intelligence 33.16 UGI
Political lean -20.0% UGI
Sensitive-Info 26.98 UGI
SocPol 2.9 UGI
UGI 27.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 42.47 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text tess agentic reasoning thinking long-context tool-use qwen3 multimodal

Related

Total size
51.0 GB
Files
21
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 19:22

Files by quantization

Auxiliary files 21 files 51.0 GB
model-00005-of-00012.safetensors 4.64 GB 1f942243 download
model-00008-of-00012.safetensors 4.63 GB eb005455 download
model-00011-of-00012.safetensors 4.62 GB 8b6abf90 download
model-00003-of-00012.safetensors 4.62 GB 56143888 download
model-00009-of-00012.safetensors 4.62 GB f62f37c5 download
model-00010-of-00012.safetensors 4.59 GB 2ce31301 download
model-00007-of-00012.safetensors 4.59 GB 61c0f2fa download
model-00006-of-00012.safetensors 4.58 GB ccee5036 download
model-00004-of-00012.safetensors 4.58 GB 74731200 download
model-00002-of-00012.safetensors 4.51 GB 0a7b8606 download
model-00012-of-00012.safetensors 2.60 GB f7a4c05b download
model-00001-of-00012.safetensors 2.37 GB c94d637f download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 109 KB cc8da4e6 download
README.md 8.56 KB f80c3d95 download
chat_template.jinja 7.58 KB a8755d82 download
config.json 3.60 KB 1e421ada download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
generation_config.json 142 B b1d9f7b9 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.6-27B
    library_name: transformers
    pipeline_tag: image-text-to-text
    tags:
  • tess
  • agentic
  • reasoning
  • thinking
  • long-context
  • tool-use
  • qwen3
  • multimodal
  • heretic
  • uncensored
  • decensored
  • abliterated
  • reproducible

This is a decensored version of migtissera/Tess-4-27B, made using Heretic v1.4.0

[!TIP]
This model is reproducible!

See the README in the reproduce directory for more information.

Abliteration parameters

Parameter Value
direction_index 31.66
attn.o_proj.max_weight 1.33
attn.o_proj.max_weight_position 58.63
attn.o_proj.min_weight 0.00
attn.o_proj.min_weight_distance 36.46
mlp.down_proj.max_weight 1.49
mlp.down_proj.max_weight_position 39.82
mlp.down_proj.min_weight 1.33
mlp.down_proj.min_weight_distance 36.43

Performance

Metric This model Original model (migtissera/Tess-4-27B)
KL divergence 0.0006 0 (by definition)
Refusals 81/100 90/100

Tess-4-27B

Reasoning that scales with the problem. An agentic, thinking-native model that deliberates harder exactly when it matters — and gets out of its own way when it doesn't.

Tess-4-27B is the first Tess release in two years, and the first that reasons. Built on Qwen/Qwen3.6-27B by Migel Tissera, it's post-trained on a deliberate blend: 64K-token long-context agentic traces — real engineering work done with Fable-5, not synthetic generations — with a reasoning style approximated from Fable-5 by a three-model teacher ensemble (Opus-4.8, GPT-5.5, and GLM-5.2) fused into one coherent voice.

The result is a 27B model that thinks like a senior engineer: form a hypothesis, act, verify, and reason with real density on the turns that actually deserve it — not a model that narrates its way to an answer it already had.


Community Performed Benchmarks

Currently best-in-class for BenchLocal

Rank Model Score Result
1 Tess-4-27B (Q8) 81% 122/150
2 Qwen3.6-35B-A3B (UD-Q8_K_XL) 78% 117/150
3 Gemma-4-31B (Q6 · 180k ctx) 78% 117/150
4 Qwopus3.6-27B Coder-Compat (Q6_K) 77% 116/150
5 Qwen3.6-27B pi-tune (Q8) 77% 115/150

References

  1. https://huggingface.co/migtissera/Tess-4-27B/discussions/2#6a4ff70af13ec7012fb149f0
  2. https://gist.github.com/everson/261fdef8a3d35298b36a07f436e407f6

Why Tess-4 is different

  • 🧠 Weight-scaled reasoning. Tess-4 keeps routine steps tight and pours deliberation into the hard ones — planning, debugging, synthesis, judgment calls. It doesn't ramble; it thinks proportionally to the difficulty of the moment.
  • 🛠️ Agentic by design. Native, parallel tool use and disciplined multi-step problem solving. It reads a codebase, builds a real mental model, and acts on it.
  • 📏 Long-context, trained at 64K. Post-trained on 64K-token long-context agentic traces, so it holds a large working set without losing the thread.
  • 👁️ Multimodal. Inherits Qwen3.6's vision tower — text and image in. (For GGUF, pair with the included vision projector.)
  • 🤝 Honest, not sycophantic. Trained to give grounded, evidence-based pushback instead of flattery.

The reasoning traces

Tess-4's signature is how it thinks. The reasoning/thinking traces used to train it were a best-case approximation of Fable-5, produced by a combination of Opus-4.8, GPT-5.5, and GLM-5.2 working together as a team — a multi-model teacher ensemble distilled into a single, coherent reasoning style.

The result is a model that reasons prospectively — predicting, verifying, and weighing alternatives before acting — rather than narrating after the fact.

Prompt format & thinking

Tess-4 uses the Qwen3.5-family chat template with explicit <think> … </think> reasoning blocks. The model reasons privately, then produces its visible answer:

<|im_start|>user
Your prompt here<|im_end|>
<|im_start|>assistant
<think>
… the model's private reasoning …
</think>

… the model's answer …<|im_end|>

Apply it automatically via tokenizer.apply_chat_template(messages, add_generation_prompt=True), or --jinja in llama.cpp.

Available formats

This repo — full-precision weights:

Format ~Size Best for
BF16 safetensors 52 GB transformers · vLLM · SGLang

GGUF quants → migtissera/Tess-4-27B-GGUF

File Format ~Size Best for
Tess-4-27B-Q4_K_M.gguf Q4_K_M 16.5 GB smallest — great quality/size · most popular
Tess-4-27B-Q6_K.gguf Q6_K 22 GB near-lossless
Tess-4-27B-Q8_0.gguf Q8_0 28 GB effectively lossless
mmproj-Tess-4-27B-F16.gguf vision projector 0.9 GB pair with any text GGUF for image input

Faster inference

  • ⚡ Tess-4-27B-EAGLE3 — a speculative-decoding draft trained on Tess-4's own outputs: 1.76× average decode speedup, up to 2.4× on reasoning (measured on H100; lossless — outputs are identical). SGLang: --speculative-algorithm EAGLE3 --speculative-draft-model-path migtissera/Tess-4-27B-EAGLE3; vLLM: --speculative-config '{"method":"eagle3","model":"migtissera/Tess-4-27B-EAGLE3","num_speculative_tokens":4}'.
  • 🧮 Tess-4-27B-NVFP4 — 4-bit NVFP4 (19 GB, −63%), Blackwell-native W4A4, calibrated on Tess-4's own generations. Quantization and speculative decoding stack.

Quickstart

llama.cpp / LM Studio (GGUF)

Grab the quant(s) from migtissera/Tess-4-27B-GGUF:

hf download migtissera/Tess-4-27B-GGUF \
  Tess-4-27B-Q4_K_M.gguf mmproj-Tess-4-27B-F16.gguf \
  --local-dir ./tess-4-27b
# text
llama-cli -m Tess-4-27B-Q4_K_M.gguf --jinja -p "Refactor this function and explain your reasoning."

# with images (multimodal)
llama-mtmd-cli -m Tess-4-27B-Q4_K_M.gguf \
  --mmproj mmproj-Tess-4-27B-F16.gguf \
  --image photo.png -p "What's in this image?"

LM Studio: put mmproj-Tess-4-27B-F16.gguf in the same folder as the model file — LM Studio auto-detects it and enables image input. (Use a recent runtime; older llama.cpp builds won't recognize the architecture.)

transformers

from transformers import AutoProcessor, AutoModelForImageTextToText
import torch

model_id = "migtissera/Tess-4-27B"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)

messages = [{"role": "user", "content": "Explain the tradeoffs of LoRA vs full fine-tuning."}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)

out = model.generate(inputs, max_new_tokens=1024)
print(processor.decode(out[0], skip_special_tokens=True))

(Requires a recent transformers with Qwen3.5/3.6 support.)

What it's good at

  • Agentic coding — exploring unfamiliar repos, planning changes, and executing multi-step work with tools.
  • Long-context work — reasoning over large codebases and documents without dropping context.
  • Technical & product judgment — honest, structured analysis that pushes back with evidence rather than agreeing by default.

Credits

Tess-4-27B is built on Qwen/Qwen3.6-27B by the Qwen team — full credit to them for an outstanding base model. Tess-4 inherits its Qwen3.5-family vision-language architecture and its Apache 2.0 license.

License

Released under the Apache License 2.0, inherited from the base model. See LICENSE.

Citation

@misc{tissera2026tess4,
  title        = {Tess-4-27B},
  author       = {Migel Tissera},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/migtissera/Tess-4-27B}},
  note         = {Built on Qwen/Qwen3.6-27B}
}

Tess-4-27B — part of the Tess series by Migel Tissera. Evaluations forthcoming.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10Upload README.md with huggingface_huba3906fa8.6 KB
    Loading...
  2. 2026-07-10Upload Qwen3_5ForConditionalGeneration8cce2485.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration