← back to catalog · registered 2026-08-22 13:56

tomvaillant/qwen3.6-27b-abliterated-journalist-GGUF

tomvaillant Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tomvaillant%2Fqwen3.6-27b-abliterated-journalist-GGUF"
Response includes
  • classification m8
  • files 8
  • hub_downloads_all_time 3,606
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
4K
200 last 30d - cooling
Likes
1
Model age
4mo ago
created 2026-05-24
Downloads over time
Now3.7K→from907↑306%
7681.8K2.9K4K907 on Jun 103.7K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 229 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K
Tags
gguf qwen3_5 llama.cpp investigative-journalism osint conversational en base_model:huihui-ai/Huihui-Qwen3.6-27B-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.6-27B-abliterated license:apache-2.0 endpoints_compatible region:us

Related

Total size
15.4 GB
Files
8
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-07-08 14:07

Files by quantization

Q4_K 1 file 15.4 GB
qwen3.6-27b-abliterated-journalist-Q4_K_M.gguf 15.4 GB 039c653c download
Auxiliary files 7 files 19.1 MB
tokenizer.json 19.1 MB 87a7830d download
tokenizer_config.json 14.8 KB c338978c download
chat_template.jinja 7.58 KB a8755d82 download
README.md 5.21 KB bd7d92bb download
config.json 3.71 KB ff20d303 download
.gitattributes 1.61 KB 5d80a2d9 download
generation_config.json 213 B 79c8cce3 download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.6-27B-abliterated
language:

  • en
    tags:
  • gguf
  • qwen3_5
  • llama.cpp
  • investigative-journalism
  • osint
  • conversational

qwen3.6-27b-abliterated-journalist-GGUF

GGUF export of tomvaillant/qwen3.6-27b-abliterated-journalist-merged, an investigative journalism and OSINT fine-tune based on huihui-ai/Huihui-Qwen3.6-27B-abliterated.

This is the large tier ship — text-only GGUF for Goose Desktop / llama.cpp on machines with ≥32GB unified memory. Smaller machines (≈16GB) run the qwen3.5-9b-abliterated-journalist-GGUF instead.

Usage

llama-cli -hf tomvaillant/qwen3.6-27b-abliterated-journalist-GGUF:Q4_K_M --jinja

Run as a local OpenAI-compatible server (thinking mode on — see note below):

llama-server -hf tomvaillant/qwen3.6-27b-abliterated-journalist-GGUF:Q4_K_M \
  --port 8081 --ctx-size 16384 --n-gpu-layers 999 --jinja

Pull via Ollama's HuggingFace passthrough:

ollama pull hf.co/tomvaillant/qwen3.6-27b-abliterated-journalist-GGUF:Q4_K_M

When wrapping the Ollama tag in a Modelfile (e.g. for opencode / Spotlight), add explicit PARAMETER stop directives — the fine-tune emits <|endoftext|> at end-of-turn while Ollama's auto-derived stop list only includes <|im_end|>:

FROM hf.co/tomvaillant/qwen3.6-27b-abliterated-journalist-GGUF:Q4_K_M
PARAMETER stop "<|im_end|>"
PARAMETER stop "<|endoftext|>"

Thinking mode is required

Earlier revisions of this card recommended --reasoning off --chat-template-kwargs '{"enable_thinking":false}' to save reasoning tokens. Don't do that on this model. The abliterated Qwen 3.6 family's /no_think codepath is damaged — verified empirically on both the Huihui base and this fine-tune at Q4_K_M with Qwen-recommended sampling (temp=0.6, top_p=0.95, top_k=20): output collapses into multilingual token soup and lock-loops within ~200 tokens. Probable cause: abliteration calibration only covered the thinking codepath, leaving the no-think branch with broken refusal-direction subtraction; Q4 quantization amplifies it.

Keep thinking on. Expect responses ~3–5× longer per turn than the 9B sibling. Bump max_output_tokens (~4096+) and any opencode-style limit.output (16384) if you see truncation inside the reasoning block.

Files

  • qwen3.6-27b-abliterated-journalist-Q4_K_M.gguf — Q4_K_M with imatrix calibration (~15 GB on disk, ~22 GB at runtime; recommended for laptops with ≥32 GB unified memory)
  • chat_template.jinja
  • tokenizer files

Training

  • Adapter: tomvaillant/qwen3.6-27b-abliterated-journalist
  • Merged checkpoint: tomvaillant/qwen3.6-27b-abliterated-journalist-merged
  • Method: LoRA with Unsloth FastModel + TRL SFT, following the official Unsloth Qwen3.5 fine-tune recipe (bf16, r=16, alpha=16, dropout=0, use_gradient_checkpointing="unsloth", optim="adamw_8bit"). Merged into bf16 safetensors via save_pretrained_merged, then converted to GGUF via llama.cpp convert_hf_to_gguf.py + llama-quantize.
  • Dataset: tomvaillant/investigative-journalism-training (687 examples, OSINT methodology)
  • Quantization: Q4_K_M with imatrix calibration on the training corpus. Imatrix gives ~1–3% perplexity recovery over stock Q4_K_M at the cost of a single one-shot calibration pass; per llama.cpp discussion #11088 the benefit is meaningful below Q5, negligible at Q6+.
  • GGUF metadata: ships with tokenizer.chat_template embedded (inlined from chat_template.jinja pre-conversion so Ollama's HF passthrough sees a real template, not the default {{ .Prompt }} fallback).

Sources And Attribution

Training data: tomvaillant/investigative-journalism-training — 687 instruction/response pairs synthesized by Claude Opus 4.6 (Anthropic) from the Buried Signals OSINT and investigative-journalism corpus: OSINT Navigator tool data, Indicator Media briefings, Buried Signals investigative skills, GIJN, Bellingcat, Verification Handbook 3, SPJ Code of Ethics, RCFP, and public manuals from UNESCO, Al Jazeera Media Institute, CiFAR, CIPE, and EJF/TEMPO Institute.

See the dataset card for the full source list, licenses, and per-partner attribution.

Intended Use

Built for local llama.cpp inference in investigative journalism and OSINT workflows. Vision tower is not included in this GGUF — to add multimodal input, layer in mmproj-BF16.gguf from unsloth/Qwen3.6-27B-GGUF (byte-identical because the vision tower was frozen during training).

Treat outputs as leads, not verified findings.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-08model card: drop personal Mac-mini reference; describe hardware tiers generic...4771c685.2 KB
    Loading...
  2. 2026-06-17Point attribution to dataset card (single source of truth)efccfd05.2 KB
    Loading...
  3. 2026-05-25README: ship correct Q4_K_M filename + reverse no_think guidance7a7fc5c6.7 KB
    Loading...
  4. 2026-05-25docs: expand SOURCES attribution (Cached Q&A, Amditis, FTM primary sources, n...048c7075.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration