← back to catalog · registered 2026-08-22 13:56

maci0/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-NVFP4

maci0 6.9B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/maci0%2FHuihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 2,066
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
213 last 30d - stable
Likes
4
Model age
3mo ago
created 2026-06-30
Downloads over time
Now2.1K→from106↑1,907%
57801.6K2.3K106 on Jul 12.1K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 55 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 w4a4 gptq quantized compressed-tensors llm-compressor vllm

Related

Total size
8.72 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 01:51

Files by quantization

Auxiliary files 12 files 8.74 GB
model.safetensors 7.42 GB 3783b91d download
model-towers.safetensors 1.30 GB 16454c7d download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 130 KB eb940ed4 download
README.md 13.0 KB 58543e43 download
chat_template.jinja 7.57 KB a585dec8 download
config.json 6.08 KB 2548b812 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.24 KB 9de16b5b download
processor_config.json 1.16 KB 33818c7f download
recipe.yaml 326 B 0ed8315b download
generation_config.json 164 B 83e369c8 download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated
base_model_relation: quantized
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • nvfp4
  • fp4
  • w4a4
  • gptq
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
  • qwen3_5
  • vision-language
  • thinking
  • uncensored
  • abliterated
  • creative
  • 1m-context

RQ-9B-QYTH-AAbliterated
Qwythos-9B Mythos · NVFP4
9B creative VL · Claude Mythos distill · abliterated (upstream) · 1M-token context.
Params9B
Active9B (dense)
Size7.5 GB
Perplexity8.25
Refusalsn/a
Context1M
MTP headbf16

TL;DR: Huihui-Qwythos-9B-Claude-Mythos, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 7.5 GB, wikitext-2 PPL 8.25, 1M-context creative writer, abliterated.

Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated NVFP4

NVFP4 (W4A4) quantization of
huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated,
packed in the compressed-tensors nvfp4-pack-quantized format with
llm-compressor. Weights are
quantized with GPTQ (error-compensated rounding) and an MSE observer, on a
domain-matched calibration blend.

Near-lossless. Fused layers (q/k/v, gate/up) share one NVFP4 global scale, so vLLM
loads it cleanly with no per-layer-scale warning or fallback. wikitext-2 perplexity for
this build: 8.25.

  • About 7.5 GB on disk versus about 19.3 GB for the bf16 source (about 39%).
  • Built for vLLM on NVIDIA Blackwell, where both the 4-bit weight and 4-bit activation
    paths are accelerated. On pre-Blackwell GPUs vLLM runs it weight-only.
  • Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).

Uncensored / abliterated model. It follows instructions without content guardrails,
including NSFW. Behaviour and alignment are inherited entirely from the base model.

Fidelity

Near-lossless versus the bf16 source, 7.5 GB vs 19.3 GB bf16 (~39%), at wikitext-2 perplexity 8.25. GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

NVFP4 is auto-detected from config.json (compressed-tensors); no quantization flag
needed. --reasoning-parser qwen3 splits the <think> block into reasoning_content.

vllm serve maci0/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated-NVFP4 \
  --served-model-name qwythos-9b-nvfp4 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder
  • Supports up to 1048576 tokens; keep at least 128K to preserve thinking quality.
    --max-model-len 131072 is a safe default; raise it if memory allows.
  • Add --language-model-only to skip the vision tower and free KV cache for text use.
  • The parser flags are not auto-detected; pass them explicitly. Drop the tool-call line
    if you do not need tool calling.

Python (OpenAI client)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
r = client.chat.completions.create(
    model="qwythos-9b-nvfp4",
    messages=[{"role": "user", "content": "Write a short mythic tale about a lantern that remembers the sea."}],
)
print(r.choices[0].message.content)

curl

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "qwythos-9b-nvfp4",
  "messages": [{"role": "user", "content": "Write a short mythic tale about a lantern that remembers the sea."}]
}'

About the base model

An abliterated (uncensored) build of empero-ai/Qwythos-9B-Claude-Mythos-5-1M, a 9B
Qwen3.5-family vision-language model tuned for creative / Mythos writing with a Claude
distillation, and a 1M-token context window.

  • 32 decoder layers: hybrid gated delta-net linear attention plus full attention, dense
    MLP, plus a vision tower for image and video input.
  • 1M context (max_position_embeddings 1048576).
  • Thinking mode by default, with an instruct toggle. In non-thinking mode the base model
    adheres strictly to the first prompt's instructions (output only, no fluff).

Quantization

Scheme NVFP4, W4A4
Weight rounding GPTQ (Hessian-based error compensation), MSE observer
Weights FP4 (E2M1), group_size=16, tensor_group, FP8 (E4M3) group scales, shared across fused layers
Activations FP4, dynamic per-group, FP8 (E4M3) scales
Quantized all language-model Linear layers
Kept in bf16 vision tower (model.visual.*), lm_head, MTP head
Untouched gated delta-net Conv1d and SSM params (A_log, dt_bias), never Linear

GPTQ is a quantization-time cost only; inference speed and format are identical to
plain round-to-nearest NVFP4, but it chooses better 4-bit values.

Calibration: 512 domain-matched samples (long reasoning + general chat + code),
max_seq_len=2048, text-only path through the VL model.

Recommended sampling

Thinking mode is the default.

  • Thinking, general / creative: temperature=1.0, top_p=0.95, top_k=20
  • Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20
  • To run non-thinking, set {%- set enable_thinking = false %} in the chat template, or
    pass extra_body={"chat_template_kwargs": {"enable_thinking": false}}.

Reproduction

Toolchain: llmcompressor==0.12.0, compressed-tensors==0.17.1, transformers==5.12.1,
torch==2.11.0+cu130, on an NVIDIA GB10 (Blackwell, sm_121). llm-compressor 0.12 shares
the NVFP4 global scale across fused layers automatically (q/k/v, gate/up).

Related

Notes

  • Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • --reasoning-parser and --tool-call-parser are not auto-detected; pass them explicitly.
  • Thinking mode is on by default; toggle it via the chat template or chat_template_kwargs.
  • No refusal guardrails; you are responsible for how you use it.

License

Apache-2.0, following the base model. Intended use and all responsibility for use follow
the base model.

Credits

Part of Rogue Quants · NVFP4 component datasheets · collection. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).

README history 15 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04uniform schema, no em-dashesccff50f13 KB
    Loading...
  2. 2026-08-04card: uniform characteristics schema (params/active/experts/etc)83142fd13 KB
    Loading...
  3. 2026-08-03trim redundant Fidelity table (covered by header)4e5315411.4 KB
    Loading...
  4. 2026-08-03card: readability - sans labels, mono values, more spacing64c33a711.6 KB
    Loading...
  5. 2026-08-03card: fix table borders + readability5338d7711.6 KB
    Loading...
  6. 2026-08-03card: die sigil + a11y table semanticsb1c30ab11.2 KB
    Loading...
  7. 2026-08-03card: DATASHEET design language6585d9d10.6 KB
    Loading...
  8. 2026-07-07Add Ornith-35B MoE abliterated to Related siblingsd917c6f12.9 KB
    Loading...
  9. 2026-07-04Add Qwopus-27B-v2 abliterated to Related siblingsd09ff2012.8 KB
    Loading...
  10. 2026-07-02Add abliterated Qwopus-27B to Related siblings7d3298912.7 KB
    Loading...
  11. 2026-07-01Enrich card: TL;DR, Quickstart (Python+curl), Related, Notese12795312.6 KB
    Loading...
  12. 2026-07-01Add Fidelity section9e839ca10.8 KB
    Loading...
  13. 2026-07-01Consistency pass: unified banner, remove em dashes, footer9c13b6a10.4 KB
    Loading...
  14. 2026-07-01Add visual flair to card7c46d059.3 KB
    Loading...
  15. 2026-06-30Add model card8083e6e4.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration