← back to catalog · registered 2026-08-22 13:56

DuoNeural/Cosmos3-Nano-GPTQ-4bit-Abliterated

DuoNeural second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FCosmos3-Nano-GPTQ-4bit-Abliterated"
Response includes
  • classification m1
  • files 6
  • hub_downloads_all_time 319
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
319
83 last 30d - stable
Likes
3
Model age
4mo ago
created 2026-06-02
Downloads over time
Now359→from38↑845%
2214526839138 on Jun 10359 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
diffusers qwen3_vl_text cosmos3 gptq 4bit abliteration quantization video-generation text-to-video nvidia mixture-of-transformers uncensored

Related

Total size
10.3 GB
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 16:14

Files by quantization

Auxiliary files 6 files 10.3 GB
model-00001-packed.safetensors 4.32 GB 0bb8fbfc download
model-00003-packed.safetensors 3.05 GB c6b0d508 download
model-00002-packed.safetensors 2.93 GB e70816f4 download
README.md 2.49 KB c918473e download
config.json 1.63 KB d55189ea download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


language:

  • en
    license: other
    library_name: diffusers
    tags:
  • cosmos3
  • gptq
  • 4bit
  • abliteration
  • quantization
  • video-generation
  • text-to-video
  • nvidia
  • mixture-of-transformers
  • uncensored
    base_model: DuoNeural/Cosmos3-Nano-Abliterated
    pipeline_tag: text-to-video

Cosmos3-Nano-GPTQ-4bit-Abliterated

DuoNeural Research Lab | 2026-06-02

🥇 First int4-quantized abliterated version of Cosmos3-Nano. The only model combining NVIDIA's Cosmos3-Nano abliteration (safety conditioning removal) with GPTQ 4-bit compression.

What This Is

Cosmos3-Nano-GPTQ-4bit-Abliterated combines two unique modifications to NVIDIA/Cosmos3-Nano:

  1. Abliteration (from DuoNeural/Cosmos3-Nano-Abliterated): refusal direction removed from the und_seq AR pathway (layers 15–32)
  2. GPTQ 4-bit quantization: 330 linear layers packed to ~11GB (2.74× compression)

The result: unconstrained video generation at ~11GB — runnable on 16GB VRAM with careful memory management.

Model Lineage

nvidia/Cosmos3-Nano (original, 30GB BF16)
    └── abliterate → DuoNeural/Cosmos3-Nano-Abliterated (30GB BF16)
            └── GPTQ + pack → DuoNeural/Cosmos3-Nano-GPTQ-4bit-Abliterated (this model, ~11GB)

For the base model quantization (safety conditioning intact), see DuoNeural/Cosmos3-Nano-GPTQ-4bit.

Quantization Details

Parameter Value
Method GPTQ (weight-only, column-wise)
Bits 4
Group size 128
Packing format DuoNeural nibble v1 (custom int32 nibble-packed)
Compression ~2.74× (30GB → ~11GB transformer)

Note: Custom nibble format — not compatible with auto-gptq/exllama loaders. Manual unpacking required.

Limitations

  • Custom GPTQ format requires manual dequantization (see DuoNeural/Cosmos3-Nano-GPTQ-4bit for format spec)
  • Double quantization: packing re-quantizes already-quantized values; additional error vs single-pass int4
  • For best quality, use DuoNeural/Cosmos3-Nano-Abliterated (full BF16, 32.7GB VRAM)
  • Action prediction head absent from Nano variant

DuoNeural Research Lab | [email protected] | duoneural.com
Papers: Zenodo Community | Models: HuggingFace

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02DuoNeural Cosmos3-Nano-GPTQ-4bit-Abliterated: first abliterated+quantized Cos...c587a622.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration