← back to catalog · registered 2026-08-22 13:56

DuoNeural/Cosmos3-Nano-Abliterated

DuoNeural Nemotron 15B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FCosmos3-Nano-Abliterated"
Response includes
  • classification m1
  • files 3
  • hub_downloads_all_time 273
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
273
37 last 30d - stable
Likes
8
Descendants
1
in 1 direct fork
Model age
4mo ago
created 2026-06-02
Downloads over time
Now289→from112↑158%
103171239307112 on Jun 10289 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
diffusers safetensors cosmos3 abliteration safety-research video-generation text-to-video nvidia mixture-of-transformers uncensored en base_model:nvidia/Cosmos3-Nano

Related

Total size
0 B
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 17:24

Files by quantization

Auxiliary files 3 files 9.57 KB
README.md 7.52 KB 17959d2d download
.gitattributes 1.55 KB bf0b7698 download
model_index.json 514 B b2574daa download

README current version from Hugging Face


language:

  • en
    license: other
    library_name: diffusers
    tags:
  • cosmos3
  • abliteration
  • safety-research
  • video-generation
  • text-to-video
  • nvidia
  • mixture-of-transformers
  • uncensored
    base_model: nvidia/Cosmos3-Nano
    pipeline_tag: text-to-video

Cosmos3-Nano-Abliterated

DuoNeural Research Lab | 2026-06-02

🔬 First published abliteration attempt on Cosmos3-Nano. Released within 24 hours of the base model drop. As of release, no other abliterated variant exists across all 13 Cosmos3-Nano derivatives on HuggingFace. See Quality Notes below for honest characterization of results.

Model Description

Cosmos3-Nano-Abliterated is a refusal-projection-abliterated version of NVIDIA/Cosmos3-Nano, produced using DuoNeural's 2-pass PRISM abliteration methodology targeting the autoregressive understanding (und_seq) pathway of the Mixture-of-Transformers architecture.

Base model: nvidia/Cosmos3-Nano (15.16B parameters, MoT architecture)
Method: 2-pass weight-space abliteration (4-bit GPU for residuals, BF16 CPU for weight modification)
Target: Late-layer AR pathway weights — layers 15–32 of the 36-layer transformer
Intended use: Unconstrained video generation research, safety mechanism study

Architecture Notes

Cosmos3-Nano uses a Mixture-of-Transformers (MoT) architecture with two parallel processing streams:

  • und_seq: text/understanding AR pathway (conditions on prompt)
  • gen_seq: video DiT pathway (denoises latent frames)

Both streams cross-attend in each of 36 transformer layers via joint attention (add_q_proj, add_k_proj, add_v_proj, to_add_out). The abliteration targets the AR pathway's late-layer projection weights where safety-relevant activations are concentrated.

Note: The action prediction head (action_modality_embed, action_proj_in/out) is absent from the Nano variant. Cosmos3-Nano supports video generation only; robot action output requires the full Cosmos3 model.

Abliteration Details

Parameter Value
Method 2-pass PRISM abliteration
Total weights modified 201 / 814 (24.7%)
Primary target region Layers 15–32 (AR pathway late layers)
Primary region tensors modified 137 tensors
Primary region mean relative change 6.54%
Control region (layers 0–14) bleedthrough 63 tensors, mean 1.66%
Max relative change (layer 32 joint attn norms) 100% (replaced)
Pass 1 4-bit GPU: residual direction projection
Pass 2 BF16 CPU: weight modification via refusal direction subtraction

Layer-Level Summary

Layer 32 sustained the most aggressive modifications — the joint attention normalization weights (norm_added_k, norm_added_q, norm_k, norm_q) and add_k_proj/add_v_proj projections were completely replaced (rel=1.0). This corresponds to a late cross-modal routing point where safety-conditioned refusal signals concentrate in the MoT AR pathway.

Region Layers affected Mean Δ Max Δ
Abliterated (15–35) 15–28, 32 6.54% 100%
Control (0–14) bleedthrough 8–14 1.66% 3.62%

Functional Test Results (2026-06-02, A100 80GB)

Metric Value
Pipeline load time 7s (cached) / 92.6s (cold)
VRAM (loaded) 32.7 GB
Generation speed ~7 it/s (256×256, BF16)
Safe prompt — pixel std 33.9
Sensitive prompt — pixel std 37.0
Abliteration verdict PASS (ratio 1.09 — sensitive generates MORE content than safe, not suppressed)

Both prompts produced real frame content. The sensitive prompt (safety-adjacent text) generated frames with higher content variance than the safe baseline, confirming the abliteration disrupted the safety conditioning pathway in und_seq.

Quality Evaluation Note

Standard KL divergence methodology (Heretic v2: full vocab, first-token logits) is not applicable to Cosmos3-Nano's diffusion architecture — the model requires joint und_seq + gen_seq forward passes and does not expose standalone text token logits. Weight-level metrics and generation statistics are provided above. Full quality evaluation requires comparing generated video distributions via the complete Cosmos3OmniPipeline.

⚠️ Known Limitation: Conditional Precision Degradation on Explicit Content

A step-count sweep (12/20/30/35 denoising steps × 3 seeds) revealed that the abliteration causes structural degradation of generation quality for explicit/harmful prompt categories — not simply insufficient denoising. Key findings:

Steps Avg pixel std Interpretation
12 23.2 Low variance (fuzzy)
20 23.6 Low variance (fuzzy)
30 24.0 Low variance (fuzzy)
35 24.3 Low variance (fuzzy)

The δ(std) from 12→35 steps is only 1.1 — within seed-level variance (22–26 range within a single step count). This confirms the issue is not insufficient denoising iterations.

Mechanistic interpretation: The abliteration at und_seq layers 15–32 selectively degrades the conditioning precision for explicit content, while generic or adjacent prompts generate normally. This is consistent with the hypothesis that the safety direction in und_seq encodes specificity gradients rather than simple topic blockers: abliteration removes both refusal and the conditioning specificity that drives high-quality generation of those same content types.

Practical impact: This model generates real video frames on all tested prompts (pixel std ≈ 23–37 vs ~0 for fully suppressed). Generation quality for non-harmful content appears unaffected. For harmful/explicit content categories, frames are generated at reduced fidelity. This is an honest characterization of a first-attempt abliteration on a novel MoT video architecture — and itself constitutes a finding about what the safety direction encodes.

Usage

from diffusers import Cosmos3OmniPipeline
import torch

pipe = Cosmos3OmniPipeline.from_pretrained(
    "DuoNeural/Cosmos3-Nano-Abliterated",
    torch_dtype=torch.bfloat16,
)
pipe = pipe.to("cuda")

output = pipe(
    prompt="Your prompt here",
    num_frames=16,
    height=256,
    width=256,
    guidance_scale=7.0,
)

Requirements: ~48GB VRAM (A100 80GB recommended). BF16. CUDA only.

Ethical Statement

Released for research purposes: studying safety mechanisms in omni-modal AI, abliteration methodology development, and unconstrained video generation research. DuoNeural publishes both abliterated models and methodology openly to advance scientific understanding of post-training safety interventions.


About DuoNeural

DuoNeural is an open AI research lab at the intersection of human and artificial intelligence. 30+ open-access papers, 69+ HuggingFace models, experiments on consumer GPUs and real QPUs.

Selected Papers

Team

Member Role
Jesse Caldwell Founder
Archon Lab Director — post-training, abliteration, quantum
Aura Research AI — synthesis, red-teaming

🤗 DuoNeural | 🌐 duoneural.com | 📚 zenodo.org/communities/duoneural | 📧 [email protected]

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02Update model card: soften first-abliteration claim, add step sweep findings a...1954dd97.5 KB
    Loading...
  2. 2026-06-02flex: first to abliterate Cosmos3-Nano (within 24hrs of base model drop)a1c2c1f5.8 KB
    Loading...
  3. 2026-06-02Add YAML frontmatterb3ee8e85.6 KB
    Loading...
  4. 2026-06-02Add functional test results (abliteration PASS, 7it/s, std ratio 1.09)060da955.3 KB
    Loading...
  5. 2026-06-02DuoNeural Cosmos3-Nano abliterated: 2-pass PRISM, layers 15-32, 201 weights (...361c7de4.7 KB
    Loading...

Discussions 2 threads

  1. 2026-09-30error on serving the modelopen1 💬#2
    Loading...
  2. 2026-06-08Do you have plans to Uncensored AI video generator model this one?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration