license: apache-2.0
language:
- en
- multilingual
tags: - nemotron
- nvidia
- omni
- multimodal
- vision
- image-to-text
- gguf
- uncensored
- mamba2
- moe
- mixture-of-experts
- hybrid-architecture
- aeon-7
pipeline_tag: image-text-to-text
base_model: AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16
Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored — GGUF
🎯 GGUF Conversion of the original AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16
📦 Files
| File | Format | Size | Description |
|---|---|---|---|
Nemotron-3-Nano-Omni-AEON-Uncensored-BF16.gguf |
BF16 | ~59GB | Full precision LLM weights |
Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf |
Q6_K | ~32GB | Quantized LLM weights (recommended) |
mmproj-Nemotron-3-Nano-Omni-BF16.gguf |
BF16 | ~1.5GB | Vision encoder projection |
🧠 Model Architecture
Source Model
- Creator: AEON-7
- Original: Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16
- Base Architecture: NVIDIA NemotronH (Nemotron-3-Nano)
LLM Backbone — NemotronHForCausalLM
- Layers: 52 hybrid layers
- Hidden Size: 2,688
- MoE: 128 routed experts, 6 active per token, 1 shared expert
- Hybrid Pattern:
MEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEMEM*EMEMEMEMEM= Mamba2 (SSM)E= Expert (Mixture of Experts)*= Attention
- Context Length: 131,072 tokens
- Vocab Size: 131,072
Vision Encoder — ViT-Huge (RADIO v2.5)
- Layers: 32
- Hidden Size: 1,280
- Image Size: 512×512
- Patch Size: 16
Audio Encoder — Parakeet
- Layers: 24
- Hidden Size: 1,024
- Sample Rate: 16kHz
🚀 Usage with llama.cpp
Text-Only Inference
./llama-cli \
-m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
-p "Hello, how are you?" \
-n 256
Vision (Image Understanding)
./llama-llava-cli \
-m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
--mmproj mmproj-Nemotron-3-Nano-Omni-BF16.gguf \
--image image.jpg \
-p "Describe this image in detail."
📝 Conversion Details
| Step | Tool | Details |
|---|---|---|
| LLM Extraction | Python safetensors |
Extracted language_model.* from omni checkpoint |
| GGUF Conversion | convert_hf_to_gguf.py |
llama.cpp — NemotronHForCausalLM → BF16 GGUF |
| Quantization | llama-quantize |
BF16 → Q6_K |
| mmproj | convert_hf_to_gguf.py --mmproj |
NemotronNanoV2VLModel — ViT + projector → BF16 |
🙏 Credits & Attribution
This is a GGUF conversion only. All model weights originate from:
| Role | Entity |
|---|---|
| 🧠 Original Model | AEON-7 |
| 📐 Architecture | NVIDIA NemotronH |
| 🔧 GGUF Conversion | @hotdogs |
| 🛠️ Tools | llama.cpp by Georgi Gerganov & contributors |
💡 Note: This repo contains only GGUF format files converted from the original AEON-7 model. For the original PyTorch/SafeTensors weights, please visit the source model.
⚠️ Disclaimer
This is an uncensored model. Use responsibly and in compliance with applicable laws and regulations.