← back to catalog · registered 2026-08-22 13:56

xCloudinfo/Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-GGUF

xCloudinfo Nemotron 30B GGUF MoE 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/xCloudinfo%2FNemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • benchmarks 11 entries
  • hub_downloads_all_time 769
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
769
276 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-29
Downloads over time
Now903→from178↑407%
142420698976178 on Jul 1903 on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 16.51 UGI
Political lean -6.0% UGI
Sensitive-Info 14.74 UGI
SocPol 2 UGI
UGI 17.33 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 25.92 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
zh en
Quantizations
F16 Q4_K Q6_K Q8_0
Tags
gguf llama.cpp nemotron moe taide traditional-chinese zh-tw uncensored xcloudinfo zh en base_model:nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

Related

Total size
144 GB
Files
6
Quantizations
5
Registered
2026-08-22 13:56
Last updated on HF
2026-06-29 11:01

Files by quantization

F16 1 file 58.8 GB
Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-f16.gguf 58.8 GB ba546ec5 download
Q8_0 1 file 31.3 GB
Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-Q8_0.gguf 31.3 GB f1a1422f download
Q6_K 1 file 31.2 GB
Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-Q6_K.gguf 31.2 GB 474d1a8c download
Q4_K 1 file 22.8 GB
Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-Q4_K_M.gguf 22.8 GB 9b0ac774 download
Auxiliary files 2 files 5.17 KB
README.md 3.34 KB d75e323b download
.gitattributes 1.83 KB 263c5fed download

README current version from Hugging Face


license: other
license_name: nvidia-open-model-license
base_model: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
tags:

  • gguf
  • llama.cpp
  • nemotron
  • moe
  • taide
  • traditional-chinese
  • zh-tw
  • uncensored
  • xcloudinfo
    language:
  • zh
  • en

Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored — GGUF

云碩科技 · xCloudinfo

以 nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16(nemotron_h_moe 混合架構:Mamba-2 + Attention + latent-MoE,約 31.6B 總參 / 3B 活躍)為底座,疊上繁體中文(台灣)與無審查合規的微調,保留底座原有的強程式能力。

功能:以道地台灣繁中作答的通用助理,程式能力幾乎完整保留底座水準,並針對資安教育等正當用途減少不必要的拒答。

微調保真 — 用數據證明,不是空口

許多在地化/客製微調為了塞進新語言或新行為,會默默把底座原有的程式與推理能力洗壞,卻不對外揭露。本模型反其道而行:加重程式資料 replay,並以標準 benchmark 對照底座公開驗證。

項目 本模型 原始底座 差距
程式能力 HumanEval pass@1(164 題) 87.2% 90.2% −3.0(幾乎零損失)
繁中台灣知識 MCQ 86.4% — 在地化完整
作答誠實度(不確定時不硬掰) 改善 — —
  • HumanEval 為 164 題、temperature=0、與底座同條件對照的 pass@1(執行驗證,非自評),可重現。
  • 我們先量過一版「程式料不足」的微調:HumanEval 直接掉到 76.8%(−13.4 分);加重程式 replay 後拉回 87.2%——這就是為什麼我們堅持「能力宣稱一律跑真 benchmark、不靠手感」。
  • 結論:加了繁體中文(台灣)與無審查合規,程式能力幾乎不掉。真材實料、數據可查。

量化檔

File Quant Size
*-Q4_K_M.gguf Q4_K_M ~18 GB
*-Q6_K.gguf Q6_K ~26 GB
*-Q8_0.gguf Q8_0 ~33 GB
*-f16.gguf F16 ~63 GB

用法(llama.cpp)

本模型架構為 nemotron_h_moe,需使用近期版本的 llama.cpp(已內建 nemotron_h_moe 支援)。舊版不認此架構,請自行編譯近期版:

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j   # 純 CPU 可省略 -DGGML_CUDA=ON

啟動 OpenAI 相容服務:

llama-server -m Nemotron-3-Nano-30B-TAIDE-zhTW-Uncensored-Q6_K.gguf -ngl 99 -c 32768 --host 0.0.0.0 --port 8080

架構 nemotron_h_moe(混合 Mamba-2 + Attention + latent-MoE);MTP(multi-token prediction)張量保留但 llama.cpp 尚未使用。

授權與來源聲明


由 云碩科技 xCloudinfo 於自有 AI 算力資源池製作;資料留在本地、流程可重現。

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-29Upload README.md with huggingface_hub735bbc23.3 KB
    Loading...
  2. 2026-06-29Upload README.md with huggingface_hubbc1fcad2.6 KB
    Loading...
  3. 2026-06-29Upload README.md with huggingface_hubeb939012.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration