← back to catalog · registered 2026-09-27 01:57

edp1096/Huihui-GLM-5.3-Flash-abliterated-NVFP4

edp1096 Glm GGUF multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/edp1096%2FHuihui-GLM-5.3-Flash-abliterated-NVFP4"
Response includes
  • classification unknown
  • files 47
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-27

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
Model Optimizer safetensors glm5_next nvidia ModelOpt GLM-5 quantized FP4 fp4 image-text-to-text conversational base_model:huihui-ai/GLM-5.3-Flash-abliterated-GGUF

Related

Total size
190 GB
Files
47
Quantizations
1
Registered
2026-09-27 01:57
Last updated on HF
2026-09-27 02:00

Files by quantization

Auxiliary files 47 files 190 GB
model-00003-of-00033.safetensors 8.33 GB f184e8b7 download
model-00001-of-00033.safetensors 8.33 GB 5c7ee3e2 download
model-00002-of-00033.safetensors 8.33 GB 7afab148 download
model-00005-of-00033.safetensors 5.57 GB bc763b5e download
model-00004-of-00033.safetensors 5.57 GB dd08e62f download
model-00019-of-00033.safetensors 5.51 GB b96eec11 download
model-00017-of-00033.safetensors 5.51 GB 7be78c24 download
model-00025-of-00033.safetensors 5.51 GB c2c44e89 download
model-00027-of-00033.safetensors 5.51 GB 86e8a97c download
model-00015-of-00033.safetensors 5.51 GB 447d1460 download
model-00013-of-00033.safetensors 5.51 GB d1d5c86f download
model-00031-of-00033.safetensors 5.51 GB 3c7461c5 download
model-00029-of-00033.safetensors 5.51 GB a705d67e download
model-00033-of-00033.safetensors 5.51 GB c6fa7921 download
model-00023-of-00033.safetensors 5.51 GB 7e309f80 download
model-00020-of-00033.safetensors 5.51 GB 781f82cd download
model-00021-of-00033.safetensors 5.51 GB 5de82be7 download
model-00006-of-00033.safetensors 5.51 GB 41416899 download
model-00030-of-00033.safetensors 5.51 GB 044b086f download
model-00008-of-00033.safetensors 5.51 GB 92dc3046 download
model-00028-of-00033.safetensors 5.51 GB 14ceb9b9 download
model-00032-of-00033.safetensors 5.51 GB 98ceb03c download
model-00022-of-00033.safetensors 5.51 GB 8f27fac3 download
model-00007-of-00033.safetensors 5.51 GB 0e7163d4 download
model-00009-of-00033.safetensors 5.51 GB 4770820b download
model-00010-of-00033.safetensors 5.51 GB fb146474 download
model-00011-of-00033.safetensors 5.51 GB 3ed00bde download
model-00024-of-00033.safetensors 5.51 GB e3832443 download
model-00026-of-00033.safetensors 5.51 GB c9c1529e download
model-00016-of-00033.safetensors 5.51 GB 295d41e3 download
model-00018-of-00033.safetensors 5.51 GB 2a4dbeb3 download
model-00014-of-00033.safetensors 5.51 GB 37cd788e download
model-00012-of-00033.safetensors 5.51 GB eeaa3799 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 15.5 MB 26765b26 download
.quant_summary.txt 602 KB 388d3ad2 download
transfer-manifest.json 41.0 KB 24fd46d4 download
config.json 16.3 KB d0e8e75d download
README.md 10.7 KB 938d4859 download
chat_template.jinja 10.4 KB fb94d40d download
README.nvidia.md 10.00 KB 08fa91a5 download
hf_quant_config.json 7.81 KB 559a76e5 download
runtime-qualification.json 2.64 KB 612b67c1 download
.gitattributes 1.60 KB a09db2ea download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 215 B 46d04e68 download

README current version from Hugging Face


pipeline_tag: image-text-to-text
base_model:

  • nvidia/GLM-5.3-Flash-NVFP4
  • huihui-ai/GLM-5.3-Flash-abliterated-GGUF
    license: mit
    library_name: Model Optimizer
    tags:
  • nvidia
  • ModelOpt
  • GLM-5
  • quantized
  • FP4
  • fp4

Huihui-GLM-5.3-Flash-abliterated-NVFP4

Base:

Conversion: DQ(NVIDIA NVFP4) + DQ(Huihui GGUF) - DQ(Unsloth GGUF). Original activation scales are retained. GGUF residuals remain; equivalence to Huihui BF16 is not claimed. See transfer-manifest.json.

Validated on two DGX Sparks (TP2): 1,047,622 input tokens without prefix-cache reuse, five retrieval records correct, and nine API regression tests passed. See runtime-qualification.json.


Model Overview

Description:

The NVIDIA GLM-5.3-Flash NVFP4 model is the quantized version of ZAI's GLM-5.3-Flash model, which is an auto-regressive language model that uses an optimized transformer architecture. GLM-5.3-Flash is a natively multimodal Mixture-of-Experts (MoE) model for reasoning, coding, and agentic tasks; it uses a hybrid sparse and linear attention architecture with Manifold-Constrained Hyper-Connections (mHC) to support a long context. For more information, please check here. The NVIDIA GLM-5.3-Flash NVFP4 model is quantized with Model Optimizer.

This model is ready for commercial or non-commercial use.

Third-Party Community Consideration

This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party's requirements for this application and use case; see link to Non-NVIDIA (GLM-5.3-Flash) Model Card from ZAI.

License/Terms of Use:

MIT

Deployment Geography:

Global

Use Case:

Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots, RAG systems, and other AI-powered applications.

Release Date:

Huggingface 09/09/2026 via https://huggingface.co/nvidia/GLM-5.3-Flash-NVFP4

References

Nvidia Model Optimizer: https://github.com/NVIDIA/Model-Optimizer

Model Architecture:

Architecture Type: Transformers

Network Architecture: GLM-5.3-Flash (Glm5NextForConditionalGeneration)

Number of Model Parameters: 320B in total and 18B activated

This model was developed based on GLM-5.3-Flash

Input:

Input Type(s): Text, Image, Video

Input Format(s): String, Red, Green, Blue (RGB), Video (MP4/WebM)

Input Parameters: One-Dimensional (1D), Two-Dimensional (2D), Three-Dimensional (3D)

Other Properties Related to Input: Context length up to 1M

Output:

Output Type(s): Text

Output Format: String

Output Parameters: 1D (One-Dimensional): Sequences

Other Properties Related to Output: None

Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA's hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.

Software Integration:

Supported Runtime Engine(s):

  • vLLM
  • SGLang

Supported Hardware Microarchitecture Compatibility:

  • NVIDIA Blackwell

Preferred Operating System(s):

  • Linux

The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.

Model Version(s):

The model version is NVFP4 1.0 version and is quantized with nvidia-modelopt v0.47.0

Training and Evaluation Datasets:

Calibration Dataset:

** Link: cnn_dailymail, Nemotron-Post-Training-Dataset-v2

** Data Collection Method by dataset: Automated.

** Labeling Method by dataset: Automated.

** Properties: The cnn_dailymail dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The Nemotron-Post-Training-Dataset-v2 is a post-training dataset curated by NVIDIA containing multi-turn conversations across diverse topics.

Training Dataset:

** Data Modality: Undisclosed

** Data Collection Method by dataset: Undisclosed

** Labeling Method by dataset: Undisclosed

** Properties: Undisclosed

Evaluation Dataset:

Datasets: GPQA Diamond, SciCode, MMMU Pro, AA-LCR, IFBench, Terminal Bench 2.1

Data Collection Method by dataset: Hybrid: Automated, Manually-Collected

Labeling Method by dataset: Hybrid: Manually-Labeled, Automated

Properties: We evaluated the model on text-based reasoning, coding, agentic tool-use, and multimodal benchmarks: GPQA Diamond contains 448 graduate-level multiple-choice questions written by domain experts in biology, physics, and chemistry. MMMU Pro is the more challenging version of the Massive Multi-discipline Multimodal Understanding benchmark, measuring college-level multimodal reasoning across diverse disciplines with expanded answer choices and a vision-only input setting.SciCode evaluates scientific coding capabilities. AA-LCR (Artificial Analysis Long Context Recall) evaluates a model's ability to accurately retrieve and recall information from long input contexts. IFBench is a benchmark for evaluating instruction-following capabilities across diverse and structured task constraints. Terminal-Bench 2.1 is an open-source evaluation framework designed to test AI agents on 89 complex, real-world tasks inside sandboxed command-line and container environments.

Inference:

Acceleration Engine: vLLM, SGLang

Test Hardware: NVIDIA Blackwell GB200

Post Training Quantization

This model was obtained by quantizing the weights and activations of GLM-5.3-Flash to NVFP4 data type, ready for inference with vLLM. Only the weights and activations of the linear operators within transformer blocks in sparse MoE shared experts and dense MLP are quantized. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 3.33x.

** modelopt PTQ recipe: nvfp4_experts_dense_mlp-kv_fp8_cast

Usage

vLLM

To serve this checkpoint with vLLM, start from vllm/vllm-openai:glm53-flash-arm64-cu130 and run:

pip install -U "transformers>=5.16.1" && \
vllm serve /checkpoint \
  --served-model-name nvidia/GLM-5.3-Flash-NVFP4 \
  --host 0.0.0.0 --port 8000 \
  --tensor-parallel-size 4 \
  --data-parallel-size 1 \
  --enable-expert-parallel \
  --enable-ep-weight-filter \
  --reasoning-parser glm45 \
  --kv-cache-dtype fp8 \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}' \
  --max-num-batched-tokens 8192 \
  --enable-chunked-prefill \
  --max-num-seqs 32 \
  --gpu-memory-utilization 0.90

SGLang

To serve this checkpoint with SGLang on 4 NVIDIA Blackwell GPUs, run:

sglang serve \
  --model-path nvidia/GLM-5.3-Flash-NVFP4 \
  --quantization modelopt_fp4 \
  --tp-size 4 \
  --dsa-prefill-backend trtllm \
  --dsa-decode-backend trtllm \
  --kv-cache-dtype fp8_e4m3 \
  --moe-runner-backend flashinfer_cutlass \
  --reasoning-parser glm45 \
  --tool-call-parser glm47 \
  --mem-fraction-static 0.85 \
  --host 0.0.0.0 \
  --port 30000

Evaluation

The accuracy benchmark results are presented in the table below:

Precision GPQA Diamond SciCode MMMU Pro AA-LCR IFBench Terminal Bench 2.1
BF16 0.9217 0.5621 0.7688 0.71 0.613 0.8258
NVFP4 0.9211 0.5769 0.763 0.7106 0.6054 0.8315

Baseline: GLM-5.3-Flash-BF16. Benchmarked with temperature=1.0, top_p=0.95. GPQA Diamond, SciCode, MMMU-Pro, AA-LCR and IFBench used max_new_tokens=327,680; Terminal Bench 2.1 used max_new_tokens uncapped.

Model Limitations:

The base model was trained on data that contains toxic language and societal biases originally crawled from the internet. Therefore, the model may amplify those biases and return toxic responses especially when prompted with toxic prompts. The model may generate answers that may be inaccurate, omit key information, or include irrelevant or redundant text producing socially unacceptable or undesirable text, even if the prompt itself does not include anything explicitly offensive.

Ethical Considerations

NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

Please make sure you have proper rights and permissions for all input image and video content; if image or video includes people, personal health information, or intellectual property, the image or video generated will not blur or maintain proportions of image subjects included.

Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.