← back to catalog · registered 2026-08-22 13:56

cloud19/G4-MeroMero-26B-FP8-Dynamic-Uncensored

cloud19 Gemma 24B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cloud19%2FG4-MeroMero-26B-FP8-Dynamic-Uncensored"
Response includes
  • classification m-uncensored
  • files 9
  • hub_downloads_all_time 5,354
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
Likes
0
Model age
3mo ago
created 2026-06-21
Downloads over time
Now6.7K→from35↑18,943%
02.4K4.9K7.3K35 on Jun 246.7K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors gemma4 fp8 llm-compressor vllm moe base_model:llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic base_model:quantized:llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic license:other compressed-tensors region:us

Related

Total size
26.7 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-21 14:54

Files by quantization

Auxiliary files 9 files 26.7 GB
model.safetensors 26.7 GB 483c1fd1 download
tokenizer.json 30.7 MB a2619fe1 download
config.json 211 KB aa9622f3 download
chat_template.jinja 16.9 KB 7fa57f97 download
tokenizer_config.json 2.75 KB 24203607 download
README.md 2.45 KB 1df29e4c download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 230 B 0b2957da download
generation_config.json 204 B f2d58f06 download

README current version from Hugging Face


license: other
base_model: llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic
tags:

  • fp8
  • llm-compressor
  • vllm
  • moe

G4-MeroMero-26B FP8 Dynamic Uncensored

This repository contains an FP8 Dynamic quantized version of the original model llmfan46/G4-MeroMero-26B-A4B-it-uncensored-heretic.

Quantization Methodology

The quantization process was conducted using the llm-compressor library from Neural Magic, executing a data-free one-shot dynamic quantization.

To preserve the cognitive logic, output consistency, and potential multimodal projection capabilities of the model, specific critical layers were excluded from quantization and preserved in their native BFloat16 precision:

  • Input Embeddings (embed_tokens)
  • Output Prediction Head (lm_head)
  • Vision Tower / Encoder (vision_tower)
  • Vision Projector (multi_modal_projector)
  • MoE Routers / Gate layers (gate)

All standard internal Linear projections within the transformer blocks were quantized to FP8 Dynamic (weights are statically quantized to FP8 per-channel, and activations are dynamically quantized to FP8 per-token). This approach yields a 2x reduction in checkpoint size while maintaining near-lossless generation quality (~99.5% retention compared to the BF16 baseline).

Inference Recommendations

For optimal throughput, latency, and memory footprint, we recommend utilizing the vLLM engine, which natively supports dynamic FP8 activation scaling and execution on modern GPU architectures (such as NVIDIA Blackwell B200 / Hopper H100).

Recommended vLLM Launch Configuration

To maximize generation performance and optimize KV cache utilization, launch the vLLM server with the following settings:

python3 -m vllm.entrypoints.openai.api_server \
    --model cloud19/G4-MeroMero-26B-FP8-Dynamic-Uncensored \
    --served-model-name gemma-rp-uncensored \
    --dtype bfloat16 \
    --quantization compressed-tensors \
    --load-format safetensors \
    --kv-cache-dtype fp8 \
    --trust-remote-code

High-Performance Backends

To achieve maximum inference speed, ensure the following environment variables and libraries are active:

  • FlashInfer Sampler: Improves sampling speeds significantly.
  • SageAttention: Set VLLM_ATTENTION_BACKEND="sageattn" to utilize the SageAttention backend for faster context handling, or fall back to flash_attn on supported hardware.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-21Upload folder using huggingface_hub219c0bd2.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration