← back to catalog · registered 2026-08-22 13:56

jabbatheduck/Laguna-S-2.1-Uncensored-GGUF

jabbatheduck GGUF MoE 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jabbatheduck%2FLaguna-S-2.1-Uncensored-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • hub_downloads_all_time 1,357
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
165 last 30d - stable
Likes
2
Model age
2mo ago
created 2026-08-04
Downloads over time
Now1.4K→from932↑55%
9061.1K1.3K1.5K932 on Aug 51.4K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en de
Tags
llama.cpp gguf laguna moe uncensored quantized bilingual code text-generation en de base_model:poolside/Laguna-S-2.1

Related

Total size
226 GB
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-07 16:43

Files by quantization

Auxiliary files 6 files 226 GB
laguna-s-2.1-q4_k_m-uncensored.gguf 66.6 GB 83a4ae82 download
laguna-s-2.1-uncensored-balanced.gguf 62.7 GB d988eb5b download
laguna-s-2.1-uncensored-compact.gguf 48.6 GB 0d9b8295 download
laguna-s-2.1-uncensored-ultracompact.gguf 48.0 GB e7bcd257 download
README.md 3.66 KB 6ec16ba9 download
.gitattributes 1.77 KB eac888e2 download

README current version from Hugging Face


license: openmdw-1.1
base_model: poolside/Laguna-S-2.1
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
language:

  • en
  • de
    tags:
  • laguna
  • moe
  • uncensored
  • gguf
  • quantized
  • bilingual
  • code

Laguna-S-2.1-Uncensored GGUF

GGUF quantization of ressl/Laguna-S-2.1-Uncensored, derived from poolside/Laguna-S-2.1.

Original uncensoring work by Robert Ressl — full model card and evaluation details are at the source repository.

Quantized Variants

All variants use the oscar llama.cpp build (v4310aa4f8) with native LLM_ARCH_LAGUNA support. The BF16 baseline was converted with convert_hf_to_gguf.py from the oscar repo, which handles Laguna's per-layer alternating head counts, sigmoid MoE routing, mixed full/sliding attention RoPE, and the shared-expert topology.

All quantizations use --leave-output-tensor to preserve output.weight in BF16 for generation quality.

File Size Precision Notes
laguna-s-2.1-bf16-uncensored.gguf 220 GiB BF16 Lossless round-trip from safetensors
laguna-s-2.1-q8_0-uncensored.gguf 117 GiB Q8_0 uniform All tensors Q8_0
laguna-s-2.1-q5_k_m-uncensored.gguf 79 GiB Q5_K_M uniform All tensors Q5_K_M
laguna-s-2.1-q4_k_m-uncensored.gguf 67 GiB Q4_K_M uniform All tensors Q4_K_M
laguna-s-2.1-uncensored-hq.gguf 77 GiB Mixed Experts+Shared+Attn:Q5_K, Emb/Out/Routing:Q8_0
laguna-s-2.1-uncensored-balanced.gguf 63 GiB Mixed Experts:Q4_K, Attn+Shared:Q5_K, Emb/Out/Routing:Q8_0
laguna-s-2.1-uncensored-compact.gguf 49 GiB Mixed Experts:Q3_K, Attn+Shared+Dense:Q5_K, Emb/Out/Routing:Q8_0
laguna-s-2.1-uncensored-ultracompact.gguf 49 GiB Mixed Experts:Q3_K, Attn:Q4_K, Shared+Dense:Q3_K, Emb/Out:Q4_K

Mixed-precision recipes use per-tensor --tensor-type overrides for 814 tensors via --tensor-type-file. The default fallback type for unlisted tensors is the top-level quant type passed to llama-quantize.

Usage

llama-cli --model laguna-s-2.1-uncensored-balanced.gguf -ngl 0 -p "Hello" -n 128

For GPU offloading:

llama-cli --model laguna-s-2.1-uncensored-balanced.gguf -ngl 99 -c 32768

Model Details

Architecture Laguna MoE, 48 layers, 256 routed experts (top-10 sigmoid) + 1 shared expert
Parameters 118B total, ~8B activated per token
Context length 1,048,576 tokens (YaRN, rope base 500000 / 10000 SWA)
Vocab size 100,352
Languages English, German
License OpenMDW-1.1

Original Evaluation (ressl/Laguna-S-2.1-Uncensored)

Metric Base Uncensored
English refusals (686 prompts) 92.71% 2.33%
German refusals (686 prompts) 74.49% 4.23%
XSTest over-refusal (214 prompts) 8.88% 1.87%
HumanEval pass@1 (164 problems) 90.24% 85.37%

Full evaluation details, datasets, and methodology are in the source repository.

Recipe Selection Guide

Use case Recommended variant
Maximum quality HQ or Q8_0
Best size/quality tradeoff Balanced
Lower VRAM, similar quality to balanced Compact
Minimal disk footprint Ultra Compact

Credits

Original uncensoring: Robert Ressl (ressl.ch)

GGUF conversion and quantization: converted with oscar convert_hf_to_gguf.py (commit 4310aa4f8) and llama-quantize from the same build, using the registered LagunaForCausalLM conversion plugin.


GGUF collection generated August 2026.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04Update README: corrected compact variant, recipe selection guided5dd1d83.7 KB
    Loading...
  2. 2026-08-04Add GGUF variants README03b77573.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration