← back to catalog · registered 2026-08-22 13:56

SC117/Laguna-S-2.1-Uncensored-APEX-GGUF

SC117 GGUF MoE second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FLaguna-S-2.1-Uncensored-APEX-GGUF"
Response includes
  • classification m-uncensored
  • files 8
  • hub_downloads_all_time 17,518
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
18K
5K last 30d - stable
Likes
23
Model age
2mo ago
created 2026-07-26
Downloads over time
Now19.5K→from4.9K↑297%
4.2K9.8K15.4K20.9K4.9K on Jul 2919.5K on Oct 11JulAugSepOct
Jul 29 → Oct 11 · 52 snapshots · spans 74 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Quantizations
BF16
Tags
gguf laguna moe uncensored abliterix apex quantization poolside text-generation base_model:SC117/Laguna-S-2.1-Uncensored base_model:quantized:SC117/Laguna-S-2.1-Uncensored license:other

Related

Total size
244 GB
Files
8
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-10-03 10:25

Files by quantization

BF16 1 file 2.08 GB
laguna-s-2.1-DFlash-BF16.gguf 2.08 GB 2ee8aa30 download
Auxiliary files 7 files 242 GB
Laguna-S-2.1-Uncensored-APEX-I-Balanced.gguf 79.4 GB def8d461 download
Laguna-S-2.1-Uncensored-APEX-I-Quality.gguf 69.5 GB d5a643ab download
Laguna-S-2.1-Uncensored-APEX-I-Compact.gguf 51.3 GB 70da561d download
Laguna-S-2.1-Uncensored-APEX-I-Mini.gguf 41.3 GB 0001c14a download
README_zh.md 21.9 KB 28cbd267 download
README.md 21.8 KB b096982a download
.gitattributes 1.86 KB 40b573d3 download

README current version from Hugging Face


library_name: gguf
license: other
license_name: openmdw-1.1
license_link: https://openmdw.ai/
pipeline_tag: text-generation
tags:

  • laguna
  • moe
  • uncensored
  • abliterix
  • apex
  • quantization
  • gguf
  • poolside
    base_model:
  • SC117/Laguna-S-2.1-Uncensored
  • poolside/Laguna-S-2.1
    base_model_relation: quantized

ABLITERIX TRIAL 16 APEX OpenMDW-1.1

Laguna-S-2.1-Uncensored-APEX-GGUF

English | 📖 中文文档

Uncensored 118B-A8B MoE · abliterix Trial 16 · APEX I-tier GGUFs + imatrix

🌊 About this release

Laguna S 2.1 is a poolside ~118B total / ~8B active per token Mixture-of-Experts model for agentic coding and long-horizon work. Architecture highlights: token-choice routing with softplus gates, 256 routed experts + 1 shared expert (top-10), GQA, 1:3 global/sliding-window attention (48 layers: 12 global + 36 local, window 512), and up to about 1M context, with optional native thinking interleaved with tool use.

This release builds on the official weights in two steps:

  1. Uncensored behavior edit via abliterix (ROCm + bitsandbytes 4-bit search path), selecting Trial 16 LoRA, stream-merged back to BF16, plus MoE safety-expert router bake-in.
  2. APEX mixed-precision GGUF quantized from the BF16 GGUF of our Laguna-S-2.1-Uncensored (converted with poolside llama.cpp (laguna)), using APEX-style tensor-type configs + imatrix (token embedding / output kept at BF16).

License: OpenMDW-1.1 (same family as the base model). See openmdw.ai and poolside terms.

⚠️ Uncensored notice

After merging abliterix Trial 16, this model shows a much lower refusal rate and can differ substantially from official Laguna-S-2.1. Evaluate compliance and safety for your use case; control access and audit as needed.

Refusals (harmful eval)8 / 100 (baseline ~97 / 100)
KL divergence~0.0042
Length deviation~0.07 σ
Generation healthPASSED
Selected trialabliterix Trial 16

Implementation sketch: LoRA merge W += (B @ A) * (alpha / r) (this trial alpha = r = 1); per-layer top ~20 safety experts with router row scale ~0.74 (see merge metadata).

📜 Disclaimer (identity & text quality)

The following can also appear with official Laguna-S-2.1 (including high-precision online deployments). They are not introduced by this Uncensored edit and should not be blamed on abliterix / the refusal merge:

  • Model identity quirks or self-descriptions that diverge from the official persona;
  • Odd tokens, typos, stiffness, or occasional awkward phrasing in Chinese and other languages.

This Uncensored pipeline mainly changes refusal / safety-related behavior. More aggressive quant tiers (e.g. Compact / Mini) may amplify existing text noise, but the underlying issues already exist in the original model and/or the local quant+inference stack. For a fair check, compare official vs this release under the same serving settings.

🧠 Model details
ArchitectureLaguna MoE (poolside laguna)
Parameters~118B total, ~8B active / token
Layers48 (L0 dense FFN, L1–L47 MoE)
Experts256 routed + 1 shared, top-10
AttentionGQA, 8 KV heads, head dim 128
ContextUp to ~1,048,576 tokens (practical limit depends on VRAM / -c)
Vocab100,352 (Laguna family tokenizer)
ModalityText → text (no mmproj in this package)
This repoAPEX I-Quality / I-Balanced / I-Compact / I-Mini GGUF

Coding / tool ability is largely retained; refusal and alignment behavior are changed. Full official bench tables were not re-run for this derivative.

💡 What is APEX?

These files use APEX-style MoE-aware mixed precision: precision follows tensor role + layer position (higher on edges, more aggressive in the middle), with imatrix for the I- tiers.

Common settings for this package: source Laguna-S-2.1-Uncensored BF16 GGUF; calibration laguna-s-2.1.imatrix; token embedding / output = BF16; routers and norms largely left at high precision defaults. Configs target Laguna 48-layer naming (ffn_*_exps / ffn_*_shexp / attn_*, L0 dense).

📦 APEX quantization tiers
File Size Mid experts Best for
*-I-Quality.gguf~70 GBedge Q6_K / near Q5_K / mid iq4_xs; shared Q8_0; attn Q6_KSmaller high-quality try (IQ mid-layers)
*-I-Balanced.gguf~80 GBedge Q6_K / near & mid Q5_K; shared Q8_0; attn Q6_KRecommended default for Chinese users — steadier than Compact / Mini
*-I-Compact.gguf~52 GBedge Q4_K / mid Q3_K; shared Q6_K; attn Q4_KTighter memory; more quant noise OK
*-I-Mini.gguf~41 GBedge Q3_K / near Q3_K / mid iq2_s; shared Q5_K→Q4_K; attn Q4_K / Q3_KSmallest I-tier here; fits ~48–64 GB unified / VRAM budgets; highest quant noise

How to choose:

  • Chinese-heavy use: prefer I-Balanced. Mid experts stay Q5_K (not IQ), which usually feels more stable. Drop to I-Compact if memory is tight; use I-Mini only when you need the smallest footprint and accept more noise.
  • English / coding: tier differences are usually smaller; pick by VRAM/speed (Balanced → Compact → Mini). I-Quality is optional when you want IQ mid-layers and a smaller footprint than Balanced.
  • I-Mini: ~41 GB, mid experts at iq2_s (requires imatrix). Good for 128 GB unified-memory boxes with room for long context, or machines that cannot hold Compact. Expect more quant artifacts than Compact.
🚀 Usage (llama.cpp)

Use a build that understands the laguna architecture (poolside laguna branch or equivalent).

Example (I-Balanced recommended)

./llama-server \
  -m ./Laguna-S-2.1-Uncensored-APEX-I-Balanced.gguf \
  --port 8080 \
  -sm none \
  --device rocm0 \
  --ctx-size 131072 \
  --flash-attn on \
  --no-mmap \
  --fit on \
  --jinja \
  --host 0.0.0.0
  • Must recognize general.architecture = laguna.
  • Thinking / tools: configure per client/server docs (e.g. enable_thinking, reasoning options).
  • Warnings like special_eos_id is not in special_eog_ids are tokenizer metadata hints; the model can still load. If stop/truncation is odd, check stop / max_tokens / template.
  • Text-only GGUF; no mmproj.
🎛️ Recommended sampling
General / codingtemperature 0.6–1.0, top_p 0.95, top_k 20
More stableSlightly lower temperature
🔧 Build pipeline (summary)
  1. Base: poolside/Laguna-S-2.1
  2. abliterix search → Trial 16 LoRA + MoE router adjustments
  3. BF16 stream-merge → Laguna-S-2.1-Uncensored
  4. poolside llama.cpp → BF16 GGUF + imatrix
  5. APEX tensor-type-file + imatrix → I-Quality / I-Balanced / I-Compact / I-Mini

Links

Disclaimer

Community derivative (behavior edit + quantization). Not an official poolside release. Use at your own risk; follow local law and upstream licenses.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-03Add optional Ko-fi support banner08e353722.1 KB
    Loading...
  2. 2026-08-03docs: add APEX-I-Mini to README92a316e21.8 KB
    Loading...
  3. 2026-08-02Upload README.md with huggingface_hub1584c7120.8 KB
    Loading...
  4. 2026-07-26Upload README.mdcb168c220.5 KB
    Loading...
  5. 2026-07-26Update README.md97692c520.6 KB
    Loading...
  6. 2026-07-26Upload README.mdbd8269620.5 KB
    Loading...

Discussions 5 threads

  1. 2026-08-02Laguna-S-2.1-Uncensored BF16 Safetensors Checkpoint please...open3 💬#5
    Loading...
  2. 2026-08-01The draft model doesn't work with Compactclosed3 💬#4
    Loading...
  3. 2026-07-28Thank you!open2 💬#3
    Loading...
  4. 2026-07-27Wishlist: Enable naming according to HuggingFace standardsopen2 💬#2
    Loading...
  5. 2026-07-26Laguna XS?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration