library_name: gguf
license: other
license_name: openmdw-1.1
license_link: https://openmdw.ai/
pipeline_tag: text-generation
tags:
- laguna
- moe
- uncensored
- abliterix
- apex
- quantization
- gguf
- poolside
base_model: - SC117/Laguna-S-2.1-Uncensored
- poolside/Laguna-S-2.1
base_model_relation: quantized
Laguna-S-2.1-Uncensored-APEX-GGUF
English | 📖 中文文档
Uncensored 118B-A8B MoE · abliterix Trial 16 · APEX I-tier GGUFs + imatrix
Laguna S 2.1 is a poolside ~118B total / ~8B active per token Mixture-of-Experts model for agentic coding and long-horizon work. Architecture highlights: token-choice routing with softplus gates, 256 routed experts + 1 shared expert (top-10), GQA, 1:3 global/sliding-window attention (48 layers: 12 global + 36 local, window 512), and up to about 1M context, with optional native thinking interleaved with tool use.
This release builds on the official weights in two steps:
- Uncensored behavior edit via abliterix (ROCm + bitsandbytes 4-bit search path), selecting Trial 16 LoRA, stream-merged back to BF16, plus MoE safety-expert router bake-in.
- APEX mixed-precision GGUF quantized from the BF16 GGUF of our Laguna-S-2.1-Uncensored (converted with poolside llama.cpp (laguna)), using APEX-style tensor-type configs + imatrix (token embedding / output kept at BF16).
License: OpenMDW-1.1 (same family as the base model). See openmdw.ai and poolside terms.
After merging abliterix Trial 16, this model shows a much lower refusal rate and can differ substantially from official Laguna-S-2.1. Evaluate compliance and safety for your use case; control access and audit as needed.
| Refusals (harmful eval) | 8 / 100 (baseline ~97 / 100) |
| KL divergence | ~0.0042 |
| Length deviation | ~0.07 σ |
| Generation health | PASSED |
| Selected trial | abliterix Trial 16 |
Implementation sketch: LoRA merge W += (B @ A) * (alpha / r) (this trial alpha = r = 1); per-layer top ~20 safety experts with router row scale ~0.74 (see merge metadata).
The following can also appear with official Laguna-S-2.1 (including high-precision online deployments). They are not introduced by this Uncensored edit and should not be blamed on abliterix / the refusal merge:
- Model identity quirks or self-descriptions that diverge from the official persona;
- Odd tokens, typos, stiffness, or occasional awkward phrasing in Chinese and other languages.
This Uncensored pipeline mainly changes refusal / safety-related behavior. More aggressive quant tiers (e.g. Compact / Mini) may amplify existing text noise, but the underlying issues already exist in the original model and/or the local quant+inference stack. For a fair check, compare official vs this release under the same serving settings.
| Architecture | Laguna MoE (poolside laguna) |
| Parameters | ~118B total, ~8B active / token |
| Layers | 48 (L0 dense FFN, L1–L47 MoE) |
| Experts | 256 routed + 1 shared, top-10 |
| Attention | GQA, 8 KV heads, head dim 128 |
| Context | Up to ~1,048,576 tokens (practical limit depends on VRAM / -c) |
| Vocab | 100,352 (Laguna family tokenizer) |
| Modality | Text → text (no mmproj in this package) |
| This repo | APEX I-Quality / I-Balanced / I-Compact / I-Mini GGUF |
Coding / tool ability is largely retained; refusal and alignment behavior are changed. Full official bench tables were not re-run for this derivative.
These files use APEX-style MoE-aware mixed precision: precision follows tensor role + layer position (higher on edges, more aggressive in the middle), with imatrix for the I- tiers.
Common settings for this package: source Laguna-S-2.1-Uncensored BF16 GGUF; calibration laguna-s-2.1.imatrix; token embedding / output = BF16; routers and norms largely left at high precision defaults. Configs target Laguna 48-layer naming (ffn_*_exps / ffn_*_shexp / attn_*, L0 dense).
| File | Size | Mid experts | Best for |
|---|---|---|---|
*-I-Quality.gguf | ~70 GB | edge Q6_K / near Q5_K / mid iq4_xs; shared Q8_0; attn Q6_K | Smaller high-quality try (IQ mid-layers) |
*-I-Balanced.gguf | ~80 GB | edge Q6_K / near & mid Q5_K; shared Q8_0; attn Q6_K | Recommended default for Chinese users — steadier than Compact / Mini |
*-I-Compact.gguf | ~52 GB | edge Q4_K / mid Q3_K; shared Q6_K; attn Q4_K | Tighter memory; more quant noise OK |
*-I-Mini.gguf | ~41 GB | edge Q3_K / near Q3_K / mid iq2_s; shared Q5_K→Q4_K; attn Q4_K / Q3_K | Smallest I-tier here; fits ~48–64 GB unified / VRAM budgets; highest quant noise |
How to choose:
- Chinese-heavy use: prefer I-Balanced. Mid experts stay Q5_K (not IQ), which usually feels more stable. Drop to I-Compact if memory is tight; use I-Mini only when you need the smallest footprint and accept more noise.
- English / coding: tier differences are usually smaller; pick by VRAM/speed (Balanced → Compact → Mini). I-Quality is optional when you want IQ mid-layers and a smaller footprint than Balanced.
- I-Mini: ~41 GB, mid experts at iq2_s (requires imatrix). Good for 128 GB unified-memory boxes with room for long context, or machines that cannot hold Compact. Expect more quant artifacts than Compact.
Use a build that understands the laguna architecture (poolside laguna branch or equivalent).
Example (I-Balanced recommended)
./llama-server \ -m ./Laguna-S-2.1-Uncensored-APEX-I-Balanced.gguf \ --port 8080 \ -sm none \ --device rocm0 \ --ctx-size 131072 \ --flash-attn on \ --no-mmap \ --fit on \ --jinja \ --host 0.0.0.0
- Must recognize
general.architecture = laguna. - Thinking / tools: configure per client/server docs (e.g.
enable_thinking, reasoning options). - Warnings like
special_eos_id is not in special_eog_idsare tokenizer metadata hints; the model can still load. If stop/truncation is odd, check stop / max_tokens / template. - Text-only GGUF; no mmproj.
| General / coding | temperature 0.6–1.0, top_p 0.95, top_k 20 |
| More stable | Slightly lower temperature |
- Base:
poolside/Laguna-S-2.1 - abliterix search → Trial 16 LoRA + MoE router adjustments
- BF16 stream-merge →
Laguna-S-2.1-Uncensored - poolside llama.cpp → BF16 GGUF + imatrix
- APEX tensor-type-file + imatrix → I-Quality / I-Balanced / I-Compact / I-Mini
Links
- Original model: https://huggingface.co/poolside/Laguna-S-2.1
- Announcement: https://poolside.ai/blog/introducing-laguna-s-2-1
- OpenRouter: https://openrouter.ai/poolside/laguna-s-2.1
- Official GGUF: https://huggingface.co/poolside/Laguna-S-2.1-GGUF
- poolside llama.cpp (laguna): https://github.com/poolsideai/llama.cpp/tree/laguna
- abliterix: https://github.com/wuwangzhang1216/abliterix
- APEX: https://github.com/mudler/apex-quant
- License: https://openmdw.ai/
Disclaimer
Community derivative (behavior edit + quantization). Not an official poolside release. Use at your own risk; follow local law and upstream licenses.