library_name: mlx
license: openmdw-1.1
license_link: https://huggingface.co/AmixDigital/Laguna-S-2.1-Uncensored-oQ4e/blob/main/LICENSE
pipeline_tag: text-generation
base_model: SC117/Laguna-S-2.1-Uncensored
base_model_relation: quantized
tags:
- mlx
- oq
- quantized
- moe
- laguna
- uncensored
- apple-silicon
- macos
Laguna-S-2.1-Uncensored-oQ4e
An independent mixed-precision MLX edition for local software development and agentic coding on Apple Silicon.
About This Release
Laguna-S-2.1-Uncensored-oQ4e is an independent MLX oQ level 4 enhanced quantization of SC117/Laguna-S-2.1-Uncensored. The original ancestor is poolside/Laguna-S-2.1.
The release targets local software-development and agentic-coding workloads on Apple Silicon while retaining the direct source's uncensored behavior and native thinking configuration.
- Source behavior: inherited from the SC117 derivative; no new behavior edit was performed by AmixDigital.
- Published artifact: 13 MLX Safetensors shards produced through the oMLX graphical interface.
- Optional acceleration: DFlash is paired at runtime and is not merged into these weights.
Important Notices
- This is a community quantization, not an official release by Poolside or SC117.
- No independent benchmark suite or long-context evaluation has been completed for this release.
- Only the documented 128 GB test environment has been validated; minimum memory is unmeasured.
- No DFlash draft weights are included.
Upstream results are not quantization results. Refusal and divergence measurements reported by SC117 were not rerun here and must not be attributed to these MLX artifacts.
Uncensored Behavior and Safety
The direct source reduces refusal behavior through an upstream Abliterix-based modification. It may generate harmful, explicit, biased, illegal, misleading, or otherwise unsafe content. This release is not suitable for all audiences.
Model Details
| Architecture | Laguna Mixture-of-Experts, text-to-text |
|---|---|
| Parameters | 118B total; approximately 8B active per token upstream |
| Layers | 48 |
| Experts | 256 routed + 1 shared, top-10 routing |
| Context | Configured up to 1,048,576 tokens; not long-context tested here |
| Quantization | oQ level 4 enhanced, MLX affine mixed precision |
| Weights | 63.017 GiB · 13 shards |
| State | Quantized target weights; optional draft distributed separately |
Repository Files
The table groups related files while covering every functional artifact in this repository.
| File | Purpose | Required by | Notes |
|---|---|---|---|
model-00001-of-00013.safetensors … model-00013-of-00013.safetensors | Quantized target weights | MLX runtime | 13 shards · 63.017 GiB |
model.safetensors.index.json | Tensor-to-shard index | MLX loader · Hugging Face | Includes verified conceptual parameter metadata |
config.json | Architecture and quantization configuration | MLX · oMLX | Laguna custom code · affine mixed precision |
configuration_laguna.py · modeling_laguna.py | Custom Laguna implementation | Custom-code loaders | Carry Apache-2.0 notices |
tokenizer.json · tokenizer_config.json · special_tokens_map.json | Tokenizer data and settings | Prompt encoding and decoding | Inherited from the direct source |
chat_template.jinja | Chat, thinking, and tool-message formatting | Conversational inference | Native thinking remains enabled by default |
generation_config.json | Generation defaults and companion reference | oMLX runtime | DFlash remains external and optional |
oq_imatrix_report.json | Machine-readable calibration evidence | Quantization audit | Personal cache path redacted; limitations retained |
LICENSE · LICENSE-APACHE-2.0 | Redistribution terms | Recipients and redistributors | OpenMDW-1.1 materials · Apache-2.0-noticed code |
README.md · .gitattributes | Model card and Hub storage rules | Hugging Face Hub | English Lumen card · LFS tracking |
Provenance and Lineage
- Published artifact: this repository's MLX oQ4e quantized target weights, prepared by AmixDigital.
- Direct source: SC117/Laguna-S-2.1-Uncensored at revision 0e9c665.
- Original ancestor: poolside/Laguna-S-2.1 at revision 00af5a5.
The optional DFlash draft is a runtime companion—not an ancestor, source checkpoint, or component of the published weights.
Transformation Details
Upstream modification. SC117 describes the direct source as an Abliterix Trial 16 behavior edit of Poolside's BF16 model, including a LoRA merge, MoE router adjustments, and restoration of the default thinking template. This work was not performed by AmixDigital.
MLX quantization. The published artifacts were created through the graphical interface of oMLX 0.5.4. No CLI quantization command was used.
| Quantization | MLX affine mixed precision; 4-bit group-size-64 default with 5-, 6-, and 8-bit overrides |
|---|---|
| Effective storage | Approximately 4.605 bits per conceptual parameter, calculated from 67,664,427,590 weight bytes and 117,561,977,600 parameters |
| Stored dtypes | BF16 tensor values and U32 packed quantization data in Safetensors |
| Precision overrides | 386 configured tensor overrides: 105 at 5-bit, 1 at 6-bit, and 280 at 8-bit; 245 use group size 64 and 141 use group size 128 |
| Highest-precision tensors | lm_head and model.embed_tokens are configured at 8-bit, group size 64 |
| Calibration preset | oqe_code_multilingual |
| Collection | 1,024 processed samples · sequence length 512 |
| Expert coverage | 36,084 / 36,096 routes activated · 12 zero-count experts |
| Report status | coverage_sufficient: false · collection_sufficient: false |
Calibration limitation: the imatrix did not activate every expert route. The maintainer reports that oMLX used a quantized proxy because the BF16 source did not fit in available unified memory.
Requirements and Compatibility
| macOS | 26.5.2 · tested |
|---|---|
| Apple Silicon | MacBook Pro, M4 Max, 16-core CPU · tested |
| Unified memory | 128 GB · tested and recommended |
| Observed memory | Approximately 70 GB or less · maintainer observation, not instrumented |
| oMLX | 0.5.4 · tested |
mlx-lm | 0.31.3 · installed in test environment |
mlx-vlm | 0.6.3 · installed in test environment |
| MLX core | Exact version not recorded |
Minimum unified memory has not been measured. Disk size alone is not a safe minimum-memory estimate.
Usage
Verified oMLX workflow
- Add the obtained model directory to oMLX 0.5.4.
- Load the model on a compatible Apple Silicon Mac.
- Keep thinking enabled when the task benefits from reasoning.
- Run a short software-development prompt before increasing context or enabling optional acceleration.
The maintainer's smoke test requested a landing-page implementation. The model loaded successfully, produced a coherent result, and exposed thinking behavior.
Prompting, Thinking and Tool Use
The chat template defaults enable_thinking to true. Thinking was observed, but tool calling was not tested on this quantization.
Acceleration and Companion Models
This target can be paired with the separately distributed poolside/Laguna-S-2.1-DFlash draft. DFlash proposes token blocks that the target verifies; it does not replace or modify the target weights.
Open the DFlash companion repository →
- Obtain the official draft repository separately.
- Enable DFlash in the target's oMLX settings.
- Select the draft directory and reload the target.
Runtime pairing was completed in oMLX 0.5.4. Throughput, acceptance rate, equivalence, and long-context behavior were not measured; no acceleration factor is claimed.
Evaluation and Performance
No independent benchmark suite has been completed for this release.
| Check | Result | Limit |
|---|---|---|
| Model load | Successful | One Mac configuration |
| Generation | Coherent landing page | Single qualitative prompt |
| Thinking | Observed | No structured evaluation |
| Tool calling | Not tested | Inherited configuration only |
| Long context | Not tested | Configured maximum unvalidated |
| DFlash | Runtime pairing completed | Speed and acceptance unmeasured |
Benchmarks from Poolside or SC117 describe their own artifacts and must not be interpreted as measurements of this quantization.
Intended Uses
- Local software development and code-oriented assistance on compatible Apple Silicon Macs.
- Agentic-coding experiments with human review and constrained tool permissions.
- General text generation where unvalidated general-purpose quality is understood.
Out-of-Scope Uses
- Unsupervised high-stakes medical, legal, financial, employment, safety, or infrastructure decisions.
- Autonomous destructive tool execution without sandboxing, review, and recovery controls.
- Uses that violate law, rights, license terms, or platform policies.
- Claims of Windows, Linux, CUDA, GGUF, vLLM, or server-GPU compatibility.
Risks, Biases and Limitations
- Uncensored output may be harmful, explicit, biased, fabricated, or unsafe.
- Quantization may add variance or degradation beyond upstream limitations.
- Twelve of 36,096 expert routes were not activated during calibration.
- Quality, tool use, long context, minimum memory, and general stability remain unmeasured.
- Custom model code requires security review before remote-code execution.
Recommendations and Mitigations
- Evaluate the exact model and prompts before deployment.
- Keep human review for code changes, tool use, and consequential outputs.
- Sandbox generated code and restrict credentials and filesystem access.
- Apply access controls and content filtering where necessary.
- Start with short contexts and monitor memory pressure on untested hardware.
License and Responsible Use
The published Model Materials are distributed under the OpenMDW-1.1 License. The complete inherited agreement is included as LICENSE. Redistribution must retain the agreement and applicable copyright and origin notices.
The redistributed Python implementation files also carry Apache-2.0 notices. A complete copy is included as LICENSE-APACHE-2.0.
Users remain responsible for third-party rights, applicable law, and determining whether the model and its outputs are appropriate for a particular use. AmixDigital and the maintainers are not responsible for illegal, abusive, or otherwise unauthorized uses of the model or its outputs.
Acknowledgements
- Poolside for the original Laguna S 2.1 and official DFlash companion.
- SC117 for the direct uncensored source and behavior-edit provenance.
- Abliterix contributors for the upstream behavior-edit method.
- oMLX contributors for MLX serving and oQ quantization tooling.
Model Card Contact
Contact AmixDigital on Hugging Face or open a discussion in this model repository to report an error or propose a correction.