license: apache-2.0
base_model: nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2
base_model_relation: quantized
quantized_by: Honkware
library_name: exllamav3
pipeline_tag: text-generation
tags:
- exl3
- exllamav3
- quantized
quantization_format: exl3
inference: false
bits_per_weight: 4.0
Qwen3.8 · 27B · EfficientThink · Uncensored · K3 · Opus5 · Grok4.6 · GPT5.6Sol · SFT · SimPO · DFlash2
EXL3 · 4.0 bpw · 16.3 GB · Dense
[!NOTE]
An ExLlamaV3 build ofnerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2at 4.0 bits per weight. See Quants for sibling repos at other bit‑widths or browse the collection.
Quants
| BPW | Size | Status |
|---|---|---|
| 4.0 | 16.3 GB | this repo |
Inference
| Loader | Use it for |
|---|---|
| TabbyAPI | OpenAI‑compatible HTTP server. Drop‑in for OpenAI clients. |
| text‑generation‑webui | Local chat UI. Pick the ExLlamaV3 loader from the model dropdown. |
| ExLlamaV3 | Direct Python API for embedding the model in your own code or pipeline. |
Download
pip install -U huggingface_hub
hf download \
Honkware/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-exl3-4.0bpw \
--local-dir ./Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-exl3-4.0bpw
Quantization recipe (advanced, embedded in quantization_config.json)
| Setting | Value |
|---|---|
| Format | EXL3 |
| Bits per weight | 4.0 |
| Head bits | 6 |
| Calibration rows | 250 |
| Calibration data | exllamav3 bundled mix (c4, code, multilingual, technical, tiny, wiki) |
| Codebook | mul1 |
| Out‑scales | always |
| Parallel mode | enabled |
License & use
[!IMPORTANT]
Use and license follow the base model.
Quantization adds no additional restrictions. Refer to the upstream repository for terms, citation, and safety documentation.
{org}/{model}-exl3-{bpw}bpw