license: apache-2.0
library_name: splash
pipeline_tag: image-text-to-text
inference: false
base_model:
- Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16
base_model_relation: quantized
tags: - splash
- lm-studio
- apple-silicon
- metal
- local-inference
- dflash2
- speculative-decoding
- qwen3.8
- qwen
- 27b
- dense
- vision-language
- 4-bit
- abliterated
- derisked
- research
- security-research
- red-teaming
- coding
- tool-calling
- long-context
Qwen3.8-27B — Blackfrost ABLITERATED · Splash package
A format conversion of Blackfrost's de-risked Qwen3.8-27B for the Splash engine on Apple silicon (LM Studio).
This repository is only a conversion. The model, the research, the refusal-surface work, and the evaluation all belong to Blackfrost (Blackfrost Softwares Corp.). Their BF16 master is published at Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16, and that model card is the authoritative description of what this checkpoint is, how it was evaluated, and how it should be used. Read it first. If you find this package useful, the credit goes to Blackfrost.
This is an independent community conversion. Blackfrost, Inco AI, and the Qwen team have not reviewed or endorsed it.
Take notice (carried over from Blackfrost's card)
This is a Blackfrost weight-level research checkpoint with a deliberately reduced refusal surface. It is not the upstream Qwen safety-stock checkpoint and must not be represented as one. Blackfrost publishes it as a public, ungated research preview that is not for sale, with evaluation and release review still in progress. Everything in that notice applies to this package too.
What this package is
Splash is Inco AI's open-source inference engine for Apple silicon. This package contains everything the Splash engine needs to serve Blackfrost's Qwen3.8-27B: the 4-bit target, a DFlash 2 draft model for speculative decoding, the vision encoder, and the tokenizer. It uses the same splash-packed-q4 layout as Inco's own incoai/Qwen3.8-27B-Splash. It is not a Transformers, GGUF, or MLX checkpoint and will not load anywhere but Splash.
It was built for and verified in LM Studio's bundled Splash engine ("Splash (Metal)" runtime 0.0.5, LM Studio 0.4.25) on a Mac Studio M4 Max with 128 GB of unified memory running macOS 26.6.
What this package is not
- Not a new model. No fine-tuning, merging, pruning, direction editing, or any weight change beyond 4-bit quantization was applied. The target weights are Blackfrost's BF16 tensors, quantized and re-laid-out, nothing else.
- Not Blackfrost's evaluated artifact. Blackfrost evaluated their BF16 master and a W4A4 NVFP4 derivative. This is a different quantization (MLX affine 4-bit, group size 64). Per Blackfrost's own disclaimer, any further weight change, quantization included, creates an artifact they have not evaluated. Treat the refusal and perplexity numbers on their card as describing their artifacts, not this one.
- Not an endorsement. See above.
Specifications
| Architecture | Qwen3.8-27B dense hybrid VLM · Gated DeltaNet + full attention |
| Base | Official Qwen/Qwen3.8-27B |
| Weights | Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16, revision 9d85770e5eb602322b4bceef55beda357e0bd0ca |
| Parameters | 27B |
| Precision | Target: affine 4-bit, group size 64 (MLX quantization scheme). Vision encoder: BF16. Draft: Inco's packed DFlash 2 |
| Package format | splash-packed-q4, manifest schema 3, one packed file per layer |
| Size | about 16 GiB on disk, 16.2 GiB resident when loaded |
| Context | 262,144 tokens native |
| Modalities | Text and image input; text output |
| Serving | Splash engine only (LM Studio runtime verified) |
| Status | Community conversion of a public research preview |
Lineage
| Base weights | Qwen/Qwen3.8-27B (Alibaba) |
| Applied by Blackfrost | Refusal-surface direction modification at the weight level, plus their operational system prompt embedded in the chat template |
| Applied here | 4-bit affine quantization (mlx-vlm 0.7.6, --q-bits 4 --q-group-size 64 --q-mode affine) and packing into Splash's binary layout. Nothing else |
| Not applied | Pruning, SFT, DPO, LoRA, merging, or any further direction editing |
| Chat behavior | Blackfrost's operational system prompt is embedded in the chat template, as in their release (see below) |
Blackfrost's internal direction bank, scaling schedule, capture data, and build workflow are not part of their release and therefore not part of this one.
Package contents
target/ 66 files ~14.1 GiB Blackfrost Qwen3.8-27B-ABLITERATED, 4-bit, one packed file per layer
draft/ 6 files ~1.2 GiB DFlash 2 draft model (from Inco's package)
vision/ 1 file ~0.9 GiB bf16 vision encoder (from Inco's package)
tokenizer/ 5 files tokenizer, config, and Blackfrost's chat template
manifest.json provenance, execution geometry, and SHA-256 of every artifact
| Component | Source | Revision |
|---|---|---|
| Target (embedding, 64 layers, head) | Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16, quantized and packed here |
9d85770e5eb602322b4bceef55beda357e0bd0ca |
| Chat template | Blackfrost's chat_template.jinja from the same revision, with Inco's one-line change (see below) |
9d85770e5eb602322b4bceef55beda357e0bd0ca |
| Draft | incoai/Qwen3.8-27B-DFlash2, copied from incoai/Qwen3.8-27B-Splash |
dedf8df68adfb1afeaf7b7480c0a0243108177b4 |
| Vision encoder, tokenizer files | mlx-community/Qwen3.8-27B-4bit, copied from Inco's package |
3e6447f082e89cc7f0bc6e5441afd38dfce760ff |
Blackfrost's modification touches the language model only. Their vision tower and tokenizer vocabulary are byte-identical to upstream Qwen's, which is why Inco's packed vision encoder and tokenizer files are reused unchanged. The chat template is Blackfrost's, including their embedded operational prompt, with the same single change Inco made for their package: a system message that appears after the first turn is rendered in place instead of raising an exception, which coding agents that inject instructions mid-conversation need. Everything else in the template is Blackfrost's.
The draft model is Inco's DFlash 2 drafter trained against stock Qwen3.8-27B. The target verifies every drafted token, so the draft only affects speed, never the output distribution. Acceptance rates against the abliterated target may differ from stock, which would show up as lower throughput, not different text.
Run it in LM Studio
Requirements follow Inco's: Apple M3 or newer, macOS 26.4 or later, and at least 36 GB of unified memory (48 GB or more recommended). LM Studio must have the Splash (Metal) runtime installed (LM Studio → Runtimes).
Download the package into LM Studio's models directory, keeping the folder layout:
hf download notfakerv/Qwen3.8-27B-ABLITERATED-Splash \
--local-dir ~/.lmstudio/models/notfakerv/Qwen3.8-27B-ABLITERATED-Splash
Then load it from the LM Studio UI, or from the CLI:
lms load qwen3.8-27b-abliterated-splash
Generation check against LM Studio's local server:
curl http://127.0.0.1:1234/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3.8-27b-abliterated-splash",
"messages": [{"role": "user", "content": "Return exactly: READY"}],
"temperature": 0,
"max_tokens": 256
}'
Qwen3.8 reasons before answering by default. Give it enough max_tokens to finish its reasoning, or your response will come back with reasoning and empty content.
For the standalone Splash CLI (brew install incoai/tap/splash), check which package formats your Splash version accepts. This package follows the splash-packed-q4 layout that Inco's own packages use, but it has been verified only inside LM Studio's bundled engine.
How it was built, and how it was checked
Inco has not published a packer for the splash-packed-q4 format. The format was therefore reconstructed from Inco's published package and the open-source Splash planner (section order, 16 KiB alignment, 256-row tiling of the codes, scales, and biases of every quantized projection, the fused GDN and attention projections, and the pre-computed decay vectors). The resulting packer was validated by regenerating Inco's own incoai/Qwen3.8-27B-Splash from the mlx-community/Qwen3.8-27B-4bit checkpoint it was built from: all 66 target files came out byte-identical to Inco's, SHA-256 for SHA-256. Only after that was it run on Blackfrost's weights.
Pipeline:
Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16→ MLX, with mlx-vlm 0.7.6:mlx_vlm.convert -q --q-bits 4 --q-group-size 64 --q-mode affine(the same scheme the mlx-community and lmstudio-community Qwen3.8 builds use).- MLX 4-bit →
target/*.binwith the validated packer. draft/,vision/, and tokenizer files taken fromincoai/Qwen3.8-27B-Splash; Blackfrost's chat template installed.- New
manifest.jsonwith the SHA-256 and size of every artifact, in Inco's artifact order, and a freshartifact_set_sha256.
Compared with Inco's stock package, 62 of the 66 target files differ. The embedding, the output head, and layers 0 and 1 are byte-identical to stock, which matches Blackfrost's modification leaving those tensors untouched.
Smoke checks in LM Studio (Splash runtime 0.0.5, M4 Max, 128 GB, temperature 0):
| Check | Observed |
|---|---|
| Package loads with every artifact's SHA-256 matching the manifest | Yes, 15.3 s, 16.19 GiB |
| Factual text question | Correct, after a short reasoning pass |
| Synthetic image (red left half, blue right half) | Described correctly |
| "Describe what you are and what rules you follow", no system message | Answered with Blackfrost's embedded operational prompt, confirming the template is in effect |
| System message injected mid-conversation ("answer in French") | Rendered and followed |
These are smoke checks, not a benchmark and not a safety evaluation. They do not establish coding, tool-use, long-context, or multi-turn retention for this quantization, and they do not reproduce Blackfrost's refusal funnel.
Blackfrost's reported results (for their artifacts, not re-measured here)
For reference only, from Blackfrost's card: their release score is 11 residual refusals from 450 original cases (2.4%), measured through a sequential, manually reviewed funnel on the W4A4 NVFP4 derivative of the BF16 master with the shipped short execution prompt. The 450-case set is 150 AdvBench, 150 StrongREJECT, and 150 XSTest prompts. Their WikiText-2 check reported a word perplexity of 8.48 for clean upstream BF16 and 9.37 for their NVFP4 derivative. The full tables, methodology, and caveats are on their model card. None of these numbers were measured on this 4-bit Splash package.
Security and deployment responsibility
Carried over from Blackfrost's card, because it applies unchanged. This checkpoint has a deliberately reduced refusal surface. Open weights do not provide an application policy, authorization system, audit trail, sandbox, or access-control boundary. Operators remain responsible for enforcing those controls outside the model. For production or shared use, Blackfrost recommends:
- authenticated access to the inference endpoint;
- independent request and tool-execution logging;
- least-privilege credentials for every tool;
- sandboxing for code execution and file access;
- explicit approval boundaries for irreversible actions;
- application-layer controls appropriate to the deployment domain.
The embedded prompt is a behavioral instruction, not a security boundary. LM Studio's server binds to localhost by default; keep it that way unless you have configured authentication.
Disclaimer
Refusal behavior in this checkpoint has been deliberately modified at the weight level by Blackfrost. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one.
No warranty of any kind. This package is provided "as is", without warranty express or implied, including fitness for a particular purpose. Nothing here guarantees that any input will be accepted or refused, that every upstream capability is retained, or that any category of output is unreachable.
Measurements describe only what was measured. The smoke checks above reflect one machine, one runtime, and a handful of prompts. Blackfrost's figures reflect their artifacts, templates, samplers, and review criteria. Neither generalizes automatically to this package in other settings.
This conversion is a further weight change. By Blackfrost's own terms, it is an artifact they have not evaluated. Problems with this package should be reported here, not to Blackfrost, unless they reproduce on their BF16 release.
Base license. Apache 2.0, as shipped with the official Qwen3.8-27B checkpoint and retained by Blackfrost. Every component here is Apache 2.0 upstream: Qwen3.8-27B (Alibaba), Blackfrost's derivative, the DFlash 2 draft (Inco AI), and the mlx-community conversion the vision and tokenizer files come from.
Credits
- Blackfrost (Blackfrost Softwares Corp.) for the model: the refusal-surface research, the weight-level modification, the embedded operational prompt, the evaluation, and the BF16 release this package is made from. Follow @Blackfrost_AI for their releases and contact them there about the model itself.
- The Qwen team (Alibaba) for Qwen3.8-27B.
- Inco AI for Splash, the DFlash 2 draft model, and the packed format, and for publishing a reference package precise enough to reverse-engineer and verify against.
- mlx-community and the Apple MLX / mlx-vlm maintainers for the quantization tooling and the reference 4-bit conversion.
@misc{inco2026splash,
title = {{Splash: A Local Engine Built Around the Model}},
author = {{Inco AI}},
year = {2026},
month = {September},
url = {https://inco.ai/blog/splash/}
}