license: apache-2.0
base_model:
- XHToken/Spark-X2.5-4B
- SC117/Spark-X2.5-4B-abliterated-FIT-GGUF
library_name: mlx
pipeline_tag: text-generation
tags: - mlx
- spark2_5
- quantized
- 8-bit
- abliterated
- uncensored
language: - en
- zh
Spark-X2.5-4B-abliterated-MLX-8bit
8-bit MLX quantization of SC117's abliterated Spark X2.5 4B tune, for Apple silicon. 4.1 GB on disk, ~4.5 GB peak memory at inference.
- Base model: XHToken/Spark-X2.5-4B (Apache-2.0)
- Abliteration: SC117/Spark-X2.5-4B-abliterated-FIT-GGUF (converted from its full-precision BF16 tier)
- Loader: XHToken/Spark-MLX-LLM (registers the
spark2_5architecture for MLX LM)
How it was made
The abliterated tune ships as GGUF only, so this conversion round-trips it back to full precision and then quantizes for MLX:
Spark-X2.5-4B-abliterated-BF16.gguf, sha256-verified against SC117'sSHA256SUMS.txt, mapped back to a Hugging Face layout checkpoint. All 290 tensors matched by name and shape against the original XHToken shard headers. llama.cpp'sconversion/spark2_5.pyconfirms the fused QKV and gate tensors are straight copies with no permutation, so nothing was reordered.- Converted and quantized with
spark-mlx-convertfrom Spark-MLX-LLM (MLX LM's converter with the Spark2_5 architecture registered): affine q8, group size 64, remaining parameters in float16.
Validation. Greedy chat-templated decoding is token-for-token identical between the source BF16 GGUF (llama.cpp on CPU) and this quantized model (MLX on Metal) across the full overlap, and identical to the intermediate 16-bit conversion. Effective 8.503 bits per weight including scales and biases.
Quantization layout
| tensors | count | stored as |
|---|---|---|
| projections + embedding | 181 | packed q8 (uint32) + fp16 scales/biases, group 64 |
layernorms + model.norm |
73 | fp16 |
per-head self_attn.g_proj sigmoid gates |
36 | fp16 |
The head-wise attention output gates stay in fp16 on purpose. They scale each attention head through a sigmoid and are sensitive to low-bit quantization; keeping them at 16 bits preserves tool-calling reliability.
Running it
Spark2_5 is not in stock mlx-lm yet. Use the Spark-MLX-LLM extension, which registers the architecture without modifying the installed MLX LM:
spark-mlx-server --model ./Spark-X2.5-4B-abliterated-MLX-8bit --host 127.0.0.1 --port 8080
spark-mlx-chat --model ./Spark-X2.5-4B-abliterated-MLX-8bit
spark-mlx-generate --model ./Spark-X2.5-4B-abliterated-MLX-8bit -p "你好" -m 128
If an official MLX LM release ships a native Spark2_5 module, the wrappers pick it up automatically.
Measured on an M4 Pro (24 GB): ~54 tokens/s generation, ~4.5 GB peak memory.
Model details
Architecture unchanged from the base: 4B dense, 36 layers, hybrid 3:1 sliding-window(512):full attention, 1M native context, 16 attention heads / 4 KV heads, head_dim 256, GELU MLP, tied embeddings, chat template with thinking enabled by default. Sampler card: temp 1.0, top_p 0.95.
Weights are the abliterated (refusal-direction-suppressed) tune. Expect less refusal behavior than the instruct model; you are responsible for how you use it.
Files
SHA256SUMS.txt covers every model file in this repo.
Credits
XHToken for Spark-X2.5-4B and the MLX loader. SC117 for the abliteration and the FIT GGUF release. Apple's MLX and the MLX LM project.