← back to catalog · registered 2026-09-28 01:57

bumblebuttpow/Spark-X2.5-4B-abliterated-MLX-8bit

bumblebuttpow 4B GGUF
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bumblebuttpow%2FSpark-X2.5-4B-abliterated-MLX-8bit"
Response includes
  • classification m8
  • files 10
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
mlx safetensors spark2_5 quantized 8-bit abliterated uncensored text-generation conversational custom_code en zh

Related

Total size
4.07 GB
Files
10
Quantizations
1
Registered
2026-09-28 01:57
Last updated on HF
2026-09-28 01:16

Files by quantization

Auxiliary files 10 files 4.08 GB
model.safetensors 4.07 GB 6141943f download
tokenizer.json 9.65 MB 2591c719 download
model.safetensors.index.json 45.4 KB 1cc7b858 download
chat_template.jinja 3.74 KB fa06c3a2 download
README.md 3.50 KB 2511c4b4 download
config.json 2.45 KB aad81e69 download
.gitattributes 1.48 KB a6344aac download
SHA256SUMS.txt 424 B 3a960ab5 download
tokenizer_config.json 420 B fcc86fb2 download
generation_config.json 281 B d47d1131 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • XHToken/Spark-X2.5-4B
  • SC117/Spark-X2.5-4B-abliterated-FIT-GGUF
    library_name: mlx
    pipeline_tag: text-generation
    tags:
  • mlx
  • spark2_5
  • quantized
  • 8-bit
  • abliterated
  • uncensored
    language:
  • en
  • zh

Spark-X2.5-4B-abliterated-MLX-8bit

8-bit MLX quantization of SC117's abliterated Spark X2.5 4B tune, for Apple silicon. 4.1 GB on disk, ~4.5 GB peak memory at inference.

How it was made

The abliterated tune ships as GGUF only, so this conversion round-trips it back to full precision and then quantizes for MLX:

  1. Spark-X2.5-4B-abliterated-BF16.gguf, sha256-verified against SC117's SHA256SUMS.txt, mapped back to a Hugging Face layout checkpoint. All 290 tensors matched by name and shape against the original XHToken shard headers. llama.cpp's conversion/spark2_5.py confirms the fused QKV and gate tensors are straight copies with no permutation, so nothing was reordered.
  2. Converted and quantized with spark-mlx-convert from Spark-MLX-LLM (MLX LM's converter with the Spark2_5 architecture registered): affine q8, group size 64, remaining parameters in float16.

Validation. Greedy chat-templated decoding is token-for-token identical between the source BF16 GGUF (llama.cpp on CPU) and this quantized model (MLX on Metal) across the full overlap, and identical to the intermediate 16-bit conversion. Effective 8.503 bits per weight including scales and biases.

Quantization layout

tensors count stored as
projections + embedding 181 packed q8 (uint32) + fp16 scales/biases, group 64
layernorms + model.norm 73 fp16
per-head self_attn.g_proj sigmoid gates 36 fp16

The head-wise attention output gates stay in fp16 on purpose. They scale each attention head through a sigmoid and are sensitive to low-bit quantization; keeping them at 16 bits preserves tool-calling reliability.

Running it

Spark2_5 is not in stock mlx-lm yet. Use the Spark-MLX-LLM extension, which registers the architecture without modifying the installed MLX LM:

spark-mlx-server --model ./Spark-X2.5-4B-abliterated-MLX-8bit --host 127.0.0.1 --port 8080
spark-mlx-chat --model ./Spark-X2.5-4B-abliterated-MLX-8bit
spark-mlx-generate --model ./Spark-X2.5-4B-abliterated-MLX-8bit -p "你好" -m 128

If an official MLX LM release ships a native Spark2_5 module, the wrappers pick it up automatically.

Measured on an M4 Pro (24 GB): ~54 tokens/s generation, ~4.5 GB peak memory.

Model details

Architecture unchanged from the base: 4B dense, 36 layers, hybrid 3:1 sliding-window(512):full attention, 1M native context, 16 attention heads / 4 KV heads, head_dim 256, GELU MLP, tied embeddings, chat template with thinking enabled by default. Sampler card: temp 1.0, top_p 0.95.

Weights are the abliterated (refusal-direction-suppressed) tune. Expect less refusal behavior than the instruct model; you are responsible for how you use it.

Files

SHA256SUMS.txt covers every model file in this repo.

Credits

XHToken for Spark-X2.5-4B and the MLX loader. SC117 for the abliteration and the FIT GGUF release. Apple's MLX and the MLX LM project.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.