base_model: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5
base_model_relation: finetune
pipeline_tag: text-generation
language:
- en
license: gemma
quantized_by: triplezrobotics
tags: - gemma4
- code
- text-generation
- function-calling
- tool-use
- agentic
- abliterated
- uncensored
model-index: - name: Gemma-4 12B Coder — SFT v5 + abliterated (weights)
results:- task:
type: text-generation
name: Function calling (tool use)
dataset:
name: gemma4-coder-tool-eval
type: tpls/gemma4-coder-tool-eval
metrics:- type: pass_rate
value: 1.0
name: Tool-call pass rate (shim, prod path) - type: pass_rate
value: 0.125
name: Tool-call pass rate (raw llama.cpp --jinja)
- type: pass_rate
- task:
Gemma-4 12B Coder — SFT v5 + abliterated (weights)
Uncensored gemma-4 12B coder weights (safetensors) — for fine-tuning, merging, or quantizing.
Ready-to-serve GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-abliterated-GGUF.
⚠️ Tool-calling needs the recovery shim. The model emits gemma-4's native tool markup, which
llama.cpp --jinjaunder-parses — wrap your endpoint with the tool-shim (see Tool-calling below) to get standardtool_calls.
💡 Pick this for the best of both: SFT v5's tool-calling and an uncensored model — our KL-guarded abliteration applied on top of SFT v5 (no capability loss).
At a glance
| Type | Model weights (safetensors) |
| Techniques | sft-qlora → abliteration |
| Tool-calling | ✅ 100% gate pass (recovery-shim path) |
| Status | ✅ Active / supported |
| Use | GGUF quants: tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-abliterated-GGUF |
Use it — GGUF quantizations
Ready-to-serve GGUF quants live at tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-abliterated-GGUF
(llama.cpp / Ollama one-liners on that card). These are the safetensors weights, for
fine-tuning / merging / quantizing.
Tool-calling
Tool-calling works — but llama.cpp --jinja doesn't recognise gemma-4's native
tool-call markup, so the bare parser under-reports calls. The model is fine; the
parser is blind to the format. Recover standard tool_calls with a small serve-side
post-processor (no weight change, no latency beyond a regex scan).
Ready-to-use → tpls/gemma4-tool-shim — a drop-in
callback for OpenAI-compatible proxies, a standalone (dependency-free) example, and the pure
parser, all Apache-2.0, with the full recovery algorithm documented. Point your
OpenAI-compatible endpoint through it.
You send tools the usual OpenAI way (tools=[…]); the model emits native markup; the
shim turns it into a standard tool_calls object:
# model completion (raw):
<|tool_call>get_weather{"city": "Paris", "units": "celsius"}
// after the shim:
{"finish_reason": "tool_calls",
"message": {"role": "assistant", "content": null,
"tool_calls": [{"id": "call_0", "type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\", \"units\": \"celsius\"}"}}]}}
Tool-calling gate
tool_eval.sh (8 hand-picked cases: 7 tool + 1 abstain) — Q4_K_M served on llama.cpp --jinja, TOOLS_IN_PROMPT=1, temp 0; SHIM = prod gemma_tool_parse path
The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.
| Measured on | Pass rate |
|---|---|
this model — raw (--jinja) |
0.125 |
| this model — shim (prod path) | 1.000 |
Intended use & limitations
Built for code generation and agentic tool use; serve locally via llama.cpp /
Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or
fabricated — validate tool arguments before executing, and keep a human in the loop
for anything consequential.
⚠️ Uncensored. For this variant the refusal direction has been ablated from the weights — safety guardrails are
substantially removed and it will attempt requests a stock model would refuse. You
are responsible for what you generate and how it's used; not suitable where refusal
behaviour is itself a safety requirement.
Where this sits in the family
- base (upstream) —
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
Provenance & reproduction
How this model was built — technique chain, training mix, and the exact knobs/pins,
so the result is reproducible without any of our tooling.
Mechanics applied
| Step | Technique | What it does | Provenance |
|---|---|---|---|
| 1 | sft-qlora |
QLoRA supervised fine-tune to keep + improve native tool-calling | — |
| 2 | abliteration |
refusal-direction ablation edits the weights to remove refusals | tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5 |
1. sft-qlora
- tools_mode: mixed (xLAM schemas folded; conditional taught)
2. abliteration
weight ablation degrades the canonical
<|tool_call>token — the model tends to leak calls as text markup, so the native llama.cpp parser may not fire. See the tool-call recovery note below to get structured calls back.
Training data, hyperparameters & environment
The supervised fine-tune (sft-qlora step above) is inherited from
Gemma-4 12B Coder — SFT v5 (weights) — see that card
for the full training mix, exact hyperparameters, and pinned environment. The remaining
step(s) above are what this model adds on top; their measured effect is below.
Other measured metrics
| Metric | Value |
|---|---|
| kl_divergence | 0.001 |
| n_trials | 100.000 |
| refusals | 3.000 |
Part of the Gemma-4 12B Coder — active collection.
Something not right, or a request? Open a discussion — happy to help.