license: apache-2.0
base_model: BlossomsAI/Qwen2.5-Coder-7B-Instruct-Uncensored
library_name: transformers
pipeline_tag: text-generation
tags:
- code
- coding
- uncensored
- instruct
- finetuned
- qlora
- qwen2.5
- gguf
- llama.cpp
model-index: - name: Pyrex 8B Instruct Uncensored
results:- task:
type: text-generation
dataset:
type: openai_humaneval
name: HumanEval
metrics:- type: pass@1
value: 41.5
name: pass@1 (greedy)
- type: pass@1
- task:
Pyrex 8B Instruct Uncensored — a sharp, uncensored coding model, fine-tuned from a coding-specialized base. Strong at code, honest by design.
Overview
Pyrex 8B Instruct Uncensored is a QLoRA supervised fine-tune of
BlossomsAI/Qwen2.5-Coder-7B-Instruct-Uncensored.
This is a real, trained model — not a rename — built through our own private,
reproducible pipeline on Hugging Face GPU compute.
It writes clean, correct, and efficient code, fixes bugs, explains technical
problems, and handles agentic / tool-use tasks. It does not refuse answers
on sensitive topics: no safety-lobotomizing data appears anywhere in the
training recipe.
Key strengths
| Area | Detail |
|---|---|
| Coding | Trained primarily on high-quality code instructions (OpenCoder real-user data, CodeAlpaca, evol-codealpaca) |
| Uncensored | Direct, honest answers without refusals |
| Tool-use / agentic | Function-calling data in the mix — works well in agent loops and OpenAI-compatible APIs |
| Portable | Full-precision safetensors + GGUF quants for Ollama, llama.cpp, LM Studio |
Benchmarks
Measured with greedy pass@1 on HumanEval (164 unseen problems), executing
the official unit tests in a sandbox. No sampling luck, no leaked answers.
| Model | HumanEval pass@1 |
|---|---|
| Base (untuned) | 29.3% (48/164) |
| Pyrex 8B | 41.5% (68/164) |
+12.2 points over the stock base on unseen problems.
Quick start
GGUF (recommended for local)
Quantized builds live in
Pyrex-8B-Instruct-Uncensored-GGUF:
| File | Size | Use |
|---|---|---|
pyrex-q4_k_m.gguf |
4.4 GB | Recommended daily driver |
pyrex-q5_k_m.gguf |
5.1 GB | Higher quality |
pyrex-q8_0.gguf |
8.1 GB | Near-lossless |
pyrex-f16.gguf |
15.2 GB | Full precision |
# Ollama (Modelfile ships in the GGUF repo)
ollama create pyrex-8b -f Modelfile
ollama run pyrex-8b "Write a Python function that merges overlapping intervals."
# llama.cpp
./llama-cli -m pyrex-q4_k_m.gguf \
-p "<|im_start|>user\nYour question here\n<|im_end|>\n<|im_start|>assistant\n" \
-n 512 -c 8192
Python (transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ImposterOnline/Pyrex-8B-Instruct-Uncensored"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Explain async/await in Python."}]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
out = model.generate(**tok(text, return_tensors="pt"), max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))
Chat format is chatml (<|im_start|>user / <|im_end|>).
Training recipe
| Parameter | Value |
|---|---|
| Base model | Qwen2.5-Coder-7B-Instruct-Uncensored |
| Method | QLoRA (4-bit NF4) + LoRA r=48, α=96 |
| Data | ~38k cleaned examples (code + tool-use + general) |
| Sequence length | 3072 |
| Loss masking | Assistant-only (completion-only SFT) |
| Optimizer | AdamW, lr 2e-4, cosine, warmup 3% |
| Epochs | 1 |
| Hardware | Hugging Face A100 GPU job |
The full pipeline is config-driven and reproducible: prepare_data.py →train_qlora.py → eval_humaneval.py → publish.sh (see the companion repoArhan-w/pyrut).
Deployment
Any OpenAI-compatible server works (vLLM, llama.cpp, Ollama). Function calling
is part of training, so it suits agentic workflows too.
License
Apache-2.0. Built on Qwen2.5-Coder (Apache-2.0).