← back to catalog · registered 2026-08-22 13:56

littlelearner/unfiltered-1.3b-grpo-math-expert

littlelearner Qwen 1.3B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/littlelearner%2Funfiltered-1.3b-grpo-math-expert"
Response includes
  • classification unknown
  • files 8
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
87
↑ 340% in 90 days
Likes
1
Model age
3mo ago
created 2026-07-06
Downloads over time
Now823→from187↑340%
155399643887187 on Aug 16823 on Oct 11AugSepOct
Aug 16 → Oct 11 · 49 snapshots · spans 56 days

Metadata

License
other
Languages
en
Tags
transformers safetensors qwen3 text-generation littlelearner unbounded instruct reinforcement-learning conversational en arxiv:2608.13545 license:other

Related

Total size
2.53 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 13:34

Files by quantization

Auxiliary files 8 files 2.53 GB
model.safetensors 2.53 GB 60f0f0e6 download
tokenizer.json 4.41 MB 067f19db download
README.md 2.54 KB e95a5268 download
.gitattributes 1.48 KB a6344aac download
config.json 1.40 KB 9e77ad7f download
tokenizer_config.json 527 B 20005d91 download
generation_config.json 235 B 58b51113 download
chat_template.jinja 208 B 7d325450 download

README current version from Hugging Face


license: other
language:

  • en
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • qwen3
  • text-generation
  • littlelearner
  • unbounded
  • instruct
  • reinforcement-learning

unfiltered-1.3b-grpo-math-expert

1.36B fully-unbounded chat model post-trained with GRPO on top of SFT.

Part of the LittleLearner scale-up study (pedagogically-controlled knowledge exposure): Qwen3 dense LMs trained on a corpus filtered to U.S. K-5 material (bounded) vs an unfiltered FineWeb-Edu corpus (unbounded), to measure what an interpretable knowledge boundary costs and grants.

Note: This checkpoint was post-trained with GRPO on mathematical reasoning tasks to probe achievable performance on MathCAMPS. As a result, its behavior is specialized toward mathematical reasoning and may not preserve general-purpose chat capabilities; responses may also exhibit a tendency toward math-oriented reasoning or output.

Model

  • Architecture: Qwen3 dense (Qwen3ForCausalLM).
  • Size: 1.358B params, hidden 2048, 26 layers, 16 query / 8 KV heads, FFN 5632. Context: 4096.
  • Tokenizer: custom 64k byte-level BPE with per-digit splitting (ChatML special tokens).
  • Pretraining: 88B tokens on unfiltered FineWeb-Edu (score >= 2, no grade filter). WSD schedule, sharded Muon, MXFP8, Megatron-Core on 8xB200.
  • SFT: supervised fine-tuned on unbounded chat data (lr 1e-5, 3 epochs).
  • RL (GRPO): segmented policy re-banding on a verifiable-answer pool: temperature-1.0 segments to a plateau, then one segment at temperature 1.3 (a one-shot unlock).

Evaluation

MathCAMPS:

  • K-5 pass@64 70.6 / pass@1 45.7
  • beyond-K-5 pass@64 45.6 / pass@1 16.5

Usage

# transformers (chat)
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "manueldeprada/littlelearner-1.3b-unbounded-grpo"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="cuda")
msgs = [{"role": "user", "content": "Liam has 3 apples and buys 4 more. How many apples does he have?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
# vLLM
from vllm import LLM
repo = "manueldeprada/littlelearner-1.3b-unbounded-grpo"
llm = LLM(repo)
msgs = [{"role": "user", "content": "Liam has 3 apples and buys 4 more. How many apples does he have?"}]
print(llm.chat(msgs)[0].outputs[0].text)

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Update README.md1737bbc2.5 KB
    Loading...
  2. 2026-08-12Update README.mda6102e22.5 KB
    Loading...
  3. 2026-08-12Update README.md7d748c62.5 KB
    Loading...
  4. 2026-07-08Update model carddb3f0d42.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration