← back to catalog · registered 2026-08-22 13:56

littlelearner/unfiltered-0.6b-grpo-math-expert

littlelearner Qwen 600M
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/littlelearner%2Funfiltered-0.6b-grpo-math-expert"
Response includes
  • classification unknown
  • files 8
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
73
↑ 315% in 90 days
Likes
1
Model age
3mo ago
created 2026-07-09
Downloads over time
Now834→from201↑315%
169412655897201 on Aug 16834 on Oct 11AugSepOct
Aug 16 → Oct 11 · 49 snapshots · spans 56 days

Metadata

License
other
Languages
en
Tags
transformers safetensors qwen3 text-generation littlelearner unbounded instruct reinforcement-learning conversational en arxiv:2608.13545 license:other

Related

Total size
1.15 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 13:35

Files by quantization

Auxiliary files 8 files 1.15 GB
model.safetensors 1.15 GB ff72327f download
tokenizer.json 4.41 MB 067f19db download
README.md 2.56 KB 10f77b4d download
.gitattributes 1.48 KB a6344aac download
config.json 1.24 KB e7c707f5 download
tokenizer_config.json 556 B c584a96d download
generation_config.json 235 B 40455263 download
chat_template.jinja 208 B 7d325450 download

README current version from Hugging Face


license: other
language:

  • en
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • qwen3
  • text-generation
  • littlelearner
  • unbounded
  • instruct
  • reinforcement-learning

unfiltered-0.6b-grpo-math-expert

0.617B fully-unbounded chat model post-trained with GRPO on top of SFT.

Part of the LittleLearner scale-up study (pedagogically-controlled knowledge exposure): Qwen3 dense LMs trained on a corpus filtered to U.S. K-5 material (bounded) vs an unfiltered FineWeb-Edu corpus (unbounded), to measure what an interpretable knowledge boundary costs and grants.

Note: This checkpoint was post-trained with GRPO on mathematical reasoning tasks to probe achievable performance on MathCAMPS. As a result, its behavior is specialized toward mathematical reasoning and may not preserve general-purpose chat capabilities; responses may also exhibit a tendency toward math-oriented reasoning or output.

Model

  • Architecture: Qwen3 dense (Qwen3ForCausalLM).
  • Size: 0.617B params, hidden 1536, 20 layers, 12 query / 6 KV heads, FFN 4096. Context: 4096.
  • Tokenizer: custom 64k byte-level BPE with per-digit splitting (ChatML special tokens).
  • Pretraining: 88B tokens on unfiltered FineWeb-Edu (score >= 2, no grade filter). WSD schedule, sharded Muon, MXFP8, Megatron-Core on 8xB200.
  • SFT: supervised fine-tuned on unbounded chat data (lr 1e-5, 3 epochs).
  • RL (GRPO): segmented policy re-banding on a verifiable-answer pool: 4 segments at rollout/training temperature 1.0, then 2 more at temperature 1.3 (no further unlock observed at this scale).

Evaluation

MathCAMPS:

  • K-5 pass@64 57.7 / pass@1 33.8
  • beyond-K-5 pass@64 35.7 / pass@1 9.7

Usage

# transformers (chat)
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "manueldeprada/littlelearner-0.6b-unbounded-grpo"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="cuda")
msgs = [{"role": "user", "content": "Liam has 3 apples and buys 4 more. How many apples does he have?"}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))
# vLLM
from vllm import LLM
repo = "manueldeprada/littlelearner-0.6b-unbounded-grpo"
llm = LLM(repo)
msgs = [{"role": "user", "content": "Liam has 3 apples and buys 4 more. How many apples does he have?"}]
print(llm.chat(msgs)[0].outputs[0].text)

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Update README.md819453e2.6 KB
    Loading...
  2. 2026-08-12Update README.mdd6b86ba2.5 KB
    Loading...
  3. 2026-08-12Update README.md6738aae2.5 KB
    Loading...
  4. 2026-07-09Update model carda2ba81d2.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration