← back to catalog · registered 2026-08-22 13:56

WithinUsAI/Mellum2-Thinker.Uncensored-12B-A2.5B-gguf

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/WithinUsAI%2FMellum2-Thinker.Uncensored-12B-A2.5B-gguf"
Response includes
  • classification m-uncensored
  • files 3
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
4
Model age
4mo ago
created 2026-06-05
Downloads over time
Now0→from0↑0%
00110 on Jun 100 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Metadata

Quantizations
Q4_K
Tags
arxiv:2605.31268 region:us

Related

Total size
7.52 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-06-09 23:40

Files by quantization

Q4_K 1 file 7.52 GB
Mellum2-Thinker-Uncensored-12B-A2.5B-Q4_K_M.GGUF 7.52 GB edb79aac download
Auxiliary files 2 files 6.69 KB
README.md 5.12 KB d6ae7a22 download
.gitattributes 1.57 KB 722ddd1c download

README current version from Hugging Face

Mellum2-Thinker.Uncensored-12B-A2.5B-GGUF

Repository: WithinUsAI/Mellum2-Thinker.Uncensored-12B-A2.5B-GGUF

Overview

Mellum2-Thinker.Uncensored-12B-A2.5B-GGUF is an uncensored community derivative of JetBrains Mellum2-12B-A2.5B-Thinking converted to GGUF format for efficient local inference.

This release preserves the original reasoning-oriented behavior of Mellum2 Thinking while reducing alignment restrictions and refusals wherever possible. The model is intended for research, experimentation, creative writing, roleplay, agentic workflows, coding, reasoning, and unrestricted local AI deployments.

Like the original Mellum2 Thinking model, the model produces reasoning traces within <think>...</think> blocks before generating a final answer. (Hugging Face)


Highlights

  • 🧠 Explicit reasoning with <think> traces
  • ⚡ MoE architecture with only ~2.5B active parameters per token
  • 📚 131K context length
  • 💻 Strong coding and software engineering capabilities
  • 🤖 Agent-friendly reasoning and planning
  • 🔓 Reduced alignment restrictions compared to the original release
  • 🦙 GGUF format for llama.cpp, KoboldCpp, LM Studio, Jan, Open WebUI, and Ollama-compatible ecosystems
  • 🏠 Designed for local and offline deployments

Model Architecture

Mellum2-Thinker.Uncensored inherits the architecture of the original Mellum2 Thinking model:

Attribute Value
Architecture Mixture-of-Experts (MoE)
Total Parameters 12B
Active Parameters 2.5B
Experts 64
Active Experts per Token 8
Layers 28
Hidden Size 2304
Context Length 131,072
Attention Sliding Window + Full Attention
Vocabulary Size 98,304
Precision BF16 Source
Format GGUF

(Hugging Face)


Intended Use

Mellum2-Thinker.Uncensored is best suited for:

  • Advanced reasoning
  • Multi-step problem solving
  • Agent frameworks
  • Coding assistance
  • Software engineering workflows
  • Autonomous task planning
  • Creative writing
  • Storytelling
  • Worldbuilding
  • Roleplay
  • Research
  • Knowledge exploration

The model is particularly effective when explicit reasoning and chain-of-thought style outputs are desired.


Prompt Format

Chat Format

<|im_start|>system
You are a helpful assistant.
<|im_end|>

<|im_start|>user
Explain recursion.
<|im_end|>

<|im_start|>assistant

Thinking Example

User: Solve this problem.

Assistant:
<think>
Step-by-step reasoning...
</think>

Final answer...

Quantization Information

This repository contains GGUF quantizations for local inference.

Typical recommendations:

Quant Recommended RAM/VRAM
Q4_K_M 8-10 GB

Actual memory requirements vary by context length and backend.


Performance Characteristics

Mellum2 was designed as a high-efficiency focal reasoning model where only 2.5B parameters are activated per token despite containing 12B total parameters. This allows significantly faster inference than similarly sized dense models while retaining strong reasoning and coding capabilities. (arXiv)


Differences From The Original Release

This repository is not an official JetBrains release.

Changes include:

  • Conversion to GGUF format
  • Community packaging for local inference
  • Reduced refusal behavior
  • Reduced alignment constraints
  • Intended for unrestricted research and experimentation
  • Preservation of reasoning-focused behavior

No affiliation with JetBrains is implied.


License

This derivative is based on Mellum2, which was released under the Apache 2.0 License. Please review the original license and ensure compliance with all applicable terms.

Original Model:

JetBrains/Mellum2-12B-A2.5B-Thinking

Original Technical Report:

Mellum2 Technical Report


Acknowledgements

Special thanks to JetBrains for releasing Mellum2 as an open-weight model and making advanced reasoning-focused MoE architectures available to the open-source AI community. (Hugging Face)

Maintained by: WithinUsAI

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-09Update README.mde55757e5.1 KB
    Loading...
  2. 2026-06-09Update README.md78b81e75.4 KB
    Loading...
  3. 2026-06-09Create README.mda11f2755.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration