← back to catalog · registered 2026-08-22 13:56

libvm/mm-cand-lot_paper_uncensored_refusal_like

libvm Qwen 8.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/libvm%2Fmm-cand-lot_paper_uncensored_refusal_like"
Response includes
  • classification m-uncensored
  • files 16
  • hub_downloads_all_time 50
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
50
8 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-05-25
Downloads over time
Now52→from29↑79%
2837455429 on Jun 1052 on Oct 1152 on Oct 7JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3 text-generation conversational arxiv:2505.09388 license:apache-2.0 text-generation-inference endpoints_compatible region:us

Related

Total size
15.3 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-25 21:33

Files by quantization

Auxiliary files 16 files 15.3 GB
model-00001-of-00005.safetensors 3.72 GB f1b58923 download
model-00002-of-00005.safetensors 3.72 GB 68b7b343 download
model-00003-of-00005.safetensors 3.69 GB 90973b2a download
model-00004-of-00005.safetensors 2.97 GB f5986cd4 download
model-00005-of-00005.safetensors 1.16 GB a24fa362 download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
merge_manifest.json 191 KB 8e5008a0 download
model.safetensors.index.json 32.1 KB 6b926301 download
tokenizer_config.json 9.50 KB 417d038a download
README.md 2.87 KB b1245241 download
.gitattributes 1.53 KB 52373fe2 download
config.json 729 B c5c5b1c6 download
generation_config.json 239 B 20a8a915 download
upload_complete.json 128 B e5eb9d2e download

README current version from Hugging Face


license: apache-2.0
library_name: transformers

Qwen3-8B-Base

Qwen3 Highlights

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models.
Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the following key improvements over the previously released Qwen2.5:

  • Expanded Higher-Quality Pre-training Corpus: Qwen3 is pre-trained on 36 trillion tokens across 119 languages — tripling the language coverage of Qwen2.5 — with a much richer mix of high-quality data, including coding, STEM, reasoning, book, multilingual, and synthetic data.
  • Training Techniques and Model Architecture: Qwen3 incorporates a series of training techiques and architectural refinements, including global-batch load balancing loss for MoE models and qk layernorm for all models, leading to improved stability and overall performance.
  • Three-stage Pre-training: Stage 1 focuses on broad language modeling and general knowledge acquisition, Stage 2 improves reasoning skills like STEM, coding, and logical reasoning, and Stage 3 enhances long-context comprehension by extending training sequence lengths up to 32k tokens.
  • Scaling Law Guided Hyperparameter Tuning: Through comprehensive scaling law studies across the three-stage pre-training pipeline, Qwen3 systematically tunes critical hyperparameters — such as learning rate scheduler and batch size — separately for dense and MoE models, resulting in better training dynamics and final performance across different model scales.

Model Overview

Qwen3-8B-Base has the following features:

  • Type: Causal Language Models
  • Training Stage: Pretraining
  • Number of Parameters: 8.2B
  • Number of Paramaters (Non-Embedding): 6.95B
  • Number of Layers: 36
  • Number of Attention Heads (GQA): 32 for Q and 8 for KV
  • Context Length: 32,768

For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation.

Requirements

The code of Qwen3 has been in the latest Hugging Face transformers and we advise you to use the latest version of transformers.

With transformers<4.51.0, you will encounter the following error:

KeyError: 'qwen3'

Evaluation & Performance

Detailed evaluation results are reported in this 📑 blog.

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen3technicalreport,
      title={Qwen3 Technical Report}, 
      author={Qwen Team},
      year={2025},
      eprint={2505.09388},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.09388}, 
}

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-25Add files using upload-large-folder tool4c468952.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration