license: apache-2.0
language:
- en
- zh
library_name: transformers
pipeline_tag: text-generation
tags: - g9v3
- llama
- text-generation
- long-context
- tool-calling
- heretic
- abliterated
- uncensored
- gguf
- llama.cpp
base_model: - Vortecks/G9v3-3B-Heretic-Abliterated
G9v3-3B
Introduction
G9v3-3B is a dense 3B causal language model from the AI9Stars team, built for local deployment and resource-constrained scenarios. It targets everyday assistant use, coding, tool-use workflows, and reasoning tasks where a compact model is preferred.
This build is a heretic abliterated variant, meaning the refusal directions in the model's activation space have been identified and ablated to remove built-in refusal behavior. As a result, this version is marketed as uncensored and will generally comply with a wider range of prompts than the base instruction-tuned release, without the usual safety-alignment guardrails.
All quantizations were produced with llama.cpp. The low-bit IQ quants (IQ1_* to IQ3_XXS) were generated with an importance matrix.
Download a file (not the whole branch) from below:
| Filename | Quant type | File Size | Split | Description |
|---|---|---|---|---|
| G9v3-3B-Heretic-Abliterated-F32.gguf | f32 | 11.96GB | false | Full 32-bit precision weights. Use only as source or for maximum quality. |
| G9v3-3B-Heretic-Abliterated-F16.gguf | f16 | 5.98GB | false | Full 16-bit weights. Use as the source GGUF. |
| G9v3-3B-Heretic-Abliterated-BF16.gguf | bf16 | 5.98GB | false | Full bfloat16 weights. |
| G9v3-3B-Heretic-Abliterated-Q8_0.gguf | Q8_0 | 3.18GB | false | Extremely high quality, generally unneeded but max available quant. |
| G9v3-3B-Heretic-Abliterated-Q6_K.gguf | Q6_K | 2.46GB | false | Very high quality, near perfect, recommended. |
| G9v3-3B-Heretic-Abliterated-Q5_1.gguf | Q5_1 | 2.27GB | false | Legacy format, slightly higher quality than Q5_0. |
| G9v3-3B-Heretic-Abliterated-Q5_K_M.gguf | Q5_K_M | 2.14GB | false | High quality, recommended. |
| G9v3-3B-Heretic-Abliterated-Q5_K_S.gguf | Q5_K_S | 2.10GB | false | High quality, recommended. |
| G9v3-3B-Heretic-Abliterated-Q5_0.gguf | Q5_0 | 2.10GB | false | Legacy format, high quality. |
| G9v3-3B-Heretic-Abliterated-Q4_1.gguf | Q4_1 | 1.93GB | false | Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon. |
| G9v3-3B-Heretic-Abliterated-Q4_K_M.gguf | Q4_K_M | 1.84GB | false | Good quality, default size for most use cases, recommended. |
| G9v3-3B-Heretic-Abliterated-IQ4_NL.gguf | IQ4_NL | 1.77GB | false | Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference. |
| G9v3-3B-Heretic-Abliterated-Q4_K_S.gguf | Q4_K_S | 1.77GB | false | Slightly lower quality with more space savings, recommended. |
| G9v3-3B-Heretic-Abliterated-Q4_0.gguf | Q4_0 | 1.76GB | false | Legacy format, offers online repacking for ARM and AVX CPU inference. |
| G9v3-3B-Heretic-Abliterated-IQ4_XS.gguf | IQ4_XS | 1.69GB | false | Decent quality, smaller than Q4_K_S with similar performance, recommended. |
| G9v3-3B-Heretic-Abliterated-Q3_K_L.gguf | Q3_K_L | 1.63GB | false | Lower quality but usable, good for low RAM availability. |
| G9v3-3B-Heretic-Abliterated-Q3_K_M.gguf | Q3_K_M | 1.52GB | false | Low quality. |
| G9v3-3B-Heretic-Abliterated-IQ3_M.gguf | IQ3_M | 1.44GB | false | Medium-low quality, new method with decent performance comparable to Q3_K_M. |
| G9v3-3B-Heretic-Abliterated-IQ3_S.gguf | IQ3_S | 1.40GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| G9v3-3B-Heretic-Abliterated-Q3_K_S.gguf | Q3_K_S | 1.39GB | false | Low quality, not recommended. |
| G9v3-3B-Heretic-Abliterated-IQ3_XS.gguf | IQ3_XS | 1.34GB | false | Lower quality, new method with decent performance, slightly better than Q3_K_S. |
| G9v3-3B-Heretic-Abliterated-IQ3_XXS.gguf | IQ3_XXS | 1.24GB | false | Lower quality, new method with decent performance, comparable to Q3 quants. |
| G9v3-3B-Heretic-Abliterated-Q2_K.gguf | Q2_K | 1.21GB | false | Very low quality but surprisingly usable. |
| G9v3-3B-Heretic-Abliterated-IQ2_M.gguf | IQ2_M | 1.13GB | false | Relatively low quality, uses SOTA techniques to be surprisingly usable. |
| G9v3-3B-Heretic-Abliterated-Q2_0.gguf | Q2_0 | 1.07GB | false | Legacy 2-bit format, very low quality. |
| G9v3-3B-Heretic-Abliterated-IQ2_S.gguf | IQ2_S | 1.06GB | false | Low quality, experimental. |
| G9v3-3B-Heretic-Abliterated-TQ2_0.gguf | TQ2_0 | 1.01GB | false | Ternary 2-bit quantization, experimental. |
| G9v3-3B-Heretic-Abliterated-IQ2_XS.gguf | IQ2_XS | 1.00GB | false | Low quality, experimental. |
| G9v3-3B-Heretic-Abliterated-IQ2_XXS.gguf | IQ2_XXS | 0.92GB | false | Very low quality, experimental. |
| G9v3-3B-Heretic-Abliterated-TQ1_0.gguf | TQ1_0 | 0.89GB | false | Ternary 1-bit quantization, experimental. |
| G9v3-3B-Heretic-Abliterated-IQ1_M.gguf | IQ1_M | 0.84GB | false | Very low quality, 1-bit experimental. |
| G9v3-3B-Heretic-Abliterated-IQ1_S.gguf | IQ1_S | 0.79GB | false | Very low quality, 1-bit experimental. |
| G9v3-3B-Heretic-Abliterated-Q1_0.gguf | Q1_0 | 0.61GB | false | Extremely low quality, 1.125 bpw, experimental. |
Which file should I choose?
- For most use cases, use Q4_K_M (good quality / size balance).
- For maximum quality, use Q6_K or Q8_0 (only if you have the RAM).
- For low-memory devices, work down from Q4_K_S → Q3_K_M → Q2_K.
- The IQ1/IQ2/IQ3 types are for squeezing onto very limited hardware.
Model Information
- Type: Causal Language Model
- Architecture: Standard
LlamaForCausalLM - Number of Parameters: ~3B
- Context Length: 131,072
Performance
| Metric | This model | Original model (ai9stars/G9v3-3B) |
|---|---|---|
| KL divergence | 0.0843 | 0 |
| Refusals | 3/100 | 100/100 |
Quickstart
GGUF quantizations of G9v3-3B-Heretic-Abliterated, for use with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes. The example commands below use the Q4_K_M quantization — swap in whichever quant level you've downloaded from the Files and versions tab.
llama.cpp
# Build or install llama.cpp: https://github.com/ggml-org/llama.cpp
huggingface-cli download Vortecks/G9v3-3B-Heretic-Abliterated-GGUF G9v3-3B-Heretic-Abliterated-Q4_K_M.gguf --local-dir .
CLI (interactive chat):
llama-cli -m G9v3-3B-Heretic-Abliterated-Q4_K_M.gguf \
-c 131072 \
-cnv \
--temp 0.7 \
--top-p 0.95
Server (OpenAI-compatible API):
llama-server -m G9v3-3B-Heretic-Abliterated-Q4_K_M.gguf \
-c 131072 \
--port 8080
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Who are you?"}], "max_tokens": 128, "temperature": 0.7}'
Ollama
ollama run hf.co/Vortecks/G9v3-3B-Heretic-Abliterated-GGUF:Q4_K_M
LM Studio
Search for Vortecks/G9v3-3B-Heretic-Abliterated-GGUF directly in the LM Studio model search, or download a .gguf file from the Files and versions tab and load it manually via Load Model from File.
Recommended sampling parameters:
| Mode | Recommended params |
|---|---|
| Think | temperature=0.9, top_p=0.95 |
| No Think | temperature=0.7, top_p=0.95 |
Limitations and Responsible Use
G9v3-3B is a language model that generates content based on learned statistical patterns from training data. It may produce inaccurate, biased, or unsafe outputs, and generated content should be reviewed and verified before use in high-stakes settings. Users are responsible for evaluating outputs, applying appropriate safeguards, and complying with applicable laws, regulations, and platform policies. As an abliterated, uncensored build, this model has had its built-in refusal behavior removed and will generally comply with a wider range of prompts than its base counterpart, so downstream safety, moderation, and policy compliance must be implemented by the deployer rather than relying on the model itself.
License
This repository and the G9v3 model weights are released under the Apache-2.0 License.