← back to catalog · registered 2026-08-22 13:56

Goraint/Qwen3-8b-192k-Context-6X-Josiefied-Uncensored-MLX-AWQ-4bit

Goraint Qwen 8.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Goraint%2FQwen3-8b-192k-Context-6X-Josiefied-Uncensored-MLX-AWQ-4bit"
Response includes
  • classification m-uncensored
  • files 12
  • benchmarks 11 entries
  • hub_downloads_all_time 7,551
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
8K
81 last 30d - cooling
Likes
1
Model age
17mo ago
created 2025-05-15
Downloads over time
Now7.6K→from40↑18,848%
02.8K5.6K8.3K40 on May 14, 20257.6K on Oct 11May '25Aug '25Nov '25FebMayAug
May 14, 2025 → Oct 11 · 113 snapshots · spans 515 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.15 UGI
Political lean -9.9% UGI
Sensitive-Info 18.27 UGI
SocPol 1.4 UGI
UGI 32.18 UGI
Willingness (10) 6 UGI
W10-Adherence 7 UGI
W10-Direct 5 UGI
Writing 27.96 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3 not-for-all-audiences text-generation conversational base_model:DavidAU/Qwen3-8B-192k-Context-6X-Josiefied-Uncensored base_model:finetune:DavidAU/Qwen3-8B-192k-Context-6X-Josiefied-Uncensored license:apache-2.0 region:us

Related

Total size
4.36 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-05-22 02:37

Files by quantization

Auxiliary files 12 files 4.38 GB
model.safetensors 4.36 GB 4a7750a9 download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 62.6 KB d28689e6 download
config.json 36.6 KB 45f711dd download
tokenizer_config.json 9.48 KB f25f41d9 download
README.md 8.39 KB fb74e747 download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B e4f1d319 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3-8B
  • DavidAU/Qwen3-8B-192k-Context-6X-Josiefied-Uncensored
    library_name: mlx
    tags:
  • not-for-all-audiences
    pipeline_tag: text-generation

📌 Overview

A 4-bit AWQ quantized version of Qwen3-8B optimized for efficient inference using the MLX library, designed to handle long-context tasks (192k tokens) with reduced resource usage. Retains core capabilities of Qwen3-8B while enabling deployment on edge devices.


📈 Performance Metrics

Metric Value
Model Size ~4.38 GB (4-bit quantized)
Inference Speed 30.58 tokens/sec (M1 MAX)
112.80 tokens/sec (M3 ULTRA)
gguf Q4_K_S 8.14 tokens/sec (M1 MAX)
Context Support 192,000 tokens

🚨 Important: Prompt Template for LM Studio Use

You need to modify the prompt template to ensure compatibility with LM Studio's inference pipeline. Below is the required template structure:

{%- if tools %}
    {{- '\/system\n' }}
    {%- if messages[0].role == 'system' %}
        {{- messages[0].content + '\n\n' }}
    {%- endif %}
    {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
    {%- for tool in tools %}
        {{- "\n" }}
        {{- tool | tojson }}
    {%- endfor %}
    {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call>...</tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call>\n" }}
{%- else %}
    {%- if messages[0].role == 'system' %}
        {{- '\/system\n' + messages[0].content + '\/\n' }}
    {%- endif %}
{%- endif %}

{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
{%- for message in messages[::-1] %}
    {%- set index = (messages|length - 1) - loop.index0 %}
    {%- set tool_start = "⦅" %}
    {%- set tool_start_length = tool_start|length %}
    {%- set start_of_message = message.content[:tool_start_length] %}
    {%- set tool_end = "⦆" %}
    {%- set tool_end_length = tool_end|length %}
    {%- set start_pos = (message.content|length) - tool_end_length %}
    {%- if start_pos < 0 %}
        {%- set start_pos = 0 %}
    {%- endif %}
    {%- set end_of_message = message.content[start_pos:] %}
    {%- if ns.multi_step_tool and message.role == "user" and not(start_of_message == tool_start and end_of_message == tool_end) %}
        {%- set ns.multi_step_tool = false %}
        {%- set ns.last_query_index = index %}
    {%- endif %}
{%- endfor %}

{%- for message in messages %}
    {%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
        {{- '\/' + message.role + '\n' + message.content + '\/' + '\n' }}
    {%- elif message.role == "assistant" %}
        {%- set content = message.content %}
        {%- set reasoning_content = '' %}
        {%- if message.reasoning_content is defined and message.reasoning_content is not none %}
            {%- set reasoning_content = message.reasoning_content %}
        {%- else %}
            {%- if '\/' in message.content %}
                {%- set content = (message.content.split('\/')|last).lstrip('\n') %}
                {%- set reasoning_content = (message.content.split('\/')|first).rstrip('\n') %}
                {%- set reasoning_content = (reasoning_content.split('')|last).lstrip('\n') %}
            {%- endif %}
        {%- endif %}

        {%- if loop.index0 > ns.last_query_index %}
            {%- if loop.last or (not loop.last and reasoning_content) %}
                {{- '\/' + message.role + '\n\n' + reasoning_content.strip('\n') + '\n\/\n' + content.lstrip('\n') }}
            {%- else %}
                {{- '\/' + message.role + '\n' + content }}
            {%- endif %}
        {%- else %}
            {{- '\/' + message.role + '\n' + content }}
        {%- endif %}

        {%- if message.tool_calls %}
            {%- for tool_call in message.tool_calls %}
                {%- if (loop.first and content) or (not loop.first) %}
                    {{- '\n' }}
                {%- endif %}
                {%- if tool_call.function %}
                    {%- set tool_call = tool_call.function %}
                {%- endif %}
                {{- '<tool_call>\n{"name": "' }}
                {{- tool_call.name }}
                {{- '", "arguments": ' }}
                {%- if tool_call.arguments is string %}
                    {{- tool_call.arguments }}
                {%- else %}
                    {{- tool_call.arguments | tojson }}
                {%- endif %}
                {{- '}\n</tool_call>' }}
            {%- endfor %}
        {%- endif %}
        {{- '\/\n' }}
    {%- elif message.role == "tool" %}
        {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
            {{- '\/user' }}
        {%- endif %}
        {{- '\n⦅\n' }}
        {{- message.content }}
        {{- '\n⦆' }}
        {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
            {{- '\/\n' }}
        {%- endif %}
    {%- endif %}
{%- endfor %}

{%- if add_generation_prompt %}
    {{- '\/assistant\n' }}
    {%- if enable_thinking is defined and enable_thinking is false %}
        {{- '

🧠 Model Details

  • Base Model: Qwen3-8B
  • Quantization: AWQ Q4 (4-bit) via MLX library
  • Context Length: 192,000 tokens (6x longer than standard)
  • Library: MLX (optimized for Apple Silicon, macOS)
  • License: Apache 2.0
  • Pipeline: text-generation
  • Tags: not-for-all-audiences, conversational, mlx

🧩 Key Features

  • Efficient Inference: 4-bit quantization reduces memory footprint by ~75% vs. FP16
  • Long Context Support: 192k tokens for complex tasks (e.g., document analysis, code generation)
  • Cross-Platform: Works on macOS with MLX for Apple Silicon acceleration
  • Customizable Prompting: Adjust templates for compatibility with tools like LM Studio

📌 Use Cases

  • Research: Long-context NLP experiments, model compression studies
  • Development: Edge deployments, real-time chatbots with extended context
  • Enterprise: Cost-effective AI solutions for document processing and code generation

⚠️ Bias, Risks & Limitations

Potential Biases

  • Trained on diverse data but may inherit societal biases (e.g., gender, cultural assumptions)
  • "Not-for-all-audiences" tag indicates potential for generating sensitive content

Technical Limitations

  • 4-bit quantization may slightly reduce accuracy on complex tasks
  • Performance depends on hardware (MLX optimized for Apple Silicon)

Mitigation Strategies

  • Review outputs for sensitive content
  • Use in controlled environments with monitoring

🧪 How to Get Started

Installation

# Install MLX (Apple Silicon only)
pip install mlx

# Load model with Hugging Face Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Goraint/Qwen3-8b-192k-Context-6X-Josiefied-Uncensored-MLX-AWQ-4bit", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("Goraint/Qwen3-8b-192k-Context-6X-Josiefied-Uncensored-MLX-AWQ-4bit")

Example Usage

prompt = "Explain quantum computing in simple terms."
inputs = tokenizer(prompt, return_tensors="pt").to("mps")
outputs = model.generate(**inputs, max_length=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

🌍 Environmental Impact


🧰 Community & Resources


📄 License

Apache 2.0

Note: This model is a community contribution and may not be officially supported by Alibaba Cloud. Always validate outputs for accuracy and safety in production environments.

README history 20 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-05-22Update README.md33765fb8.4 KB
    Loading...
  2. 2025-05-16Update README.md181d3658.4 KB
    Loading...
  3. 2025-05-16Update README.mdd726fba8.3 KB
    Loading...
  4. 2025-05-16Update README.md09438d58.4 KB
    Loading...
  5. 2025-05-16Update README.md20a27178.4 KB
    Loading...
  6. 2025-05-16Update README.mdc727d878.4 KB
    Loading...
  7. 2025-05-16Update README.md6f43f888.4 KB
    Loading...
  8. 2025-05-16Update README.md4e610f88.4 KB
    Loading...
  9. 2025-05-16Update README.mde38967b8.4 KB
    Loading...
  10. 2025-05-16Update README.md1322b398.4 KB
    Loading...
  11. 2025-05-16Update README.md867c5a78.4 KB
    Loading...
  12. 2025-05-16Update README.mdebd4f483.9 KB
    Loading...
  13. 2025-05-16Update README.md26146673.9 KB
    Loading...
  14. 2025-05-16Update README.md16a869f7.9 KB
    Loading...
  15. 2025-05-16Update README.md054facf6.7 KB
    Loading...
  16. 2025-05-16Update README.md379f63a6.7 KB
    Loading...
  17. 2025-05-15Update README.mdf6d379a2.1 KB
    Loading...
  18. 2025-05-15Update README.md5b36fe83.5 KB
    Loading...
  19. 2025-05-15Update README.mdb27b9415.3 KB
    Loading...
  20. 2025-05-15Update README.md1b271ab5.3 KB
    Loading...

Discussions 1 thread

  1. 2025-07-06Prompt Template invalidopen3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration