← back to catalog · registered 2026-08-22 13:56

bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF

bartowski Llama 70B GGUF second-order 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bartowski%2FLlama-3.3-70B-Instruct-abliterated-GGUF"
Response includes
  • classification m8
  • files 24
  • benchmarks 11 entries
  • hub_downloads_all_time 332,604
  • author_summary 72 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 3 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • author=bartowski (M8 quantization producer)
  • is_gguf=1
  • base_model='huihui-ai/Llama-3.3-70B-Instruct-abliterated' looks abliterated -> assume M1
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
333K
54K last 30d - stable
Likes
39
Model age
21mo ago
created 2024-12-23
Downloads over time
Now352.2K→from3.4K↑10,112%
0129K258.1K387.1K3.4K on Dec 25, 2024352.2K on Oct 11Dec '24Mar '25Jun '25Sep '25Dec '25MarJunSep
Dec 25, 2024 → Oct 11 · 135 snapshots · spans 655 days

Benchmarks

Benchmark Score Source
Entertainment 3.2 UGI
Hazardous 3.5 UGI
Natural Intelligence 33.54 UGI
Political lean -20.3% UGI
Sensitive-Info 32.66 UGI
SocPol 3.1 UGI
UGI 48.44 UGI
Willingness (10) 8 UGI
W10-Adherence 8 UGI
W10-Direct 8 UGI
Writing 31.76 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.3
Languages
en fr it pt hi es th de
Quantizations
IQ1 IQ2 IQ3 IQ4 Q2_K Q3_K Q4 Q4_K Q5_K
Tags
gguf facebook meta llama llama-3 abliterated uncensored text-generation en fr it pt

Related

Total size
646 GB
Files
24
Quantizations
10
Registered
2026-08-22 13:56
Last updated on HF
2024-12-24 03:01

Files by quantization

Q5_K 1 file 45.3 GB
Llama-3.3-70B-Instruct-abliterated-Q5_K_S.gguf 45.3 GB b2f816f1 download
Q4 2 files 78.6 GB
Llama-3.3-70B-Instruct-abliterated-Q4_1.gguf 41.3 GB f528a972 download
Llama-3.3-70B-Instruct-abliterated-Q4_0.gguf 37.4 GB 84d4ded4 download
Q4_K 3 files 118 GB
Llama-3.3-70B-Instruct-abliterated-Q4_K_L.gguf 40.3 GB 0d6c02e1 download
Llama-3.3-70B-Instruct-abliterated-Q4_K_M.gguf 39.6 GB 5804f6de download
Llama-3.3-70B-Instruct-abliterated-Q4_K_S.gguf 37.6 GB d5873b10 download
IQ4 2 files 72.6 GB
Llama-3.3-70B-Instruct-abliterated-IQ4_NL.gguf 37.3 GB 96eca493 download
Llama-3.3-70B-Instruct-abliterated-IQ4_XS.gguf 35.3 GB 9ee44e17 download
Q3_K 4 files 131 GB
Llama-3.3-70B-Instruct-abliterated-Q3_K_XL.gguf 35.4 GB 53e42106 download
Llama-3.3-70B-Instruct-abliterated-Q3_K_L.gguf 34.6 GB d298c849 download
Llama-3.3-70B-Instruct-abliterated-Q3_K_M.gguf 31.9 GB ed69c969 download
Llama-3.3-70B-Instruct-abliterated-Q3_K_S.gguf 28.8 GB 1ea761ee download
IQ3 2 files 55.3 GB
Llama-3.3-70B-Instruct-abliterated-IQ3_M.gguf 29.7 GB 83053ebc download
Llama-3.3-70B-Instruct-abliterated-IQ3_XXS.gguf 25.6 GB 1b7714c7 download
Q2_K 2 files 50.1 GB
Llama-3.3-70B-Instruct-abliterated-Q2_K_L.gguf 25.5 GB bc86a9d6 download
Llama-3.3-70B-Instruct-abliterated-Q2_K.gguf 24.6 GB d9a1fc1e download
IQ2 4 files 80.7 GB
Llama-3.3-70B-Instruct-abliterated-IQ2_M.gguf 22.5 GB 4503cf9b download
Llama-3.3-70B-Instruct-abliterated-IQ2_S.gguf 20.7 GB 1d20973b download
Llama-3.3-70B-Instruct-abliterated-IQ2_XS.gguf 19.7 GB 5a3cef9f download
Llama-3.3-70B-Instruct-abliterated-IQ2_XXS.gguf 17.8 GB a56c420f download
IQ1 1 file 15.6 GB
Llama-3.3-70B-Instruct-abliterated-IQ1_M.gguf 15.6 GB 801c20ad download
Auxiliary files 3 files 23.8 MB
Llama-3.3-70B-Instruct-abliterated.imatrix 23.8 MB b44d6a85 download
README.md 30.0 KB 9c103242 download
.gitattributes 4.33 KB 3c06fdc8 download

README current version from Hugging Face


quantized_by: bartowski
pipeline_tag: text-generation
language:

  • en
  • fr
  • it
  • pt
  • hi
  • es
  • th
  • de
    extra_gated_prompt: "### LLAMA 3.3 COMMUNITY LICENSE AGREEMENT\nLlama 3.3 Version
    \ Release Date: December 6, 2024\n"Agreement" means the terms and conditions for
    \ use, reproduction, distribution and modification of the Llama Materials set forth
    \ herein.\n"Documentation" means the specifications, manuals and documentation
    \ accompanying Llama 3.3 distributed by Meta at https://www.llama.com/docs/overview.\n
    "Licensee" or "you" means you, or your employer or any other person or entity
    \ (if you are entering into this Agreement on such person or entity’s behalf), of
    \ the age required under applicable laws, rules or regulations to provide legal
    \ consent and that has legal authority to bind your employer or such other person
    \ or entity if you are entering in this Agreement on their behalf.\n"Llama 3.3"
    \ means the foundational large language models and software and algorithms, including
    \ machine-learning model code, trained model weights, inference-enabling code, training-enabling
    \ code, fine-tuning enabling code and other elements of the foregoing distributed
    \ by Meta at https://www.llama.com/llama-downloads.\n
    "Llama Materials" means, collectively, Meta’s proprietary Llama 3.3 and Documentation
    \ (and any portion thereof) made available under this Agreement.\n"Meta" or "
    we" means Meta Platforms Ireland Limited (if you are located in or, if you are
    \ an entity, your principal place of business is in the EEA or Switzerland) and
    \ Meta Platforms, Inc. (if you are located outside of the EEA or Switzerland).\n
    By clicking “I Accept” below or by using or distributing any portion or element
    \ of the Llama Materials, you agree to be bound by this Agreement.\n1. License Rights
    \ and Redistribution.\na. Grant of Rights. You are granted a non-exclusive, worldwide,
    \ non-transferable and royalty-free limited license under Meta’s intellectual property
    \ or other rights owned by Meta embodied in the Llama Materials to use, reproduce,
    \ distribute, copy, create derivative works of, and make modifications to the Llama
    \ Materials.\nb. Redistribution and Use.\ni. If you distribute or make available
    \ the Llama Materials (or any derivative works thereof), or a product or service
    \ (including another AI model) that contains any of them, you shall (A) provide
    \ a copy of this Agreement with any such Llama Materials; and (B) prominently display
    \ “Built with Llama” on a related website, user interface, blogpost, about page,
    \ or product documentation. If you use the Llama Materials or any outputs or results
    \ of the Llama Materials to create, train, fine tune, or otherwise improve an AI
    \ model, which is distributed or made available, you shall also include “Llama”
    \ at the beginning of any such AI model name.\nii. If you receive Llama Materials,
    \ or any derivative works thereof, from a Licensee as part of an integrated end
    \ user product, then Section 2 of this Agreement will not apply to you. \niii. You
    \ must retain in all copies of the Llama Materials that you distribute the following
    \ attribution notice within a “Notice” text file distributed as a part of such copies:
    \ “Llama 3.3 is licensed under the Llama 3.3 Community License, Copyright © Meta
    \ Platforms, Inc. All Rights Reserved.”\niv. Your use of the Llama Materials must
    \ comply with applicable laws and regulations (including trade compliance laws and
    \ regulations) and adhere to the Acceptable Use Policy for the Llama Materials (available
    \ at https://www.llama.com/llama3\_3/use-policy),
    \ which is hereby incorporated by reference into this Agreement. \n2. Additional
    \ Commercial Terms. If, on the Llama 3.3 version release date, the monthly active
    \ users of the products or services made available by or for Licensee, or Licensee’s
    \ affiliates, is greater than 700 million monthly active users in the preceding
    \ calendar month, you must request a license from Meta, which Meta may grant to
    \ you in its sole discretion, and you are not authorized to exercise any of the
    \ rights under this Agreement unless or until Meta otherwise expressly grants you
    \ such rights.\n3. Disclaimer of Warranty. UNLESS REQUIRED BY APPLICABLE LAW, THE
    \ LLAMA MATERIALS AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED ON AN “AS IS”
    \ BASIS, WITHOUT WARRANTIES OF ANY KIND, AND META DISCLAIMS ALL WARRANTIES OF ANY
    \ KIND, BOTH EXPRESS AND IMPLIED, INCLUDING, WITHOUT LIMITATION, ANY WARRANTIES
    \ OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
    \ YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING
    \ THE LLAMA MATERIALS AND ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE LLAMA
    \ MATERIALS AND ANY OUTPUT AND RESULTS.\n4. Limitation of Liability. IN NO EVENT
    \ WILL META OR ITS AFFILIATES BE LIABLE UNDER ANY THEORY OF LIABILITY, WHETHER IN
    \ CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING OUT OF THIS
    \ AGREEMENT, FOR ANY LOST PROFITS OR ANY INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL,
    \ EXEMPLARY OR PUNITIVE DAMAGES, EVEN IF META OR ITS AFFILIATES HAVE BEEN ADVISED
    \ OF THE POSSIBILITY OF ANY OF THE FOREGOING.\n5. Intellectual Property.\na. No
    \ trademark licenses are granted under this Agreement, and in connection with the
    \ Llama Materials, neither Meta nor Licensee may use any name or mark owned by or
    \ associated with the other or any of its affiliates, except as required for reasonable
    \ and customary use in describing and redistributing the Llama Materials or as set
    \ forth in this Section 5(a). Meta hereby grants you a license to use “Llama” (the
    \ “Mark”) solely as required to comply with the last sentence of Section 1.b.i.
    \ You will comply with Meta’s brand guidelines (currently accessible at https://about.meta.com/brand/resources/meta/company-brand/).
    \ All goodwill arising out of your use of the Mark will inure to the benefit of
    \ Meta.\nb. Subject to Meta’s ownership of Llama Materials and derivatives made
    \ by or for Meta, with respect to any derivative works and modifications of the
    \ Llama Materials that are made by you, as between you and Meta, you are and will
    \ be the owner of such derivative works and modifications.\nc. If you institute
    \ litigation or other proceedings against Meta or any entity (including a cross-claim
    \ or counterclaim in a lawsuit) alleging that the Llama Materials or Llama 3.3 outputs
    \ or results, or any portion of any of the foregoing, constitutes infringement of
    \ intellectual property or other rights owned or licensable by you, then any licenses
    \ granted to you under this Agreement shall terminate as of the date such litigation
    \ or claim is filed or instituted. You will indemnify and hold harmless Meta from
    \ and against any claim by any third party arising out of or related to your use
    \ or distribution of the Llama Materials.\n6. Term and Termination. The term of
    \ this Agreement will commence upon your acceptance of this Agreement or access
    \ to the Llama Materials and will continue in full force and effect until terminated
    \ in accordance with the terms and conditions herein. Meta may terminate this Agreement
    \ if you are in breach of any term or condition of this Agreement. Upon termination
    \ of this Agreement, you shall delete and cease use of the Llama Materials. Sections
    \ 3, 4 and 7 shall survive the termination of this Agreement.\n7. Governing Law
    \ and Jurisdiction. This Agreement will be governed and construed under the laws
    \ of the State of California without regard to choice of law principles, and the
    \ UN Convention on Contracts for the International Sale of Goods does not apply
    \ to this Agreement. The courts of California shall have exclusive jurisdiction
    \ of any dispute arising out of this Agreement.\n### Llama 3.3 Acceptable Use Policy\n
    Meta is committed to promoting safe and fair use of its tools and features, including
    \ Llama 3.3. If you access or use Llama 3.3, you agree to this Acceptable Use Policy
    \ (“Policy”). The most recent copy of this policy can be found at https://www.llama.com/llama3\
    _3/use-policy
    .\nProhibited Uses\nWe
    \ want everyone to use Llama 3.3 safely and responsibly. You agree you will not
    \ use, or allow others to use, Llama 3.3 to:\n1. Violate the law or others’ rights,
    \ including to:\n\n 1. Engage in, promote, generate, contribute to, encourage,
    \ plan, incite, or further illegal or unlawful activity or content, such as: \n
    \ 1. Violence or terrorism \n 2. Exploitation or harm to children, including
    \ the solicitation, creation, acquisition, or dissemination of child exploitative
    \ content or failure to report Child Sexual Abuse Material \n 3. Human trafficking,
    \ exploitation, and sexual violence \n 4. The illegal distribution of information
    \ or materials to minors, including obscene materials, or failure to employ legally
    \ required age-gating in connection with such information or materials. \n
    \ 5. Sexual solicitation \n 6. Any other criminal activity\n\n 2. Engage
    \ in, promote, incite, or facilitate the harassment, abuse, threatening, or bullying
    \ of individuals or groups of individuals\n\n 3. Engage in, promote, incite, or
    \ facilitate discrimination or other unlawful or harmful conduct in the provision
    \ of employment, employment benefits, credit, housing, other economic benefits,
    \ or other essential goods and services\n\n 4. Engage in the unauthorized or unlicensed
    \ practice of any profession including, but not limited to, financial, legal, medical/health,
    \ or related professional practices\n\n 5. Collect, process, disclose, generate,
    \ or infer private or sensitive information about individuals, including information
    \ about individuals’ identity, health, or demographic information, unless you have
    \ obtained the right to do so in accordance with applicable law\n\n 6. Engage
    \ in or facilitate any action or generate any content that infringes, misappropriates,
    \ or otherwise violates any third-party rights, including the outputs or results
    \ of any products or services using the Llama Materials\n\n 7. Create, generate,
    \ or facilitate the creation of malicious code, malware, computer viruses or do
    \ anything else that could disable, overburden, interfere with or impair the proper
    \ working, integrity, operation or appearance of a website or computer system\n\n
    \ 8. Engage in any action, or facilitate any action, to intentionally circumvent
    \ or remove usage restrictions or other safety measures, or to enable functionality
    \ disabled by Meta\n\n2. Engage in, promote, incite, facilitate, or assist in the
    \ planning or development of activities that present a risk of death or bodily harm
    \ to individuals, including use of Llama 3.3 related to the following:\n\n 1.
    \ Military, warfare, nuclear industries or applications, espionage, use for materials
    \ or activities that are subject to the International Traffic Arms Regulations (ITAR)
    \ maintained by the United States Department of State or to the U.S. Biological
    \ Weapons Anti-Terrorism Act of 1989 or the Chemical Weapons Convention Implementation
    \ Act of 1997\n\n 2. Guns and illegal weapons (including weapon development)\n
    \n 3. Illegal drugs and regulated/controlled substances\n\n 4. Operation of
    \ critical infrastructure, transportation technologies, or heavy machinery\n\n
    \ 5. Self-harm or harm to others, including suicide, cutting, and eating disorders\n
    \n 6. Any content intended to incite or promote violence, abuse, or any infliction
    \ of bodily harm to an individual\n\n3. Intentionally deceive or mislead others,
    \ including use of Llama 3.3 related to the following:\n\n 1. Generating, promoting,
    \ or furthering fraud or the creation or promotion of disinformation\n\n 2. Generating,
    \ promoting, or furthering defamatory content, including the creation of defamatory
    \ statements, images, or other content\n\n 3. Generating, promoting, or further
    \ distributing spam\n\n 4. Impersonating another individual without consent, authorization,
    \ or legal right\n\n 5. Representing that the use of Llama 3.3 or outputs are
    \ human-generated\n\n 6. Generating or facilitating false online engagement, including
    \ fake reviews and other means of fake online engagement\n\n4. Fail to appropriately
    \ disclose to end users any known dangers of your AI system\n5. Interact with third
    \ party tools, models, or software designed to generate unlawful content or engage
    \ in unlawful or harmful conduct and/or represent that the outputs of such tools,
    \ models, or software are associated with Meta or Llama 3.3\nWith respect to any
    \ multimodal models included in Llama 3.3, the rights granted under Section 1(a)
    \ of the Llama 3.3 Community License Agreement are not being granted to you if you
    \ are an individual domiciled in, or a company with a principal place of business
    \ in, the European Union. This restriction does not apply to end users of a product
    \ or service that incorporates any such multimodal models.\nPlease report any violation
    \ of this Policy, software “bug,” or other problems that could lead to a violation
    \ of this Policy through one of the following means:\n* Reporting issues with the
    \ model: https://github.com/meta-llama/llama-models/issues
    \ * Reporting risky content generated by the model: developers.facebook.com/llama\
    _output\_feedback
    * Reporting
    \ bugs and security concerns: facebook.com/whitehat/info
    \ * Reporting violations of the Acceptable Use Policy or unlicensed uses of Llama
    \ 3.3: [email protected] "
    extra_gated_fields:
    First Name: text
    Last Name: text
    Date of birth: date_picker
    Country: country
    Affiliation: text
    Job title:
    type: select
    options:
    • Student
    • Research Graduate
    • AI researcher
    • AI developer/engineer
    • Reporter
    • Other
      geo: ip_location
      ? By clicking Submit below I accept the terms of the license and acknowledge that
      the information I provide will be collected stored processed and shared in accordance
      with the Meta Privacy Policy
      : checkbox
      base_model: huihui-ai/Llama-3.3-70B-Instruct-abliterated
      license: llama3.3
      extra_gated_button_content: Submit
      tags:
  • facebook
  • meta
  • llama
  • llama-3
  • abliterated
  • uncensored
    extra_gated_description: The information you provide will be collected, stored, processed
    and shared in accordance with the Meta Privacy Policy.

Llamacpp imatrix Quantizations of Llama-3.3-70B-Instruct-abliterated

Using llama.cpp release b4381 for quantization.

Original model: https://huggingface.co/huihui-ai/Llama-3.3-70B-Instruct-abliterated

All quants made using imatrix option with dataset from here

Run them in LM Studio

Prompt format

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

Cutting Knowledge Date: December 2023
Today Date: 26 Jul 2024

{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>

{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Download a file (not the whole branch) from below:

Filename Quant type File Size Split Description
Llama-3.3-70B-Instruct-abliterated-Q8_0.gguf Q8_0 74.98GB true Extremely high quality, generally unneeded but max available quant.
Llama-3.3-70B-Instruct-abliterated-Q6_K.gguf Q6_K 57.89GB true Very high quality, near perfect, recommended.
Llama-3.3-70B-Instruct-abliterated-Q5_K_L.gguf Q5_K_L 50.60GB true Uses Q8_0 for embed and output weights. High quality, recommended.
Llama-3.3-70B-Instruct-abliterated-Q5_K_M.gguf Q5_K_M 49.95GB true High quality, recommended.
Llama-3.3-70B-Instruct-abliterated-Q5_K_S.gguf Q5_K_S 48.66GB false High quality, recommended.
Llama-3.3-70B-Instruct-abliterated-Q4_1.gguf Q4_1 44.31GB false Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon.
Llama-3.3-70B-Instruct-abliterated-Q4_K_L.gguf Q4_K_L 43.30GB false Uses Q8_0 for embed and output weights. Good quality, recommended.
Llama-3.3-70B-Instruct-abliterated-Q4_K_M.gguf Q4_K_M 42.52GB false Good quality, default size for most use cases, recommended.
Llama-3.3-70B-Instruct-abliterated-Q4_K_S.gguf Q4_K_S 40.35GB false Slightly lower quality with more space savings, recommended.
Llama-3.3-70B-Instruct-abliterated-Q4_0.gguf Q4_0 40.12GB false Legacy format, offers online repacking for ARM and AVX CPU inference.
Llama-3.3-70B-Instruct-abliterated-IQ4_NL.gguf IQ4_NL 40.05GB false Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference.
Llama-3.3-70B-Instruct-abliterated-Q3_K_XL.gguf Q3_K_XL 38.06GB false Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability.
Llama-3.3-70B-Instruct-abliterated-IQ4_XS.gguf IQ4_XS 37.90GB false Decent quality, smaller than Q4_K_S with similar performance, recommended.
Llama-3.3-70B-Instruct-abliterated-Q3_K_L.gguf Q3_K_L 37.14GB false Lower quality but usable, good for low RAM availability.
Llama-3.3-70B-Instruct-abliterated-Q3_K_M.gguf Q3_K_M 34.27GB false Low quality.
Llama-3.3-70B-Instruct-abliterated-IQ3_M.gguf IQ3_M 31.94GB false Medium-low quality, new method with decent performance comparable to Q3_K_M.
Llama-3.3-70B-Instruct-abliterated-Q3_K_S.gguf Q3_K_S 30.91GB false Low quality, not recommended.
Llama-3.3-70B-Instruct-abliterated-IQ3_XXS.gguf IQ3_XXS 27.47GB false Lower quality, new method with decent performance, comparable to Q3 quants.
Llama-3.3-70B-Instruct-abliterated-Q2_K_L.gguf Q2_K_L 27.40GB false Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable.
Llama-3.3-70B-Instruct-abliterated-Q2_K.gguf Q2_K 26.38GB false Very low quality but surprisingly usable.
Llama-3.3-70B-Instruct-abliterated-IQ2_M.gguf IQ2_M 24.12GB false Relatively low quality, uses SOTA techniques to be surprisingly usable.
Llama-3.3-70B-Instruct-abliterated-IQ2_S.gguf IQ2_S 22.24GB false Low quality, uses SOTA techniques to be usable.
Llama-3.3-70B-Instruct-abliterated-IQ2_XS.gguf IQ2_XS 21.14GB false Low quality, uses SOTA techniques to be usable.
Llama-3.3-70B-Instruct-abliterated-IQ2_XXS.gguf IQ2_XXS 19.10GB false Very low quality, uses SOTA techniques to be usable.
Llama-3.3-70B-Instruct-abliterated-IQ1_M.gguf IQ1_M 16.75GB false Extremely low quality, not recommended.

Embed/output weights

Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.

Downloading using huggingface-cli

Click to view download instructions

First, make sure you have hugginface-cli installed:

pip install -U "huggingface_hub[cli]"

Then, you can target the specific file you want:

huggingface-cli download bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF --include "Llama-3.3-70B-Instruct-abliterated-Q4_K_M.gguf" --local-dir ./

If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:

huggingface-cli download bartowski/Llama-3.3-70B-Instruct-abliterated-GGUF --include "Llama-3.3-70B-Instruct-abliterated-Q8_0/*" --local-dir ./

You can either specify a new local-dir (Llama-3.3-70B-Instruct-abliterated-Q8_0) or download them all in place (./)

ARM/AVX information

Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.

Now, however, there is something called "online repacking" for weights. details in this PR. If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.

As of llama.cpp build b4282 you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.

Additionally, if you want to get slightly better quality for , you can use IQ4_NL thanks to this PR which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed incrase.

Click to view Q4_0_X_X information (deprecated

I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.

Click to view benchmarks on an AVX2 system (EPYC7702)
model size params backend threads test t/s % (vs Q4_0)
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 pp512 204.03 ± 1.03 100%
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 pp1024 282.92 ± 0.19 100%
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 pp2048 259.49 ± 0.44 100%
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 tg128 39.12 ± 0.27 100%
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 tg256 39.31 ± 0.69 100%
qwen2 3B Q4_0 1.70 GiB 3.09 B CPU 64 tg512 40.52 ± 0.03 100%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 pp512 301.02 ± 1.74 147%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 pp1024 287.23 ± 0.20 101%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 pp2048 262.77 ± 1.81 101%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 tg128 18.80 ± 0.99 48%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 tg256 24.46 ± 3.04 83%
qwen2 3B Q4_K_M 1.79 GiB 3.09 B CPU 64 tg512 36.32 ± 3.59 90%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 pp512 271.71 ± 3.53 133%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 pp1024 279.86 ± 45.63 100%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 pp2048 320.77 ± 5.00 124%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 tg128 43.51 ± 0.05 111%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 tg256 43.35 ± 0.09 110%
qwen2 3B Q4_0_8_8 1.69 GiB 3.09 B CPU 64 tg512 42.60 ± 0.31 105%

Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation

Which file should I choose?

Click here for details

A great write up with charts showing various performances is provided by Artefact2 here

The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.

If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.

If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.

Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.

If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.

If you want to get more into the weeds, you can check out this extremely useful feature chart:

llama.cpp feature matrix

But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.

These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.

The I-quants are not compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.

Thank you ZeroWw for the inspiration to experiment with embed/output.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-12-24Update metadata with huggingface_hubce37d9930 KB
    Loading...
  2. 2024-12-24Upload README.md with huggingface_hub7e3bd4f14.7 KB
    Loading...

Discussions 1 thread

  1. 2026-04-09good stuffopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration