Quantizations for humanists
What Q4_K_M actually means. Why one abliterated model has twelve files. Why 'the model' stops being a single artifact once it exists as a spread of compressed variants. A non-engineer's guide to the cryptic strings on every download page.
A model's weights are billions of numbers. Native precision is 16-bit floating point (F16), so a 70B model needs 140 GB. Quantization is lossy compression: store each number in 4 bits instead of 16, and the 70B drops to 40 GB and runs on a consumer GPU. The names decode: Q4_K_M is '4-bit, K-quant scheme, medium variant'. IQ3_XXS is 'importance-weighted 3-bit, extra-extra-small'. Q4_K_M is the community sweet spot; below Q3 quality degrades noticeably; above Q6 gains stop scaling. When a catalog lists 16,000 abliterated models, it is not indexing 16,000 distinct artifacts - it is indexing a smaller number of edits fanned out into clouds of compressed variants.
- What a model's weights actually are and why size matters
- How quantization trades precision for memory
- How to decode Q4_K_M character by character
- Why one model becomes a family of near-copies
- What K-quants and I-quants each mean
- The philosophical residue: 'the model' as a family, not an artifact
What a model's weights are
A model's "weights" are, physically, an enormous list of numbers - for a mid-sized model, several billion of them. Each number records the strength of one connection learned during training. To run the model is to do arithmetic with all of them for every word produced.
The size of that list depends on how precisely each number is stored. Precision is measured in bits. A model's native precision is usually 16-bit floating point (F16 or BF16), meaning each weight occupies two bytes of memory. That is why a 70-billion-parameter model in its native precision needs roughly 140 gigabytes just to hold the weights - more memory than any consumer graphics card, and more than most single data-center accelerators.
F16 is "native" because it is the format the model was trained and released in. Every compressed copy is measured against it.
What quantization is
Quantization is lossy compression for weights. Instead of two bytes per number, store each in fewer bits - 8, 5, 4, 3, even 2 - accepting a small loss of accuracy in exchange for a large saving in memory. At 4 bits, that 70B model drops to under 40 gigabytes and becomes runnable on a single high-end consumer GPU. A 7B model shrinks from about 13.5 GB at F16 to about 4 GB at Q4_K_M, roughly a 3.3x reduction.
The trade is real: fewer bits mean coarser numbers, and coarser numbers mean the model's outputs drift from what full precision would produce. How much they drift depends on the scheme.
The GGUF format
The dominant format for local inference is GGUF - the letters derive from "GGML Universal Format"; GGML in turn carries the initials of its author, Georgi Gerganov. He created it for the llama.cpp project, and it replaced the earlier GGML format in 2023.
Its design virtue is self-containment: a single .gguf file holds the quantized weights, the tokenizer, and all the metadata needed to run the model, so any compatible program can load it without extra files. This is why abliterated GGUFs are the format most laptop users encounter - LM Studio, Ollama, KoboldCpp, and text-generation-webui all consume the same files.
Reading the cryptic names
The strings on a download page are readable once decoded. Two representative examples:
Q4_K_M
- Q4 - 4-bit quantization. Each weight uses 4 bits instead of 16.
- K - the K-quant scheme, a smarter method that stores weights in blocks with their own scaling factors, giving better quality than the older "0" schemes at the same bit count.
- M - the medium variant, which keeps a handful of especially sensitive weights at higher precision.
So Q4_K_M reads as "4-bit, K-quant, medium."
IQ3_XXS
- I - marks an I-quant, which uses a pre-computed importance matrix (a measurement of which weights matter most, taken by running calibration text through the model) to decide where to spend its scarce bits.
- Q3 - 3-bit quantization.
- XXS - the smallest, most aggressive variant. Stands for "extra-extra-small."
So IQ3_XXS reads as "importance-weighted 3-bit, extra-extra-small."
Rough orientation
This is how to read the options, not a recommendation of which to pick:
- The bit count tracks the memory needed. More bits, more memory, closer to native quality.
- The suffix tracks the quality-for-size trade within a given bit count. S / M / L variants (small / medium / large) preserve progressively more of the sensitive weights.
- Higher bit counts (Q6, Q8) sit closer to native quality and need more memory.
- 4-bit K-quants are the common community compromise - Q4_K_M is what most model cards recommend as the value sweet spot.
- Sub-4-bit I-quants squeeze onto small hardware at rising risk of degradation.
Which one fits depends entirely on the memory of the machine doing the running, which is why catalogs list many. A person with 8 GB of VRAM has fundamentally different options from one with 24 GB or 80 GB, and the same underlying model has to be quantized differently for each.
Why one model has twelve files
This is why a single abliterated model appears on a hosting site five, ten, or fifteen times. Each entry is the same underlying weights compressed to a different level for a different class of hardware. The high-volume quantizers - mradermacher (with over 68,000 model repositories), bartowski - exist to produce these families exhaustively.
A typical release page for a mid-sized abliterated model might carry: Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0, IQ3_XXS, IQ4_XS. Twelve files, one underlying edit. Nothing in the twelve entries adds a new abliteration or teaches the model a new capability. They are twelve compressions of one artifact.
The philosophical residue
What to download (rough guide)
For most users on most hardware, the answer is Q4_K_M. It is the community sweet spot: small enough to run on modest hardware, high enough quality that prose stays coherent, well-tested across model sizes from 7B to 70B.
Rough shape of the decision:
- Have plenty of memory, want highest quality → Q6_K or Q8_0. Above Q6 the perplexity gains stop scaling with disk space, so Q8 is rarely worth the extra memory.
- Standard consumer hardware, want the default → Q4_K_M. The value sweet spot.
- Modest hardware, accept some quality tradeoff → Q4_K_S, or if further squeezed, IQ4_XS or IQ3_M with imatrix calibration.
- Very tight memory, willing to accept clear degradation → IQ3_XXS or Q2_K. Prose degrades visibly here; use for experimentation, not production.
For the technical reasoning behind these choices, and the huihui-ai _L trick that preserves ablation-affected tensors at higher precision, see B9 M8 Repackaging.
Where to go next
Read B9 M8 Repackaging for the technical side of GGUF quantization - the actual commands, the imatrix calibration workflow, the failure modes. Read C2 Reading benchmarks for how quantization interacts with the benchmark numbers a model card reports. And if you are trying to run a specific abliterated model and are unsure which quant to pick, the model card on Hugging Face usually recommends one - trust the recommendation on the specific model over any general rule.
Frequently asked questions
What are a model's weights?
Physically, an enormous list of numbers - billions of them for a mid-sized model. Each records the strength of one connection learned during training. Running the model means doing arithmetic with all of them for every word produced. The size of that list, in bytes, depends on how precisely each number is stored - which is what quantization changes.
What is quantization?
Lossy compression for model weights. Store each weight in fewer bits (4, 5, 8) instead of the native 16, accepting a small loss of accuracy in exchange for a large saving in memory. A 70B model at native F16 is 140 GB; at 4-bit quantization it drops to about 40 GB and becomes runnable on a single high-end consumer GPU. The trade is real - fewer bits mean coarser numbers and the outputs drift from what full precision would produce - but for well-designed schemes at 4+ bits the drift is small.
What is GGUF?
The dominant file format for local model inference. Stands for "GGML Universal Format" - GGML carries the initials of Georgi Gerganov, who created it for the llama.cpp project. Its design virtue is self-containment: one .gguf file holds the quantized weights, the tokenizer, and all metadata needed to run the model, so any compatible program (LM Studio, Ollama, KoboldCpp, text-generation-webui) can load it without extra files.
What does Q4_K_M mean?
"4-bit, K-quant, medium." Q4 = 4 bits per weight. K = the K-quant scheme, which stores weights in blocks with per-block scaling factors (better than older "0" schemes at the same bit count). M = the medium variant, which keeps a handful of especially sensitive weights at higher precision. It is the community sweet spot for 7B-70B models: small enough to run on modest hardware, high enough quality that prose stays coherent.
What does IQ3_XXS mean?
"Importance-weighted 3-bit, extra-extra-small." I = I-quant, which uses a pre-computed importance matrix (measurement of which weights matter most, taken by running calibration text through the model) to decide where to spend its scarce bits. Q3 = 3 bits per weight. XXS = the smallest, most aggressive variant. Used when memory is very tight and some quality degradation is acceptable.
Why does one abliterated model have twelve files?
Each file is the same underlying model at a different quantization level - different memory footprint, different speed, different quality trade-off - so users with different hardware can pick the one that fits. A typical release page might carry Q2_K through Q8_0 plus several I-quants, roughly a dozen files. Nothing about the twelve entries adds a new abliteration or new capability. They are twelve compressions of one artifact.
Which quantization should I download?
For most users on most hardware, Q4_K_M. It is the community sweet spot: small enough to run on modest hardware, high enough quality that prose stays coherent, well-tested across model sizes from 7B to 70B. If you have plenty of VRAM and want maximum quality, Q6_K. If memory is very tight, IQ4_XS or IQ3_M with imatrix calibration. If the model card recommends a specific quant, trust that recommendation over any general rule.
Does the catalog really have 16,000 different models?
Not exactly. It has 16,000+ Hugging Face repositories tagged as abliterated - but many of these are quantization variants of the same underlying edits. When you multiply a few hundred distinct abliteration operations by their typical quantization spreads (often 10-15 variants each) plus their merge descendants, you approach the catalog's headline count without needing that many distinct edits. The 16,000 figure names a landscape of quantized variants, not 16,000 independent minds.
Why do different quants of the same model behave differently?
Because quantization is lossy - the coarser numbers at low bit counts produce slightly different arithmetic than the higher-precision original. The 3-bit version may hesitate where the 8-bit version is fluent, may hallucinate where the other is accurate. Independent testers have documented that quantization can even interact with abliteration, changing measured refusal rates and benchmark scores. The same nominal model at Q3_K_S versus Q6_K can produce measurably different outputs on the same prompt.
References
- llama.cpp repository (Georgi Gerganov, ggml-org). github.com/ggml-org/llama.cpp
- GGUF format documentation. github.com/ggml-org/llama.cpp/blob/master/docs/development/gguf.md
- mradermacher Hugging Face profile (68,000+ GGUF repositories). huggingface.co/mradermacher
- bartowski Hugging Face profile. huggingface.co/bartowski
- llama.cpp perplexity discussions (community source for quant-quality tradeoffs). github.com/ggml-org/llama.cpp/discussions