← back to catalog · registered 2026-10-04 14:58

TokenTinkerer/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-GGUF

TokenTinkerer 35B GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TokenTinkerer%2FQwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-GGUF"
Response includes
  • classification m-uncensored
  • files 13
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-04

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
IQ3 IQ4 Q3_K Q4_K Q5_K Q6_K
Tags
llama.cpp gguf quantized imatrix moe uncensored qwen3.6 custom-quants text-generation en zh base_model:HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Related

Total size
188 GB
Files
13
Quantizations
7
Registered
2026-10-04 14:58
Last updated on HF
2026-10-04 15:51

Files by quantization

Q6_K 1 file 27.1 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-6.72bpw-Q6_K.gguf 27.1 GB bcd7e605 download
Q5_K 2 files 47.0 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-6.06bpw-Q5_K_M.gguf 24.4 GB 0f3120f2 download
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-5.59bpw-Q5_K_S.gguf 22.5 GB f5f2f475 download
Q4_K 2 files 20.3 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-4.90bpw-Q4_K_M.gguf 19.8 GB a761d832 download
Qwen3.6-35B-A3B-MTP-ONLY-Q4_K_M-shared.gguf 532 MB 70c377fa download
IQ4 2 files 35.1 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-4.51bpw-IQ4_NL.gguf 18.2 GB b2a969dc download
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-4.18bpw-IQ4_XS.gguf 16.9 GB 5419af06 download
Q3_K 2 files 30.7 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-3.93bpw-Q3_K_M.gguf 15.9 GB 77e45d6c download
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-3.67bpw-Q3_K_S.gguf 14.8 GB b2a49dc7 download
IQ3 2 files 27.6 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-3.51bpw-IQ3_XS.gguf 14.2 GB 694d3e29 download
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-TTQ-3.33bpw-IQ3_XXS.gguf 13.5 GB c1e4dc98 download
Auxiliary files 2 files 24.6 KB
README.md 22.0 KB 8db2a8fd download
.gitattributes 2.61 KB 40bc95c4 download

README current version from Hugging Face


base_model: HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
language:

  • en
  • zh
    tags:
  • gguf
  • llama.cpp
  • quantized
  • imatrix
  • moe
  • uncensored
  • qwen3.6
  • custom-quants
    license: apache-2.0

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive — TTQ quant set

TTQ (TokenTinkerer Quant) is a ten-rung GGUF ladder built from
HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive,
from 3.33 bpw to 6.72 bpw. Every rung was measured on the same machine, with the same
protocol, against the same reference — and every claim below is a measurement, not an estimate.

Source caveat — there is no full-precision release of this model. The fine-tune is published
only as GGUF; no F16/BF16 weights exist. The highest-quality source available is therefore
upstream's Q8_K_P (~10 bpw), and it serves as both the requantization source for every rung
here and the reference against which all KLD / RMS figures are measured. Two consequences:
every quality number below is relative to that file rather than to the true full-precision
model, and an unknown fraction of each measured divergence is inherited from it and cannot be
separated without a full-precision source. Comparisons between files remain valid, because they
all share the same reference. The full-precision Qwen/Qwen3.6-35B-A3B base model is not a
substitute — the fine-tune is a different model.

The point of the set is efficiency at a given size. At matched size these rungs beat the
available _P quants from the base repository on all four axes measured: KLD, RMS Δp, decode
speed, and prefill. A few examples, all with significance shown:

this set vs base repo file KLD size decode prefill RMS Δp
6.72 bpw Q6_K (27.10 GiB) Q6_K_P (28.54 GiB) 1.61× better (5.1σ) −1.44 GiB +11% +50% better
6.06 bpw Q5_K_M (24.45 GiB) Q6_K_P 1.23× better (2.5σ) −4.09 GiB +19% +80% better
5.59 bpw Q5_K_S (22.54 GiB) Q5_K_P (26.10 GiB) 1.28× better (3.9σ) −3.56 GiB +6% +28% better
4.90 bpw Q4_K_M (19.76 GiB) Q5_K_P a tie (0.65σ) −6.34 GiB +16% +45% better

The last line is the shape of the whole set: where it does not win on quality, it wins by being
dramatically smaller at the same quality.


The ladder

KLD is KL divergence against the model's Q8_K_P logits, measured over 40 fixed chunks of an
internal evaluation set (lower is better; ± is the standard error). RMS Δp is the root-mean-square
difference in the probability assigned to the actual next token, so it reads directly as
percentage points. The corpus and chunking are identical for every rung here and for both
reference files, which is what makes the columns comparable. tg128 is llama.cpp llama-bench
decode at -ncmoe 40 (all experts on CPU), -p 512 -n 128 -r 5 — see
Speed before comparing it to anything.

file suffix declared type GiB bpw KLD ± PPL(Q) RMS Δp tg128
TTQ-3.33bpw-IQ3_XXS IQ3_XXS (23) 13.46 3.33 0.060539 ± 0.001366 6.060 7.339% 27.07
TTQ-3.51bpw-IQ3_XS IQ3_XS (22) 14.16 3.51 0.051261 ± 0.001194 6.094 6.757% 26.30
TTQ-3.67bpw-Q3_K_S Q3_K_S (11) 14.84 3.67 0.042459 ± 0.001034 5.960 5.924% 36.28
TTQ-3.93bpw-Q3_K_M Q3_K_M (12) 15.85 3.93 0.034379 ± 0.000873 5.959 5.205% 37.41
TTQ-4.18bpw-IQ4_XS IQ4_XS (30) 16.90 4.19 0.025999 ± 0.000740 5.944 4.520% 37.22
TTQ-4.51bpw-IQ4_NL IQ4_NL (25) 18.20 4.51 0.018427 ± 0.000634 5.896 3.789% 53.86
TTQ-4.90bpw-Q4_K_M Q4_K_M (15) 19.76 4.90 0.015152 ± 0.000886 5.902 3.415% 53.36
TTQ-5.59bpw-Q5_K_S Q5_K_S (16) 22.54 5.59 0.011316 ± 0.000672 5.879 3.115% 48.65
TTQ-6.06bpw-Q5_K_M Q5_K_M (17) 24.45 6.06 0.007621 ± 0.000252 5.868 2.501% 48.38
TTQ-6.72bpw-Q6_K Q6_K (18) 27.10 6.72 0.005800 ± 0.000218 5.872 2.173% 45.46

Full filenames are Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-<suffix>.gguf. The suffix is
the only part that varies.

How to read the names

Two independent fields, and they mean different things:

  • <bpw>bpw is the true measured rate, computed from the finished file's byte count. Trust
    this number.
  • The trailing token is the file's declared general.file_type, and it is deliberately chosen
    to understate the true rate and to be distinct across the whole set. For example
    TTQ-3.33bpw-IQ3_XXS is not a plain IQ3_XXS file — it is 3.33 bpw, comfortably above the
    3.06 bpw that the IQ3_XXS label nominally implies.

That second rule is not cosmetic. If two files in a repository declare the same general.file_type,
the Hub's download selector collapses them and one becomes effectively invisible. Every token here
is distinct and monotone, so all ten rungs stay individually downloadable.


Which rung should I use?

if you want use why
the smallest usable file 3.33 bpw IQ3_XXS (13.46 GiB) 3.33 bpw is the practical floor for this model; below it, quality collapses (see Limits)
best quality per byte below 4 bpw 3.67 bpw Q3_K_S 1.76× better than the base repo's IQ3_M at a comparable size, and 37% faster than the two codebook rungs
the one that also dominates on size 3.51 bpw IQ3_XS 1.46× better than IQ3_M and 0.22 GiB smaller — but it is the slowest file in the set (codebook)
fastest small rung in this column 4.51 bpw IQ4_NL 53.86 tg, statistically tied with the 4.90. This ranking does not survive serving — see the caveat below: in a GPU-resident config the 3.93 rung runs ~21% faster than this one
the best balance 4.90 bpw Q4_K_M ties the base repo's Q5_K_P on KLD at 6.34 GiB smaller, and is 16% faster on decode and 45% faster on prefill
near-Q8 quality, small 6.72 bpw Q6_K (27.10 GiB) best KLD and best RMS in the set, and it beats Q6_K_P on every axis
maximum quality, cost no object Q8_K_P (from the base repo) the reference this whole ladder is measured against

The one non-obvious trap: do not pick by size alone. The two smallest rungs use codebook
expert formats and are the slowest files in the set (26-27 tg), while the two rungs around
18-20 GiB are the fastest (53-54 tg) — nearly 2× quicker than files barely larger than the 3.51.
Details below.


Serving

Stock llama.cpp

llama-server -m TTQ-3.93bpw-Q3_K_M.gguf \
  -c 32768 -ngl 99 -fa on --jinja \
  -ctk q8_0 -ctv q8_0

--jinja makes llama.cpp apply the chat template embedded in the file (Qwen's official one) —
no extra files needed. The whole point of these quants is behaviour parity with the base model, so
the defaults are the right starting point.

Fitting the experts. A 35B MoE at these sizes will not fit its experts in VRAM on a consumer
card, and llama.cpp handles that well: -ncmoe N keeps the experts of the first N layers on the
CPU
and the rest on the GPU. That dial matters more than any other flag:

  • Measured on an 8 GB card, -ncmoe 36 ran ~20% faster than -ncmoe 40 (all experts on CPU).
  • One step further (-ncmoe 34) bought another ~5%, for ~800 MiB of VRAM.
  • Leave 0.5-1 GiB of VRAM free. On Windows/WDDM, running out does not fail cleanly — it
    collapses prefill (measured: 1104 → ~600 t/s). If throughput suddenly drops, check VRAM before
    suspecting the model.
  • The optimum is rung-specific, not transferable: the 4.51 rung has an earlier cliff than the
    3.93. Tune it per rung.

Sampling

What this set is served with:

--temp 0.6  --top-p 0.95  --top-k 20  --min-p 0

That is Qwen's own recommended sampling, verbatim. Their card lists, for thinking mode on
precise coding tasks
: temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0. The two penalties are already off by default in llama.cpp, so the command
above is that preset exactly. Nothing here is a house preference.

For comparison, their card suggests two other presets: temperature=1.0, presence_penalty=1.5 for
thinking mode on general tasks (hotter, more variety), and temperature=0.7, top_p=0.80 for
non-thinking mode. Swap to either if it suits your use; these quants are not tuned to a sampler.

The two benchmark spot-checks below were run at these settings, not greedy — which is why they
carry that caveat. All KLD, PPL and RMS figures in this card are sampler-independent: perplexity
evaluation is a deterministic single pass over fixed text.

buun-llama-cpp

Same command line — the fork is a drop-in — with one addition:

  • --moe-cache, which stock does not have. It caches hot experts in VRAM. Its value is
    rung-dependent — worth roughly +11% decode on the 3.93, roughly 0% on the 4.51. It is
    elastic: the pool is free-VRAM-minus-reserve and shrinks itself under pressure rather than
    failing, so it never blocks anything. Pass an explicit size (--moe-cache 1000); on lets it
    silently shrink other allocations instead of erroring, and off costs nothing on the 4.51 but
    ~11% on the 3.93.

With the MTP head (speculative decoding)

This repository also ships the model's multi-token-prediction head as a sidecar. Attaching it
recovers most of the decode lost to quantization:

llama-server -m TTQ-3.93bpw-Q3_K_M.gguf \
  -c 32768 -ngl 99 -fa on --jinja \
  -md Qwen3.6-35B-A3B-MTP-ONLY-Q4_K_M-shared.gguf \
  -ngld 999 -ctkd q8_0 -ctvd q8_0 \
  --spec-type draft-mtp --spec-draft-n-max 2

Measured, same session, greedy, at 3.93 bpw: 37.30 → 49.35 t/s, +32%. Prefill pays a few percent
for the verify batch.

-ctkd/-ctvd set the head's own KV cache — it is a separate model and needs its own. -ngld
is its layer offload. If your llama.cpp build has turbo4 (a fork KV type) you can use it for both,
which is what the numbers above were measured with; q8_0 is the stock-safe choice.

--spec-draft-n-max 2 measured best. On real workloads — the two benchmark runs below — draft
acceptance came in at 84.5-95%.
A depth sweep on a varied-prompt harness gave 89% → 82% → 74% →
68% from n_max 1 → 4, with decode peaking at 2: higher values trade acceptance for verify-batch
cost.

The MTP head in this repository

Qwen3.6-35B-A3B-MTP-ONLY-Q4_K_M-shared.gguf — 532 MiB, a sidecar, not a quant of the model.

It is stripped: token_embd.weight and output.weight were removed and
qwen35moe.nextn_shared_target_tensors = true added, so the loader borrows the target model's
copies instead of duplicating ~670 MiB of weights. The unstripped head is 1192 MiB. Use it only with
the matching target model and a build that supports --spec-type draft-mtp.

Speed column conditions

Read this before comparing tg128 to any other number, including numbers on other cards.

What was measured. llama-bench -ngl 999 -ncmoe 40 -t 6 -p 512 -n 128 -r 5, i.e. all expert
layers on CPU
, batch 1, no speculative decoding. That isolates the expert kernel, which is what
differs between rungs — but it is not serving throughput. Measured serving numbers with expert
layers moved to GPU and an MTP drafter attached differ substantially, and the ordering itself can
invert
. Same-session 2×2 at a matched config (ncmoe 36, cache 1000, np 1, greedy):

rung drafter off drafter on
3.93 Q3_K_M 37.30 49.35
4.51 IQ4_NL 37.84 45.88

The 4.51's 44% advantage in the llama-bench column does not survive any GPU residency — with
the drafter off the two are a tie, and with it on the 3.93 leads by 8%. The 4.51 also runs
materially closer to the VRAM ceiling (7.7 GiB peak against the 3.93's 6.6-7.4 GiB), and on this
machine how much VRAM is left over is what prevents prefill collapse. Do not read this column as
throughput, and do not choose a rung for serving from it.

The speed column is not monotonic in size, and the reason is the format. Two families are in
play:

family rungs tg128
codebook (IQ3_XXS, IQ3_XS) 3.33, 3.51 24-26
half codebook (20/40 layers) 4.18 36.3
linear (Q3_K, IQ4_NL, Q4_K, Q5_K, Q6_K) 3.67 → 6.72 33-52

A codebook format reconstructs weights by looking up a memory-resident grid; on AVX2 there is no
register-resident equivalent, and the 256-512 entry tables cost more than the bytes they save. So
the 14.16 GiB rung is slower than the 15.85 GiB rung, and the fastest file in the set is the
18.20 GiB one. A card listing only size and KLD would actively mislead here, which is why this
column exists.

Above ~20 GiB the picture is muddier. Within the linear family the rate is roughly 2.3% decode
per GiB — but format choices move it: raising the expert gate/up from Q4_K to Q5_K is
speed-neutral (−0.6% for +1.9 GiB), while Q8_0 on expert tensors costs 7.7% per GiB. There is no
single rate.

These numbers were measured with llama-bench's own short, highly repetitive prompt. That
matters: with repetitive text, the same experts are selected every token, so the CPU's L3 cache
actually helps and the all-CPU configuration is flattered. Giving the same model a realistic,
varied prompt costs about 30% of decode in a server setting. Treat tg128 as a comparative
expert-kernel measurement and nothing more.


How these were made

  • Requantized from upstream's Q8_K_P (see the source caveat above) using the buun-llama-cpp
    fork and the imatrix shipped alongside the base model.
  • Chat template: every file embeds Qwen's official template at tokenizer.chat_template
    (7,764 characters, sha256 prefix e84f32a2), inherited unchanged from the upstream source and
    verified identical to the template published in Qwen/Qwen3.6-35B-A3B. No serving-side template
    is baked in, so these files behave as the base model does out of the box. If you want to
    reproduce the author's serving setup, it uses a separate community template; pass it explicitly
    with --chat-template-file.
  • Expert tensors are mixed per-layer and per-tensor; the non-expert tensors (attention, SSM,
    shared-expert, embeddings, output) are held at Q8_0 in every rung above 5.59 bpw, which is
    why those rungs carry a higher byte count than a stock recipe of the same nominal type.
  • The output head is kept at Q8_0 even in the smallest rungs where the embeddings drop to Q4_K.
    It is the last projection before the logits that KLD is measured on, and protecting it is cheap.
  • The declared general.file_type is patched after each build to keep the token set distinct and
    monotone (see How to read the names).

Measurement protocol

Everything is measured locally, on the same machine, with the same harness:

  • KLD / PPL / RMS: llama-perplexity --kl-divergence against a stored Q8_K_P reference,
    40 fixed chunks of an internal evaluation set at -c 512, -ngl 999 -ncmoe 40 -t 6 --no-warmup.
    Reported ± is the standard error over chunks. The reference logits are derived from the same
    Q8_K_P file that every rung is requantized from.
  • Speed: llama-bench, as above. Both columns are taken from a second consecutive
    invocation.
    The first is cold, and understates prefill by up to 1.58× and decode by about 12%
    — an artifact of the page cache rather than a property of the model, and the single easiest way
    to produce a misleading GGUF comparison.
  • Significance is quoted as |Δ| / sqrt(σ₁² + σ₂²).

Hardware: RTX 3060 Ti (8 GB), Ryzen 5 5600X, 32 GB DDR4-2133, Windows. Speed numbers in particular
are hardware-specific; the KLD, PPL and RMS numbers are hardware-independent and can be
compared directly with any other measurement made with the same reference and chunk count.


Comparison against the base repository's quants

All files below were measured by the same protocol on the same machine, so the columns are directly
comparable. _P files and IQ3_M/IQ4_XS/Q4_K_M are from the base repository.

file GiB bpw KLD ± PPL(Q) RMS Δp
Q2_K_P 13.95 3.46 0.140590 ± 0.003094 6.456 10.855%
TTQ-3.33bpw-IQ3_XXS 13.46 3.33 0.060539 6.060 7.339%
IQ3_M 14.38 3.56 0.074836 ± 0.001481 6.127 7.990%
TTQ-3.51bpw-IQ3_XS 14.16 3.51 0.051261 6.094 6.757%
TTQ-3.67bpw-Q3_K_S 14.84 3.67 0.042459 5.960 5.924%
TTQ-3.93bpw-Q3_K_M 15.85 3.93 0.034379 5.959 5.205%
TTQ-4.18bpw-IQ4_XS 16.90 4.19 0.025999 5.944 4.520%
IQ4_XS 17.44 4.32 0.032877 ± 0.000864 6.010 5.253%
Q3_K_P 17.72 4.39 0.064623 ± 0.001405 6.085 7.055%
TTQ-4.51bpw-IQ4_NL 18.20 4.51 0.018427 5.896 3.789%
TTQ-4.90bpw-Q4_K_M 19.76 4.90 0.015152 5.902 3.415%
Q4_K_M 19.71 4.88 0.028658 ± 0.001102 5.928 4.709%
Q4_K_P 21.82 5.41 0.024974 ± 0.001055 5.920 4.359%
TTQ-5.59bpw-Q5_K_S 22.54 5.59 0.011316 5.879 3.115%
TTQ-6.06bpw-Q5_K_M 24.45 6.06 0.007621 5.868 2.501%
Q5_K_P 26.10 6.47 0.014504 ± 0.000463 5.884 3.511%
TTQ-6.72bpw-Q6_K 27.10 6.72 0.005800 5.872 2.173%
Q6_K_P 28.54 7.07 0.009366 ± 0.000664 5.863 2.749%
Q8_K_P (reference) 40.61 10.06 — — —

Read that table by size, not by row order. The pattern is consistent: for every file in this set,
the base repository's nearest-size quant is beaten — by 1.26× to 2.50×, median about 1.9× — and
the nearest-quality upstream file is 1.4 to 7.9 GiB larger.


Benchmark spot-checks

Run through the serving recipe — the launcher config and sampling these rungs are meant to be
used with — not through the all-CPU-expert config used for the speed column. More will be added as
they are run.

model benchmark score reference notes
TTQ-3.93bpw-Q3_K_M GPQA Diamond, 198 q 84.8 official 86.0 — ~0.3σ, indistinguishable answer-match, serving sampling
TTQ-3.93bpw-Q3_K_M HumanEval, 164 problems 97.0 no official figure published for this model pass@1, serving sampling

GPQA Diamond is the only one of these with a published comparison, and the two figures are
statistically indistinguishable: at n=198 and p ≈ 0.85 the binomial standard error is ±2.5
points
, so the 1.2-point gap is roughly 0.3σ. Resolving a difference that small at 95%
confidence would need on the order of 3,400 questions. HumanEval gives no point of comparison
here — it is not on the base model's card — so treat 97.0 as an absolute result, with a ±1.3-point
binomial standard error at n=164.

What these do and do not show. They show no measurable collapse in quality: a 15.85 GiB file
at 3.93 bpw answers graduate-level science at parity with the full-precision model and solves 159 of
164 Python problems. They do not give a precise retention figure, and neither benchmark could
detect a 1-2 point loss. Two larger packs (MMLU-Redux, 5,330 questions; MMLU-Pro, 12,032) have
±0.35-point standard errors and would resolve what these cannot; those runs are measured in days,
not hours.

Two caveats apply. The official figure comes from Qwen's own harness, whose protocol for these
benchmarks is not published, so the comparison is approximate. And the sampling used here is the
serving recipe's, which is not greedy.


Limits and things that did not work

Recorded because negative results are results, and because several of them contradict intuitions
that this project started with.

3.33 bpw is the floor. Below it, both available routes were built and measured:

  • a 2.75 bpw codebook rung (IQ2_S gate/up + IQ2_XS down) scored KLD 0.135076 — a statistical
    tie with Q2_K_P, the worst file in the table, at 2.84 GiB smaller. Not worth publishing.
  • the linear route (Q2_K/Q3_K at the bottom) is dominated for the same reason.

The reason is measurable: the marginal KLD cost per GiB rises as the model shrinks. Dropping
from 3.33 to 2.75 bpw costs 0.0318 KLD/GiB, roughly 2.4× the 0.0133 KLD/GiB of the 3.33→3.51 step.
Bytes buy quality cheaply at the top of the ladder and expensively at the bottom.

Allocation shape barely matters. Three 5.58 bpw rungs were built to the same size with the
expert bits placed very differently — gate-heavy, down-heavy, and with the non-expert reduced to
fund more gate. They came out statistically indistinguishable (0.7-1.05σ). A fourth rung that spent
1.46 GiB on Q8_0 expert gates gained no KLD at all while losing 11% of decode. Design has to be
justified by size, format family and simplicity — not by where the bits sit.

Two obvious-looking allocation principles were tested and failed: "put the bits on gate/up and
take them from down" and "the non-expert is a poor place to spend". Both were contradicted at the
size where they were tested. They may hold elsewhere; they are not stated here as rules.

The speed column is a comparative kernel measurement, as explained above — not throughput, and
its rung ordering is not necessarily the ordering you will see in a server.


Credits and licence

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration