← back to catalog · registered 2026-09-24 03:57

SC117/occamy-1.0-abliterated-FIT-GGUF

SC117 GGUF MoE multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2Foccamy-1.0-abliterated-FIT-GGUF"
Response includes
  • classification m8
  • files 6
  • author_summary 16 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-24

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh es fr de pt it ru ja ko vi th ar
Tags
gguf llama-cpp quantization moe mixture-of-experts qwen35moe fit-gguf imatrix abliterated uncensored multimodal vision

Related

Total size
183 MB
Files
6
Quantizations
2
Registered
2026-09-24 03:57
Last updated on HF
2026-09-24 04:10

Files by quantization

F16 1 file 858 MB
mmproj-occamy-1.0-abliterated-F16.gguf 858 MB ae1c4045 download
Auxiliary files 5 files 183 MB
occamy-1.0-abliterated-BF16-imatrix.gguf 183 MB 08a43b20 download
README.md 31.7 KB 626d2a48 download
README.zh-CN.md 30.8 KB 7bdb9820 download
SHA256SUMS.txt 1.89 KB 80d1f773 download
.gitattributes 1.77 KB c1b70340 download

README current version from Hugging Face


base_model: Accio-Lab/occamy-1.0
base_model_relation: quantized
license: apache-2.0
library_name: gguf
pipeline_tag: image-text-to-text
language:

  • en
  • zh
  • es
  • fr
  • de
  • pt
  • it
  • ru
  • ja
  • ko
  • vi
  • th
  • ar
    tags:
  • gguf
  • llama-cpp
  • quantization
  • moe
  • mixture-of-experts
  • qwen35moe
  • fit-gguf
  • imatrix
  • abliterated
  • uncensored
  • multimodal
  • vision
  • reasoning
  • long-context
  • agent

FIT-GGUF v0.4.0GATE-VERIFIED TIERS5 FIT + 4 APEX-I- + BF16ABLITERATED (INPUT-GATED)MEASURED KL + SAME-TOPVISION mmprojAPACHE-2.0

Occamy-1.0-abliterated · FIT-GGUF

Five fidelity tiers of a 256K-context 35B MoE with vision — the smallest GGUF that provably meets each KL gate, re-verified on its own bytes. Plus an independent APEX-I- family and the BF16 reference, all measured by the same protocol.

10.94 GiB · MINIverified minimum at each fidelity64.61 GiB · BF16

English · 简体中文 📖

🧭 About FIT-GGUF — the tool behind these files

Every FIT tier here was planned, executed and verified by FIT-GGUF, an open-source, deterministic tensor-level planning layer on top of llama.cpp quantization. Standard GGUF quantization asks you to pick one of a handful of presets; FIT-GGUF instead asks what quality do you want, then finds and verifies the smallest GGUF that demonstrably meets it.

Traditional GGUF gives you presets. FIT gives you a fidelity contract: macro KL ≤ tier anchor, measured against this model's own BF16 on five fixed domains.

Deterministic size prediction & byte-exact delivery✅ Validated (G2 gate, delta = 0)
Universally optimal tensor allocation⚠️ Not established — FIT claims verified fidelity contracts, not a universal quality optimum

Method, pre-registered research record and the fit CLI are all open source: github.com/Scorp1o117/FIT-GGUF

⚠️ Safety notice / 安全提示

The source model is an abliterated, refusal-removed model with no meaningful built-in guardrails, and may comply with harmful, illegal or unsafe requests. Use it only where you can provide appropriate moderation, access control and legal review. Do not deploy it to end users without your own safety layer.

源模型经过拒答方向消融,不具备可靠的内置安全护栏;请仅在合法、受控、具备审核与滥用防护的环境中使用,使用者自行承担部署责任。

🧬 Abliteration — input-gated, documented

The refusal direction was ablated locally with the abliterix pipeline using input-gated ablation. A classic rank-1 edit is ΔW = −w · v (vᵀW) — the same output direction v for every input, harmless ones included, which is where the KL bill comes from. The gated form replaces the input-side vector with a per-expert direction g chosen to maximise refusal suppression per unit of perturbation, g ∝ S_h⁻¹ μ_r: the harmful-prompt mean whitened by the harmless second-moment matrix, so the edit lands where harmless traffic does not live. 2,160 expert deltas were applied across 40 layers, and the export is re-derived from the original checkpoint plus those deltas and compared byte-for-byte.

MetricDelivered
Refusals (100 held-out harmful prompts)10 (99 before ablation)
Ablation perturbation, 3-token full-distribution KL0.0844
Previous best on this model0.2196 — the gated direction cut it by 62%

Two different KLs — do not mix them. The 0.0844 above is the abliteration's perturbation of the full-precision model. Every other KL number on this page is quantization divergence against this repository's BF16 file. They are measured on different objects and are not comparable.

📦 Everything in this repository

Five FIT tiers, four APEX-I- tiers and the BF16 reference — every file measured by the same evaluator, on the same five slices, against the same reference logits, so the columns are directly comparable across families.

FamilyFileGiBMacro KL ↓Same-top ↑Role / gate
FIT REFERENCE…-FIT-REFERENCE-24G-Q5_K_M.gguf23.550.019595.40%KL ≤ 0.02 · near-lossless
FIT QUALITY…-FIT-QUALITY-17G-IQ4_XS.gguf17.500.046492.26%KL ≤ 0.05 · best all-round
FIT BALANCED…-FIT-BALANCED-14G-IQ3_S.gguf13.800.095988.83%KL ≤ 0.10 · 16 GB cards
FIT COMPACT…-FIT-COMPACT-13G-IQ3_XXS.gguf12.760.135886.15%KL ≤ 0.15 · tighter VRAM
FIT MINI…-FIT-MINI-11G-IQ2_S.gguf10.940.199283.18%KL ≤ 0.20 · smallest usable
APEX-I-balanced…-APEX-I-BALANCED-Q6_K.gguf23.590.019795.04%Q6_K base · FIT REFERENCE is 0.04 GiB smaller and cleaner
APEX-I-quality…-APEX-I-QUALITY-Q6_K.gguf21.250.024094.46%Q6_K base · FIT QUALITY is 3.75 GiB smaller and cleaner
APEX-I-compact…-APEX-I-COMPACT-Q4_K_M.gguf15.400.070090.34%Q4_K_M base · FIT BALANCED is 1.60 GiB smaller and cleaner
APEX-I-mini…-APEX-I-MINI-Q3_K_M.gguf12.540.165484.72%Q3_K_M base · FIT MINI is 1.60 GiB smaller and cleaner
BF16…-BF16.gguf64.610100%The reference every number above is measured against
imatrix…-BF16-imatrix.gguf0.18500×512 chunks, 510 entries — re-quantize this model yourself
mmprojmmproj-…-F16.gguf0.84Vision projector, F16 — pair it with any tier above

File size ≠ RAM/VRAM usage. KV cache and compute buffers are separate — and 30 of this model's 40 layers are linear-attention/DeltaNet, so only 10 expand the cache. Pick a sane -c.

The FIT tiers are 78.55 GiB where the smallest passing standard preset in each would be 89.89 GiB11.34 GiB (12.6%) smaller. APEX-I- is an independent, hand-written per-tensor recipe family kept here as a cross-check rather than a baseline: FIT wins every comparable gate, but two different searches landing within 0.04 GiB of each other at the 0.02 anchor is a tie, not a rout.

📈 Measured quality
Macro KL against main GGUF size: native presets, the five FIT tiers, and the four APEX-I- tiers

llama.cpp b10666 · c=512 b=512 · five domains · vs the BF16 in this repo · 中文大图

The native ladder is the reference curve; the FIT tiers sit below and to the left of it at every gate. The dashed APEX-I- line is the independent recipe family from the table above, and it is the reason the 0.02 anchor is worth reading carefully: at that gate the two families are level, while everywhere below it FIT is well clear.

Read the curve, not just the tiers. The gates are deliberately capped at KL 0.02–0.20, so the big native presets are cleaner than every tier — Q6_K at 26.56 GiB is 10% cleaner than REFERENCE. If you want maximum fidelity rather than minimum size, take Q6_K or the BF16 file.

Full-size EN · 中文大图 · raw per-domain JSON

🚀 Run it

This is plain qwen35moe architecture. Unlike models that need a patched llama.cpp, these files load in any reasonably recent upstream build — no PR branch, no custom fork. The measurements here used b10666.

Text only

./llama-server -m occamy-1.0-abliterated-FIT-BALANCED-14G-IQ3_S.gguf -ngl 99 -c 32768

./llama-cli -m occamy-1.0-abliterated-FIT-MINI-11G-IQ2_S.gguf -ngl 99 -c 8192 -p "Hello" -n 256

With vision

./llama-server -m occamy-1.0-abliterated-FIT-QUALITY-17G-IQ4_XS.gguf
--mmproj mmproj-occamy-1.0-abliterated-F16.gguf -ngl 99 -c 32768

The chat template is embedded in every file — including the tool-calling and reasoning-content handling — so LM Studio, KoboldCpp, Jan and friends load them without extra configuration. Native context is 262,144 tokens, but KV cache grows with it: start at -c 32768 and work up.

🔬 Evaluation protocol & honest scope
Runtimellama.cpp b10666 · Linux x86_64 · ROCm 10.0 · AMD Ryzen AI MAX+ 395 (gfx1151)
Command shapellama-perplexity -ngl 99 -t 16 -c 512 -b 512 --kl-divergence …
ReferenceThis model's own BF16 logits (the BF16 GGUF shipped here)
Domainswiki_test · wiki_valid · Chinese · code · agent_chat (five fixed 64 KiB slices, macro mean)
Calibration12-point standard preset ladder plus gap probes; tier search by size bisection; guard profile scope exact_model
Weights bindingsource_weights_sha256 = b77f1175…f78856 — the registry entry is keyed to these exact weights

Tier verification: a FIT tier ships only if macro KL ≤ its anchor. Each shipped file was re-evaluated on its own bytes after quantization and reproduced its search-time KL exactly (G2 delta +0 on all five — the delivered byte count equals the re-finalized prediction to the byte).

Allocator scope: the stock balanced v0.3 policy with precision floors applied where the window's candidate set cannot reach a tensor, using this model's own importance matrix. No model-specific refine profile. What is claimed is deterministic size planning plus measured verification of these specific artifacts — not a universally optimal allocation.

The FIT tiers are deliberately not policy-uniform, and the ledger records which planning policy produced each one. A tier's product is the smallest artifact that reaches its anchor and can be rebuilt from the bundle; that artifact does not care which policy planned it. Policy governs where the search looks next, and the bracket is filtered by it — filtering the selection by policy would have thrown away smaller passing artifacts already in the ledger.

Sizes in the file names are measured bytes, not budgets. The tool refuses to emit an artifact whose size it cannot predict to the byte.

🧩 Included — and not included

✅ 5 gate-verified FIT tiers · ✅ 4 APEX-I- tiers for the cross-check · ✅ BF16 reference GGUF (the exact weights every measurement above is taken against) · ✅ the importance matrix used for every quantization · ✅ a vision projector converted from the base model's visual tower · ✅ chat template embedded in every file · ✅ checksums (SHA256SUMS.txt) · ✅ labelled quality curves and the raw per-domain JSON (results/)

❌ no 2-bit-and-below classes — measured collapse on this model, excluded by design · ❌ no MTP head · ❌ no GGUF above Q6_K for the text weights; take the BF16 file if you need the exact weights

The abliteration was performed locally from Accio-Lab/occamy-1.0; this repository contributes the abliterated weights' quantization plans, the artifacts, the vision projector and the measurements.

🔍 Verify & reproduce

sha256sum -c SHA256SUMS.txt

The evaluation slices are public in the FIT-GGUF repository; the calibration ladder, gap probes, search audit, the per-tier recipes and the guard profile for this release are retained in the FIT-GGUF experiment record experiments/2026-09-22-occamy-1p0-abliterated-4tier/. The five-domain .kld reference logits are not stored there — they regenerate from the BF16 GGUF shipped here and are checked against reference-manifest.json, which pins each domain's hash and byte count. Because the importance-matrix path is embedded in the GGUF metadata, the byte count of a re-quantization depends on where you put the imatrix file — the weights do not. Exact-size behaviour is scoped to the recorded source metadata and the recorded llama.cpp build; changing the converter, runtime, source layout or metadata requires revalidation.

📄 License & credits

Apache-2.0, inherited from the base model — follow the upstream license and model-card requirements.

Accio-Lab — the Occamy-1.0 model and its technical report · llama.cpp — quantization and the KL/perplexity evaluator · abliterix — the local input-gated ablation pipeline · FIT-GGUF — verified size-exact quantization. FIT-GGUF is an independent project, not affiliated with Accio-Lab or llama.cpp.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Abliteration, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.