← back to catalog · registered 2026-08-22 13:56

legraphista/Meta-Llama-3-70B-Instruct-abliterated-v3.5-IMat-GGUF

legraphista Llama 70B GGUF second-order 8K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/legraphista%2FMeta-Llama-3-70B-Instruct-abliterated-v3.5-IMat-GGUF"
Response includes
  • classification m8
  • files 24
  • benchmarks 5 entries
  • hub_downloads_all_time 56,920
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
57K
Likes
2
Model age
2.4y ago
created 2024-05-31
Downloads over time
Now106K→from726↑14,496%
038.8K77.7K116.5K726 on Jul 24, 2024106K on Oct 11Jul '24Nov '24Mar '25Jul '25Nov '25MarJul
Jul 24, 2024 → Oct 11 · 156 snapshots · spans 809 days

Benchmarks

Benchmark Score Source
BBH average 0.5126475548060708 OpenLLM-v2
IFEval instruct 0.8081534772182254 OpenLLM-v2
IFEval-Prompt 0.7412199630314233 OpenLLM-v2
MATH lvl 5 0.11858006042296072 OpenLLM-v2
MMLU-Pro 0.44522938829787234 OpenLLM-v2

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3
Quantizations
IQ1 IQ2 IQ3 IQ4 Q2_K Q3_K Q4_K
Tags
gguf quantized GGUF imatrix quantization imat static 16bit 8bit 6bit 5bit 4bit

Related

Total size
514 GB
Files
24
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2024-06-01 02:24

Files by quantization

Q4_K 2 files 77.2 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q4_K.gguf 39.6 GB b3a2c100 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q4_K_S.gguf 37.6 GB 1a89a4de download
IQ4 2 files 72.6 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ4_NL.gguf 37.3 GB 37247653 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ4_XS.gguf 35.3 GB 1d693571 download
Q3_K 3 files 95.3 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K_L.gguf 34.6 GB 13bcdbc4 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K.gguf 31.9 GB fb11a588 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K_S.gguf 28.8 GB e674fa76 download
IQ3 4 files 111 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_M.gguf 29.7 GB 637eb35a download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_S.gguf 28.8 GB 5f9f8e04 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_XS.gguf 27.3 GB 318e4b45 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_XXS.gguf 25.6 GB 3b8226f3 download
Q2_K 2 files 47.4 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q2_K.gguf 24.6 GB eb3857bf download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q2_K_S.gguf 22.8 GB 3e92dc9d download
IQ2 4 files 80.7 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_M.gguf 22.5 GB e52f3c47 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_S.gguf 20.7 GB 3df314c1 download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_XS.gguf 19.7 GB 11852bca download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_XXS.gguf 17.8 GB 20e6d6b0 download
IQ1 2 files 29.9 GB
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ1_M.gguf 15.6 GB 73c3c24f download
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ1_S.gguf 14.3 GB 42731e5d download
Auxiliary files 5 files 24.1 MB
imatrix.dat 23.8 MB 5ba7aa81 download
imatrix.dataset 273 KB 41ae1656 download
README.md 12.2 KB c5b8f717 download
imatrix.log 10.5 KB 42d02760 download
.gitattributes 6.48 KB fb4d1e1e download

README current version from Hugging Face


base_model: failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
inference: false
library_name: gguf
license: llama3
pipeline_tag: text-generation
quantized_by: legraphista
tags:

  • quantized
  • GGUF
  • imatrix
  • quantization
  • imat
  • imatrix
  • static
  • 16bit
  • 8bit
  • 6bit
  • 5bit
  • 4bit
  • 3bit
  • 2bit
  • 1bit

Meta-Llama-3-70B-Instruct-abliterated-v3.5-IMat-GGUF

Llama.cpp imatrix quantization of failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5

Original Model: failspy/Meta-Llama-3-70B-Instruct-abliterated-v3.5
Original dtype: BF16 (bfloat16)
Quantized by: llama.cpp b3058
IMatrix dataset: here


Files

IMatrix

Status: ✅ Available
Link: here

Common Quants

Filename Quant type File Size Status Uses IMatrix Is Split
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0/* Q8_0 74.98GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q6_K/* Q6_K 57.89GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q4_K.gguf Q4_K 42.52GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K.gguf Q3_K 34.27GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q2_K.gguf Q2_K 26.38GB ✅ Available 🟢 IMatrix 📦 No

All Quants

Filename Quant type File Size Status Uses IMatrix Is Split
Meta-Llama-3-70B-Instruct-abliterated-v3.5.BF16/* BF16 141.12GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.FP16/* F16 141.12GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0/* Q8_0 74.98GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q6_K/* Q6_K 57.89GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q5_K/* Q5_K 49.95GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q5_K_S/* Q5_K_S 48.66GB ✅ Available ⚪ Static ✂ Yes
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q4_K.gguf Q4_K 42.52GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q4_K_S.gguf Q4_K_S 40.35GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ4_NL.gguf IQ4_NL 40.05GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ4_XS.gguf IQ4_XS 37.90GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K.gguf Q3_K 34.27GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K_L.gguf Q3_K_L 37.14GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q3_K_S.gguf Q3_K_S 30.91GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_M.gguf IQ3_M 31.94GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_S.gguf IQ3_S 30.91GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_XS.gguf IQ3_XS 29.31GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ3_XXS.gguf IQ3_XXS 27.47GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q2_K.gguf Q2_K 26.38GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q2_K_S.gguf Q2_K_S 24.47GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_M.gguf IQ2_M 24.12GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_S.gguf IQ2_S 22.24GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_XS.gguf IQ2_XS 21.14GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ2_XXS.gguf IQ2_XXS 19.10GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ1_M.gguf IQ1_M 16.75GB ✅ Available 🟢 IMatrix 📦 No
Meta-Llama-3-70B-Instruct-abliterated-v3.5.IQ1_S.gguf IQ1_S 15.34GB ✅ Available 🟢 IMatrix 📦 No

Downloading using huggingface-cli

If you do not have hugginface-cli installed:

pip install -U "huggingface_hub[cli]"

Download the specific file you want:

huggingface-cli download legraphista/Meta-Llama-3-70B-Instruct-abliterated-v3.5-IMat-GGUF --include "Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0.gguf" --local-dir ./

If the model file is big, it has been split into multiple files. In order to download them all to a local folder, run:

huggingface-cli download legraphista/Meta-Llama-3-70B-Instruct-abliterated-v3.5-IMat-GGUF --include "Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0/*" --local-dir ./
# see FAQ for merging GGUF's

Inference

Simple chat template

<|begin_of_text|><|start_header_id|>user<|end_header_id|>

{user_prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

{assistant_response}<|eot_id|><|start_header_id|>user<|end_header_id|>

{next_user_prompt}<|eot_id|>

Chat template with system prompt

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>

{user_prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

{assistant_response}<|eot_id|><|start_header_id|>user<|end_header_id|>

{next_user_prompt}<|eot_id|>

Llama.cpp

llama.cpp/main -m Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0.gguf --color -i -p "prompt here (according to the chat template)"

FAQ

Why is the IMatrix not applied everywhere?

According to this investigation, it appears that lower quantizations are the only ones that benefit from the imatrix input (as per hellaswag results).

How do I merge a split GGUF?

  1. Make sure you have gguf-split available
  2. Locate your GGUF chunks folder (ex: Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0)
  3. Run gguf-split --merge Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0/Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0-00001-of-XXXXX.gguf Meta-Llama-3-70B-Instruct-abliterated-v3.5.Q8_0.gguf
    • Make sure to point gguf-split to the first chunk of the split.

Got a suggestion? Ping me @legraphista!

README history 20 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-06-01Upload README.md with huggingface_hub799805a12.2 KB
    Loading...
  2. 2024-06-01Upload README.md with huggingface_hub557702012 KB
    Loading...
  3. 2024-06-01Upload README.md with huggingface_hubf31918911.9 KB
    Loading...
  4. 2024-06-01Upload README.md with huggingface_huba7d9bcd11.7 KB
    Loading...
  5. 2024-06-01Upload README.md with huggingface_hubf1e774311.5 KB
    Loading...
  6. 2024-06-01Upload README.md with huggingface_hubcf8d3d611.4 KB
    Loading...
  7. 2024-06-01Upload README.md with huggingface_hub540beca11.2 KB
    Loading...
  8. 2024-05-31Upload README.md with huggingface_hubf73eef611 KB
    Loading...
  9. 2024-05-31Upload README.md with huggingface_hub29131ba10.9 KB
    Loading...
  10. 2024-05-31Upload README.md with huggingface_hubfaef34610.7 KB
    Loading...
  11. 2024-05-31Upload README.md with huggingface_hub410e25710.5 KB
    Loading...
  12. 2024-05-31Upload README.md with huggingface_hub651b97910.4 KB
    Loading...
  13. 2024-05-31Upload README.md with huggingface_hub5ef151b10.2 KB
    Loading...
  14. 2024-05-31Upload README.md with huggingface_hubcc5c20e10 KB
    Loading...
  15. 2024-05-31Upload README.md with huggingface_hubb6373a59.9 KB
    Loading...
  16. 2024-05-31Upload README.md with huggingface_huba3a56a79.7 KB
    Loading...
  17. 2024-05-31Upload README.md with huggingface_hub0fbc9819.5 KB
    Loading...
  18. 2024-05-31Upload README.md with huggingface_hub3f2f14d9.2 KB
    Loading...
  19. 2024-05-31Upload README.md with huggingface_hub1931b208.9 KB
    Loading...
  20. 2024-05-31Upload README.md with huggingface_hub4a740998.5 KB
    Loading...

Discussions 2 threads

  1. 2024-07-18Are QK and IQ quantizations made from the F16 or BF16 Gguf?open1 💬#2
    Loading...
  2. 2024-05-31Wow!open3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration