← back to catalog · registered 2026-08-22 13:56

FunnyPunch/LLAMA-3_8B_Unaligned_BETA_Long_Context

FunnyPunch Llama GGUF
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/FunnyPunch%2FLLAMA-3_8B_Unaligned_BETA_Long_Context"
Response includes
  • classification unknown
  • files 10
  • benchmarks 5 entries
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
161
↑ 481% in 90 days
Likes
0
Model age
7mo ago
created 2026-02-24
Downloads over time
Now761→from131↑481%
100341583824131 on Feb 25761 on Oct 11761 on Oct 10FebAprJunAugOct
Feb 25 → Oct 11 · 72 snapshots · spans 228 days

Benchmarks

Benchmark Score Source
BBH average 0.43844856661045534 OpenLLM-v2
IFEval instruct 0.43764988009592326 OpenLLM-v2
IFEval-Prompt 0.3049907578558225 OpenLLM-v2
MATH lvl 5 0.07854984894259819 OpenLLM-v2
MMLU-Pro 0.3464926861702128 OpenLLM-v2

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.1
Tags
gguf base_model:SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA base_model:quantized:SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA license:llama3.1 endpoints_compatible region:us conversational

Related

Total size
42.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-02-24 23:32

Files by quantization

Auxiliary files 10 files 42.4 GB
LLAMA_3_8B_UNALIGNED_BETA-q8_0-131k.gguf 7.95 GB 17fea9fb download
LLAMA_3_8B_UNALIGNED_BETA-q6_k-131k.gguf 6.14 GB d666a1e4 download
LLAMA_3_8B_UNALIGNED_BETA-q5_k_m-131k.gguf 5.34 GB 65a9ddf4 download
LLAMA_3_8B_UNALIGNED_BETA-q5_k_s-131k.gguf 5.21 GB e37b2dd4 download
LLAMA_3_8B_UNALIGNED_BETA-q4_k_m-131k.gguf 4.58 GB d17e5d69 download
LLAMA_3_8B_UNALIGNED_BETA-q4_0-mixed-kvQ5_0.gguf 4.45 GB 83fa78fa download
LLAMA_3_8B_UNALIGNED_BETA-q4_k_s-131k.gguf 4.37 GB 405b6142 download
LLAMA_3_8B_UNALIGNED_BETA-q4_0-mixed.gguf 4.35 GB 62fcc3ad download
README.md 2.65 KB 608f36ca download
.gitattributes 2.10 KB 189b96d0 download

README current version from Hugging Face


license: llama3.1
base_model:

  • SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA

For now this is just a test of various basic techniques of increasing the context window on the best 8B model there is, that has 1 big problem, the limit of context window being just 16384.
I strongly suggest not downloading, or if you do...I guess tell me how bad this is. It's first time I upload anything on HF ever. Or Git. Or do anything on the internet since writing pages in HTTP in 1999.

Go here for original: https://huggingface.co/SicariusSicariiStuff/LLAMA-3_8B_Unaligned_BETA

24.02.2026 - serious tests on Q8_0 started.

I initially had an idea, because I was using Q4_0 and similar on the phone where I have 32768 or 65536 context length set up in Layla. And Layla ignores the context limit unlike LM Studio.
But in LM Studio it wasn't viable, despite the model being extremely fast (few seconds to generate response from Q8_0, regardless how far one was into the chat), because LM Studio forces one to the limit.

Since limit in original Llama_3_8B_Unaligned_BETA is 16384, if you entered in LM Studio manually 32768, you would end up, if you were unlucky with 3276 context length and if you were lucky with 12768 context length.
And in Layla, chats with 30K tokens and still working were normal.

First I tried adapring RoPE from Wingless Imp 8B by Sicarius. However model went a bit nuts. That being said it's unknown if the issue wasn't on LM Studio side (their standard format of prompt is good, but I noticed that for example Impish Nemo hates it).
I'm currently testing Q8_0.

If anyone wants to help, please ask and I will upload one of the standard K quants or ARM quants. No Imatrix though, because I don't have a proper file to generate one.

1st Update: 24/25th.02.2026 - like most Llama 3.1 8B models, the issues seem to start around 40-50K tokens. Prompting has to be way more careful. I didn't try generate super long stories yet, because I don't know how on the software I use (LM Studio and others often have limit of 8192 per message).
Maybe I could try to use llama.cpp directly? However something tells me that results might be mixed. This model has been trained on 16K stories, yes - but that usually means that the model will go off script at 24K at latest.
Maybe some super low temperature and very strict setting, but then it won't be "creative" writing. (Don't get me started on calling writing using AI "creative" - even the best models can abstract for real only a tiny amount. 8B - not really).
I will measure perplexity today, but I need to have access to bigger stories. The biggest contiguous one that I have is maybe 180K tokens (it's one story. I Could access a bigger one though)

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-24Update README.md93541742.7 KB
    Loading...
  2. 2026-02-24Update README.md258d8291.7 KB
    Loading...
  3. 2026-02-24Update README.mda7b07ef620 B
    Loading...
  4. 2026-02-24Update README.md203b6a2514 B
    Loading...
  5. 2026-02-24Update README.md885d698421 B
    Loading...
  6. 2026-02-24initial commit0200e9d26 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration