← back to catalog · registered 2026-09-11 04:55

Solstice-AI/Qwen3.8-27B-TTURBO-Cold-Fusion-709-L-Uncensored-UltraOptimised-VariableThinking

Solstice-AI 27B GGUF multimodal second-order
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-11

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q4_K Q5_K Q6_K Q8_0
Tags
gguf solstice-ai davidau qwen qwen3.8 qwen3.8-27b cold-fusion gain project-heretic heretic uncensored abliterated

Related

Total size
87.5 GB
Files
11
Quantizations
6
Registered
2026-09-11 04:55
Last updated on HF
2026-09-11 04:34

Files by quantization

Q8_0 1 file 28.2 GB
Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q8_0.gguf 28.2 GB 4d8f97e8 download
Q6_K 1 file 22.4 GB
Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q6_K.gguf 22.4 GB e518b4c0 download
Q5_K 1 file 19.7 GB
Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q5_K_M.gguf 19.7 GB b7519bac download
Q4_K 1 file 17.2 GB
Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q4_K_M.gguf 17.2 GB 0d659ff8 download
BF16 1 file 888 MB
mmproj-BF16.gguf 888 MB a7fefa00 download
Auxiliary files 6 files 18.7 MB
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
tokenizer_config.json 31.3 KB ae7250a0 download
chat_template.jinja 23.8 KB ec69ba46 download
README.md 7.62 KB aaaf1e59 download
.gitattributes 4.07 KB 82f6fcd9 download

README current version from Hugging Face


language:

  • en
  • zh
    license: apache-2.0
    base_model: DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored
    tags:
  • solstice-ai
  • davidau
  • qwen
  • qwen3.8
  • qwen3.8-27b
  • cold-fusion
  • gain
  • project-heretic
  • heretic
  • uncensored
  • abliterated
  • fable
  • cot
  • reasoning
  • coding
  • gguf
  • llama.cpp
  • ollama
  • mtp
  • multi-token-prediction
  • speculative-decoding
  • vision
  • multimodal
  • mmproj
  • q8_0
  • q6_k
  • q5_k_m
  • q4_k_m
  • arc-challenge
  • 709-arc
    pipeline_tag: image-text-to-text
    datasets:
  • Solstice-AI/Solace-1.0-Omni

Solstice-AI Banner

Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored (GGUF UltraOptimised)

Official Solstice-AI Quantization Suite • Hardware Multi-Token Prediction (MTP) • 10-Level Cognitive Architecture • Twin-Turbo GAIN

Original Model & GAIN Merge by DavidAU • Curated Quantization, MTP Integration & Cognitive Architecture by Solstice-AI

Solstice-AI License Format Hardware MTP ARC-C 10-Level Spectrum


Executive Summary

Solstice-AI/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-GGUF-UltraOptimised-MTP-Curated is the curated, zero-bloat GGUF release of DavidAU's flagship Qwen3.8-27B Twin Turbo Cold Fusion foundation (DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored).

This curated release eliminates non-MAX duplicates, degraded extreme low-bits, and external drafters in favor of pure, Pareto-optimal checkpoints with Native Hardware Multi-Token Prediction (MTP) and Solstice-AI's 10-Level Cognitive Reasoning Architecture.


Curated "GOATed" Checkpoints

Every checkpoint in this suite is a MAX-MTP tier: retaining the critical output tensor (output.weight / lm_head) in unquantized 16-bit precision alongside Q8_0 MTP heads to preserve reasoning accuracy (Arc-C 701+ at 4-bit):

Checkpoint File Size VRAM Budget Optimal Target
...-MAX-MTP-Q4_K_M.gguf 17.23 GB 16 GB VRAM The GOAT 4-bit (16-bit lm_head, Arc-C 701, ideal for RTX 4080 / 16GB GPUs)
...-MAX-MTP-Q5_K_M.gguf 19.73 GB 24 GB VRAM The GOAT 5-bit (sweet spot of generation speed & reasoning depth)
...-MAX-MTP-Q6_K.gguf 22.38 GB 24–32 GB VRAM High-fidelity near-lossless sweet spot for RTX 3090/4090 & Apple Silicon
...-MAX-MTP-Q8_0.gguf 28.16 GB 32 GB+ VRAM Full reference precision
mmproj-BF16.gguf 0.87 GB System RAM/VRAM Spatial-temporal multimodal vision projector (images & video frames)

10-Level Cognitive Reasoning Architecture

Built directly into tokenizer_config.json and chat_template.jinja, this suite introduces a 10-level cognitive spectrum. Levels feature soft-elastic pacing (thoughts scale organically to problem difficulty without artificial token caps).

Triggering Modes In-Chat & Via API

  • In-Chat Message Tags (works across Ollama, LM Studio, OpenWebUI, LibreChat):
    • Thinking Mode: Add {REASON:<alias>} anywhere in your message (e.g., {REASON:amax}, {REASON:uhigh}, {REASON:athena}). The tag is stripped from the prompt and persists across subsequent chat turns.
    • Instant Instruct Mode (Zero Reasoning Tokens): Prefix with i (e.g., {REASON:iamax}, {REASON:iuhigh}, {REASON:iathena}) to close <think></think> immediately and generate a direct answer framed through that persona.
  • API Parameters:
    # Thinking Mode
    response = client.chat.completions.create(
        model="...",
        messages=[{"role": "user", "content": "Analyze system architecture"}],
        extra_body={"chat_template_kwargs": {"reasoning_effort": "amax"}}
    )
    
    # Instant Instruct (0 Thinking Tokens)
    response = client.chat.completions.create(
        model="...",
        messages=[{"role": "user", "content": "Fast code generation"}],
        extra_body={"chat_template_kwargs": {"enable_thinking": False, "reasoning_effort": "uhigh"}}
    )
    

The Cognitive Spectrum

Level Primary Key Technical Aliases Mythological Alias Cognitive Framework & Behavior
0 disabled none, off, direct Mortal 0 tokens: <think></think> closed immediately for instant direct output.
1 ulow ultra-low, micro Hermes Rapid instinct & sanity check. Direct path from premise to verdict (<150 tokens).
2 low compact, fast Apollo Crisp logic and premise validation with zero cognitive overhead.
3 lmed low-medium, targeted Artemis Boundary hunter: tests zero conditions, nulls, and hidden edge cases.
4 medium med, balanced Athena Strategic balance: evaluates architectural trade-offs and structural cohesion.
5 mhigh medium-high, architect Prometheus Proactive forethought: models 10x/100x scale, failure modes, and fault tolerance.
6 high deep, thorough Solstice Deep systemic derivation: multi-branch hypothesis trace and red-team falsification.
7 xhigh extreme-high (Default) Hyperion Native Qwen 3.8 continuous derivation and exhaustive semantic deconstruction.
8 uhigh ultra-high, swarm Einstein 20-Agent Swarm: Deploys 20 virtual perspective agents across Sternberg styles.
9 amax absolute-max, deep-research Oracle Deep Research Council: 1–5 complexity scaling, multi-expert panel & audit matrix.

Quickstart

Native MTP Speculative Decoding via llama.cpp

Checkpoints with -MTP- feature native dual-stream token prediction built into the weights (no external drafter file needed):

# High-speed interactive chat with native MTP
llama-cli \
  --hf-repo Solstice-AI/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-GGUF-UltraOptimised-MTP-Curated \
  --hf-file Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q4_K_M.gguf \
  --mmproj mmproj-BF16.gguf \
  -c 131072 \
  -ngl 99 \
  -p "{REASON:amax} Perform a rigorous architectural evaluation of microservices vs monoliths."

Server Deployment

llama-server \
  --hf-repo Solstice-AI/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-GGUF-UltraOptimised-MTP-Curated \
  --hf-file Qwen3.8-27B-TTURBO-Fable-C-Fusion-709-L-Uncen-NM-DAU-NEO-MAX-MTP-Q4_K_M.gguf \
  --mmproj mmproj-BF16.gguf \
  --host 0.0.0.0 --port 8080 -c 131072 -ngl 99

Citations & Acknowledgments

  • DavidAU for the phenomenal Qwen3.8-27B Twin-Turbo Cold Fusion GAIN merged base foundation.
  • Qwen Team for the foundational Qwen 3.8 architecture.
  • Solstice-AI for downstream curated quantization, MTP packaging, and the 10-level cognitive architecture.

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-11Update README: scrub DSpark, document Hardware MTP & 10-level cognitive archi...2b101767.6 KB
    Loading...
  2. 2026-09-11Update README.md5b688785.8 KB
    Loading...
  3. 2026-09-11Update README.md37be18e6.1 KB
    Loading...
  4. 2026-09-11Update README.mdc64fac46.2 KB
    Loading...
  5. 2026-09-11Update README.md61d63876.2 KB
    Loading...
  6. 2026-09-11Update README.mdb1d8c626.1 KB
    Loading...
  7. 2026-09-11Update README.md7e584466 KB
    Loading...
  8. 2026-09-11Update README.md237bbb26 KB
    Loading...
  9. 2026-09-11Apply official Solstice-AI Ultra-Optimised model card & benchmark brandingb5f1c706.8 KB
    Loading...
  10. 2026-09-11Update official Solstice-AI documentation with MTP & DSpark drafter guide1cd1ddd6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.