← back to catalog · registered 2026-08-22 13:56

jcbtc/chadrock3.6-40b-opus-deckard-uncensored-thinking-neo-code-di-imatrix-rocmfp4

jcbtc Qwen 40B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jcbtc%2Fchadrock3.6-40b-opus-deckard-uncensored-thinking-neo-code-di-imatrix-rocmfp4"
Response includes
  • classification m-uncensored
  • files 4
  • hub_downloads_all_time 1,837
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
131 last 30d - cooling
Likes
8
Model age
3mo ago
created 2026-06-15
Downloads over time
Now1.9K→from939↑100%
8921.3K1.6K2K939 on Jun 171.9K on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
llama.cpp gguf qwen3.6 qwen35 40b model-size-40b dense mtp speculative-decoding rocm rocmfp4 chadrock

Related

Total size
19.9 GB
Files
4
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-06-23 01:58

Files by quantization

F16 1 file 885 MB
mmproj-Qwen3.6-27B-f16.mmproj 885 MB 26b00d80 download
Auxiliary files 3 files 19.9 GB
CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf 19.9 GB 252c463f download
README.md 13.8 KB 58b630c8 download
.gitattributes 130 B 05631b37 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
  • PiehSoft/Qwen3.6-40B-Deckard-MTP
    base_model_relation: quantized
    language:
  • en
    library_name: llama.cpp
    pipeline_tag: text-generation
    tags:
  • qwen3.6
  • qwen35
  • 40b
  • model-size-40b
  • dense
  • gguf
  • mtp
  • speculative-decoding
  • rocm
  • rocmfp4
  • chadrock
  • llama.cpp
  • amd
  • ryzen-ai-max-395
  • strix-halo
  • deckard
  • opus
  • uncensored
  • thinking
  • neo-code
  • imatrix

CHADROCK3.6 40B Opus Deckard ROCmFP4

CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4

CHADROCK3.6 40B Opus Deckard is a ROCmFP4/MTP GGUF release of the Deckard 40B lineage for AMD Ryzen AI Max+ 395 / Strix Halo systems.

It keeps the public release identity, file name, and downloadable artifact under the Chadrock name:

CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf

The behavior lineage comes from DavidAU's Qwen3.6 40B Opus Deckard release, with the MTP graft and GGUF recipe from PiehSoft's Deckard MTP work. The Chadrock release converts that line into a Strix-focused ROCmFP4 runtime format and uses a tuned draft-MTP serving profile.

This GGUF will not run correctly with stock llama.cpp. It needs the pinned ciru-ai/ROCmFPX runner because the file uses ROCmFP4 tensor types that upstream llama.cpp does not currently understand. For image input, use the Vulkan-only ROCmFP4 build path shown below.

The model file is already provided here. You do not need to rebuild or quantize the model. Build the custom llama server once, download this GGUF, and run the profile below.

Why This Build Exists

This release is for people who want the 40B Deckard/Opus behavior line with native MTP speculative decoding and lower AMD unified-memory pressure than the available K-quants.

The goal is not a generic GGUF quant. The goal is a model/runtime pairing for Strix Halo:

  • Deckard 40B Opus / Heretic / Uncensored / Thinking lineage
  • MTP speculative decoding from the PiehSoft Deckard MTP recipe
  • ROCmFP4 Strix Lean tensor recipe
  • Vulkan target+draft serving on AMD Strix Halo
  • 262144 token context target
  • draft-mtp depth 2 profile
  • lower GPU pressure than the Q4_K_M baseline in the local sweep

Model Lineage

Qwen3.6 27B / 40B compatible hidden-size family
  -> DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
       behavior lineage:
         - Claude 4.6 Opus / Deckard style
         - Heretic / uncensored / thinking / NEO-CODE / IMatrix line
  -> PiehSoft/Qwen3.6-40B-Deckard-MTP
       adds:
         - GGUF MTP graft recipe
         - Deckard MTP quants
         - MTP head from the compatible Qwen3.6 27B MTP donor line
  -> CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4
       adds:
         - ROCmFP4 Strix Lean conversion
         - Chadrock AMD serving profile

Direct source links:

Technical Metadata

Field Value
model size 40B dense
architecture family qwen35 / Qwen3.6
runtime format Chadrock ROCmFP4 GGUF
release filename CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf
local profile chadrock-qwen36-40b-deckard-mtp-rocmfp4-vulkan-d2
context target 262144 tokens
max generation target 16384 tokens
draft mode draft-mtp, n_max=2, p_split=0.10
target backend in local profile Vulkan0
target/draft KV f16 / f16
batch / ubatch 2048 / 512
intended hardware AMD Ryzen AI Max+ 395 / Strix Halo

Vision Status

Vision is now locally validated for this exact ROCmFP4 GGUF with the included Qwen3.6 27B projector:

mmproj-Qwen3.6-27B-f16.mmproj

Validated command path:

  • runner repo: ciru-ai/ROCmFPX
  • pinned runner commit: 7aa484a2f0a504dc612a3d74a068024f3e6d6353
  • legacy vision smoke-test commit: 4795079b0
  • build mode: GGML_VULKAN=ON, GGML_HIP=OFF, GGML_CUDA=OFF
  • runtime device: Vulkan0
  • required flags: --mmproj mmproj-Qwen3.6-27B-f16.mmproj --image-min-tokens 1024 -fa off

Smoke test image result on June 15, 2026 through the live OpenAI-compatible endpoint:

- Red box
- Blue circle
- Green triangle
- Text: "RED BOX BLUE CIRCLE GREEN TRIANGLE"

The earlier failure mode was runner-side: older HIP-enabled Strix ROCmFP4 builds loaded the projector but produced wrong image descriptions. The model file and projector are usable when served by the Vulkan-only ROCmFP4 build above.

The projector is uploaded with a .mmproj extension so Hugging Face's GGUF parser keeps the primary model file as the main artifact. Pass it to llama.cpp like a normal mmproj file.

Speed: Baseline vs Chadrock

These are clean local server/API measurements on AMD Ryzen AI Max+ 395 / Strix Halo. The comparison uses the same custom llamaCHADROCK runner, the same prompt files, no prompt cache, one slot, 512 generated tokens, draft-mtp, n_max=2, f16 target/draft KV, -b 2048, -ub 512, -sm none, -ngl 999, and Vulkan0.

Baseline:

PiehSoft/Qwen3.6-40B-Deckard-MTP
Qwen3.6-40B-Deckard-MTP-Q4_K_M.gguf

Chadrock:

CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf
Model Prompt ctx PP tok/s TG tok/s TTFP Draft accepted Peak GPU pressure Wall
PiehSoft Q4_K_M baseline 4,096 178.61 14.84 22.94s 293/435 (67.4%) 30.25 GiB 57.45s
PiehSoft Q4_K_M baseline 8,192 171.34 16.03 47.83s 313/395 (79.2%) 30.25 GiB 79.76s
PiehSoft Q4_K_M baseline 16,384 159.37 15.25 102.82s 310/400 (77.5%) 30.26 GiB 136.41s
CHADROCK ROCmFP4 4,096 204.64 18.24 20.02s 306/408 (75.0%) 26.49 GiB 48.09s
CHADROCK ROCmFP4 8,192 197.70 17.68 41.45s 304/413 (73.6%) 26.49 GiB 70.41s
CHADROCK ROCmFP4 16,384 182.49 16.90 89.80s 303/414 (73.2%) 26.49 GiB 120.09s

Uplift

Context Prompt gain Decode gain TTFP reduction Wall reduction GPU pressure saved
4,096 +14.6% +22.9% 12.7% faster 16.3% faster 3.77 GiB (12.5%)
8,192 +15.4% +10.3% 13.3% faster 11.7% faster 3.77 GiB (12.5%)
16,384 +14.5% +10.9% 12.7% faster 12.0% faster 3.77 GiB (12.4%)

On Strix Halo, "GPU pressure" here is peak_vram_used_bytes + peak_gtt_used_bytes. The iGPU has a small dedicated VRAM aperture plus a large GTT/unified-memory pool, so GTT is part of the real pressure picture.

The local benchmark artifact is:

/home/crown/bench-results/llama/context-sweeps/20260615T074356Z-deckard40-clean-baseline-vs-chadrock/api/results.sqlite3

These are local server measurements, not universal llama.cpp claims. MTP throughput depends on prompt shape, driver version, clocks, KV cache settings, draft depth, and draft acceptance.

Best Settings / Advanced Setup

For the pinned score-build runner notes and copy-paste ROCmFP4 setup commands,
use the advanced Ciru setup page:

https://llm.ciru.ai/chadrock-rocmfpx/

The shared pinned runner build is:

ciru-ai/ROCmFPX commit: 7aa484a2f0a504dc612a3d74a068024f3e6d6353
historical score tag: chadrock-rocmfp4-mtp-scores-20260621

For this Deckard 40B release, use the vision-validated Vulkan-only profile
below rather than an older HIP-enabled Strix runner:

backend: Vulkan0 target + Vulkan0 draft
context: 262144
batch / ubatch: 2048 / 512
target KV: f16 / f16
draft KV: f16 / f16
MTP: draft-mtp, n_max=2, n_min=0, p_min=0.0, p_split=0.10
serving: one slot, metrics on, --mmproj mmproj-Qwen3.6-27B-f16.mmproj
vision: --image-min-tokens 1024, flash attention off
sampler: temperature=0, top_p=0.95, top_k=20
reasoning: off, reasoning_format=none, reasoning_budget=0

The projector path was validated with a real image-answer gate. Do not use the
older HIP-enabled runner for multimodal validation; it could load the projector
but produced wrong image descriptions in local testing.

Run With llama-server

Build Charlie's custom llama.cpp once, download this GGUF, then run:

/path/to/rocmfp4-llama/build-vulkan-rocmfp4/bin/llama-server \
  -m CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf \
  --alias main \
  --host 0.0.0.0 \
  --port 8080 \
  --jinja \
  -c 262144 \
  -n 16384 \
  --reasoning off \
  --reasoning-format none \
  --reasoning-budget 0 \
  -ngl 999 \
  -fa off \
  -sm none \
  -dev Vulkan0 \
  --mmproj mmproj-Qwen3.6-27B-f16.mmproj \
  --image-min-tokens 1024 \
  -b 2048 \
  -ub 512 \
  -t 16 \
  -tb 32 \
  -ctk f16 \
  -ctv f16 \
  --ctx-checkpoints 0 \

  --checkpoint-every-n-tokens -1 \

  --spec-type draft-mtp \
  --spec-draft-device Vulkan0 \
  --spec-draft-ngl all \
  --spec-draft-type-k f16 \
  --spec-draft-type-v f16 \
  --spec-draft-threads 16 \
  --spec-draft-threads-batch 32 \
  --spec-draft-n-max 2 \
  --spec-draft-n-min 0 \
  --spec-draft-p-min 0.0 \
  --spec-draft-p-split 0.10 \
  --poll 100 \
  --poll-batch 1 \
  --spec-draft-poll 1 \
  --spec-draft-poll-batch 1 \
  --temp 0 \
  --min-p 0.0 \
  --top-p 0.95 \
  --top-k 20 \
  --repeat-penalty 1.0 \
  --seed 123 \
  --cache-ram 8192 \
  --parallel 1 \
  --metrics

Use --parallel 1 for this MTP profile. Multi-slot serving changes draft-MTP behavior and is not the intended configuration.

The local profile uses Vulkan0 for target and draft on Strix Halo. The required runner is still the ROCmFP4 llama.cpp fork because its Vulkan path understands the custom ROCmFP4 GGUF tensor types.

UI Notes

The reproducible public build above uses -DLLAMA_BUILD_WEBUI=OFF, so it is an API server build. Use it with OpenAI-compatible clients pointed at /v1/chat/completions.

If you need the bundled llama.cpp WebUI, keep the same Vulkan-only backend settings but build from a revision where the WebUI assets build or provision successfully. Do not switch to the older HIP-enabled Strix runner path for multimodal validation.

Build The Required ROCmFP4 Runner

The GGUF is already provided. You only need to build the custom server once:

git clone https://github.com/ciru-ai/ROCmFPX.git
cd ROCmFPX
git checkout 7aa484a2f0a504dc612a3d74a068024f3e6d6353

cmake -S . -B build-vulkan-rocmfp4 -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DGGML_VULKAN=ON \
  -DGGML_HIP=OFF \
  -DGGML_CUDA=OFF \
  -DLLAMA_BUILD_WEBUI=OFF

cmake --build build-vulkan-rocmfp4 -j 16 --target llama-server llama-quantize

The server binary will be here:

build-vulkan-rocmfp4/bin/llama-server

Recommended Strix Halo prerequisites:

sudo apt-get update
sudo apt-get install -y git cmake ninja-build build-essential clang pkg-config \
  glslc vulkan-tools

vulkaninfo --summary

This Vulkan-only build is the recommended path for vision on this release. The older HIP-enabled Strix runner path can load the projector but was observed to give wrong image descriptions, so do not use that path for multimodal validation.

About ROCmFP4 / Chadrock

ROCmFP4 is not stock Q4, MXFP4, NVFP4, or a renamed K-quant. It is a custom GGUF tensor format with AMD-focused backend support.

Charlie's ROCmFP4 fork adds:

  • Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST tensor formats
  • tensor-aware Strix presets such as Q4_0_ROCMFP4_STRIX_LEAN
  • ROCm/HIP kernels for the new layouts
  • Vulkan shader support for ROCmFP4 decode paths
  • MTP regression guards for target/draft serving

The practical point: this model is a Chadrock runtime artifact. Use the ROCmFP4 runner named above.

File

File Size SHA256
CHADROCK3.6-40B-Opus-Deckard-Uncensored-Thinking-NEO-CODE-Di-IMatrix-ROCmFP4.gguf 20 GB 252c463f6edd267dd04bff303af6d6e7fa22279a2f50ec74e002dabba270d0d7
mmproj-Qwen3.6-27B-f16.mmproj 885 MB 26b00d800d2853627d09f0caabde10e79a1a6e5e5dc589fa62deeb1b14f3c673

Credits

  • DavidAU: Qwen3.6 40B Opus Deckard / Heretic / Uncensored / Thinking source lineage.
  • PiehSoft: Deckard 40B MTP GGUF recipe, MTP validation notes, and Q4_K_M baseline used for this comparison.
  • froggeric: Qwen3.6 27B mmproj used for upstream multimodal validation and included here as an experimental companion.
  • Unsloth: Qwen3.6 27B MTP GGUF donor reference used in the compatible MTP-head lineage.
  • Qwen: Qwen3.6 base model family.
  • charlie12345 / @Italianclownz: ROCmFP4 llama.cpp fork, Strix Halo build path, and AMD-focused MTP runtime work.

Notes

  • This is an experimental AMD ROCmFP4/MTP runtime release.
  • Stock llama.cpp, Ollama, LM Studio, and other tools will not understand the ROCmFP4 GGUF tensor types unless they incorporate the required runner support.
  • Vision is validated with mmproj-Qwen3.6-27B-f16.mmproj using the Vulkan-only ROCmFP4 build path above.
  • The benchmark comparison is intentionally server/API based because standalone llama-bench TG is only diagnostic for MTP models and can understate served draft-MTP throughput.
  • The model can be served under any alias. The local service profile uses main; the public release identity and actual file name are the full Chadrock name above.

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-23Pin Chadrock v2 runner checkpoint fix76e760113.8 KB
    Loading...
  2. 2026-06-22Pin ROCmFPX runner commit in README7537ae313.7 KB
    Loading...
  3. 2026-06-21Add best Chadrock runtime settings7fcbae413.6 KB
    Loading...
  4. 2026-06-15Document validated ROCmFP4 vision runner042b94b12.5 KB
    Loading...
  5. 2026-06-15Clarify experimental vision statusdb9f54212.1 KB
    Loading...
  6. 2026-06-15Enable vision projector in Chadrock 40B card71c2cc412.3 KB
    Loading...
  7. 2026-06-15Add model cardafaca6610.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration