license: apache-2.0
base_model:
- HauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive
- z-lab/Qwen3.5-122B-A10B-DFlash
tags: - dflash
- speculative-decoding
- draft-model
- gguf
- qwen3.5
- dgx-spark
library_name: gguf
Qwen3.5 122B HauhauCS Aggressive DFlash draft
This repository contains a trained DFlash draft for the exact target:
HauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive, tested with its Q4_K_M GGUF (approximately 74 GB).
It is not a standalone language model. Pair it with the target in a DFlash-capable speculative-decoding runtime. The target weights are not duplicated here.
Important: Hugging Face may display automatically generated
llama.cpp, Ollama, or local-app commands for repositories containing GGUF files. Those generic commands are not sufficient for this draft. Load this repository as the draft model alongside the linked HauhauCS target, using the patched DFlash runtime and launch instructions in the companion GitHub repository.
Release files
epoch-011-Q4_K_M.gguf— recommended llama.cpp draft, 462,634,656 bytes.epoch-011-BF16.gguf— BF16 GGUF draft.model.safetensors— framework checkpoint.config.json— DFlash architecture configuration.training.sanitized.json— training metadata with host-specific paths removed.SHA256SUMS— release hashes.
Local benchmark
| Configuration | Median output rate | p50 latency |
|---|---|---|
| Baseline | 25.58 tok/s | 4.19 s |
| Routed DFlash | 44.91 tok/s | 2.10 s |
- Routed draft acceptance: 77.78%.
- 8/8 deterministic output comparisons exactly matched the depth-zero baseline.
- 40/40 routed soak requests completed with 0 errors and 74.37% draft acceptance.
These results describe a small, deterministic local workload. They are not a guarantee for other prompts, runtimes, hardware, concurrency, or decoding settings.
Runtime requirements
The tested llama.cpp path used --spec-type draft-dflash, --spec-draft-model, per-request speculative.n_max, target context 262,144, two parallel slots, thinking disabled, and deterministic decoding.
The required llama.cpp and Speculators patches, launch/training scripts, exact commit pins, and reproducibility record are staged in the companion GitHub repository.
Attribution
- Target model and quantization: HauhauCS
- DFlash paper/reference implementation: z-lab/dflash
- Stock initialization draft: z-lab/Qwen3.5-122B-A10B-DFlash
- Training framework: vllm-project/speculators
- Runtime: ggml-org/llama.cpp
Limitations
- The draft is specialized for the named HauhauCS Aggressive target and tested quant. Compatibility with related Qwen3.5 checkpoints is unverified.
- Correctness in speculative decoding still depends on the target verifier and compatible runtime behavior.
- The release does not endorse or reproduce every claim made by the target model publisher.
- Users are responsible for evaluating safety and suitability in their deployment context.
License
Apache-2.0, subject to the licenses and attribution of the target model, stock DFlash draft, training framework, and runtime components.
Release pair
This repository contains the trained weights, framework checkpoint, sanitized metadata, and checksum manifest. Companion code, patches, launch scripts, and reproducibility details: D4pp3rD1scourse/qwen35-122b-hauhaucs-dflash. Use the draft only with the linked target and compatible patched runtime.