license: apache-2.0
language:
- en
- multilingual
pipeline_tag: image-text-to-text
tags: - qwen3.6
- qwen3.5
- moe
- safetensors
- bfloat16
- transformers
- vllm
- vision
- multimodal
- genesis
- uncensored
- blackwell
base_model: - LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-GGUF
Qwen3.6-35B-A3B-Uncensored-Genesis-Final — SafeTensors BF16 Reconstruction
Thanks to MarMix for doing GGUF to Safetensors conversion.
This repository contains a verified SafeTensors/BF16 reconstruction of:
LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-GGUF
Source GGUF used for reconstruction:
Qwen3.6-35B-A3B-Uncensored-Genesis-Final-Q8_K_P.gguf
This release was prepared primarily for Transformers / vLLM inference on NVIDIA GPUs, and was validated on an:
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Important provenance note
This is not the original publisher BF16 checkpoint.
It is a SafeTensors/BF16 reconstruction from the preserved Q8_K_P GGUF lineage. Information lost during Q8_K_P quantization cannot be recovered by conversion back to BF16.
The tensors present in the Genesis GGUF were reconstructed into the Hugging Face / SafeTensors layout. Architecture-required tensors that are not stored in the GGUF representation, including missing vision/MTP-related tensors, were copied from the compatible reference model
Qwen/Qwen3.6-35B-A3B.Therefore, do not describe this repository as an original or lossless BF16 release.
Upstream model
Genesis release:
https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-GGUF
Genesis is based on:
https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
Reference architecture files / missing architecture tensors used during reconstruction:
https://huggingface.co/Qwen/Qwen3.6-35B-A3B
Please see the upstream Genesis repository for the author's description of the Genesis process, acknowledgements, intended behavior, recommended prompting, and original GGUF releases.
Reconstruction verification
The conversion was performed with the Qwen3.6 MoE conversion path and then verified against both the source GGUF and the Hugging Face reference layout.
Verification summary:
- Source GGUF tensors: 733
- Reference-model tensors: 1045
- Tensors reconstructed from GGUF: 733
- Architecture-required tensors copied from reference: 352
- Total tensors in reconstructed model: 1045
- GGUF-expected tensors checked bit-exact: 693
- GGUF mismatches: 0
- Reference-copied tensors checked bit-exact: 352
- Reference-copy mismatches: 0
- Unmapped GGUF tensors: 0
- Missing converted tensors: 0
- Extra converted tensors: 0
- Verification result: CONVERSION CORRECT: YES
The generated checkpoint contains 17 SafeTensors shards.
Architecture
The reconstructed model resolves in current Transformers/vLLM as:
Qwen3_5MoeForConditionalGeneration
Key upstream architecture characteristics:
- ~35B total parameters
- ~3B active parameters per forward pass
- 256 experts
- 8 routed experts + 1 shared expert per token
- 40 layers
- Hybrid Gated DeltaNet / full-attention MoE architecture
- Native multimodal support
- Text, image, and video processing
- Native context metadata up to 262K tokens
The included Hugging Face processor loads as:
Qwen3VLProcessor
with image and video processors available.
Multimodal / vision note
The upstream GGUF release uses a separate mmproj file for llama.cpp-style multimodal inference.
Do not load the GGUF mmproj file with this SafeTensors/vLLM reconstruction.
For this repository, the Hugging Face multimodal processor and the required vision tensors are already part of the reconstructed Transformers-compatible model layout.
The vision/MTP tensors that were absent from the GGUF representation came from the compatible Qwen/Qwen3.6-35B-A3B reference model. They should not be represented as Genesis-modified tensors unless independently demonstrated.
Recommended vLLM setup
Validated environment:
- vLLM: 0.29.0
- Python: 3.13
- GPU: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
- Compute capability: SM120 / 12.0
- SafeTensors dtype: BF16
- Tested context: 131072
- Tool calling: enabled
- Reasoning parser: enabled
- Multimodal encoder: enabled
Recommended single-user workstation configuration
The following configuration is recommended for interactive single-user use with OpenCode:
# Use 127.0.0.1 for local-only access.
# Use 0.0.0.0 only if you intentionally want vLLM reachable from other machines.
vllm serve /path/to/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-Safetensors-BF16 \
--host 127.0.0.1 \
--port 11620 \
--served-model-name Qwen3.6-35B-A3B-Uncensored-Genesis-Final \
--dtype bfloat16 \
--max-model-len 131072 \
--gpu-memory-utilization 0.80 \
--max-num-seqs 32 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml
Notes and recommendations
--host 127.0.0.1keeps the API accessible only from the local machine.--host 0.0.0.0makes vLLM listen on all network interfaces. Use this only when remote access is required, and protect the server with an API key and appropriate firewall rules.--max-model-len 131072enables a maximum context length of 128K tokens.--gpu-memory-utilization 0.80leaves additional VRAM headroom for CUDA workspaces, temporary allocations, multimodal processing, and long-running workloads. The model was previously validated successfully with0.90, but0.80is a more conservative workstation setting.--max-num-seqs 32is recommended for single-user OpenCode use. The model was previously tested with512, which is more appropriate for high-concurrency server workloads.--reasoning-parser qwen3enables Qwen-style reasoning parsing.--enable-auto-tool-choiceallows the model to decide when to call tools.--tool-call-parser qwen3_xmlenables parsing of Qwen tool calls for OpenAI-compatible clients such as OpenCode.- For a dedicated inference server with many simultaneous users, higher values for
--gpu-memory-utilizationand--max-num-seqsmay be appropriate. For an interactive workstation, lower values provide more VRAM headroom and reduce unnecessary resource pressure.
For a local OpenAI-compatible endpoint that should not be anonymously accessible, also add:
--api-key YOUR_API_KEY
Binding to 127.0.0.1 is still recommended for local-only use.
Blackwell notes
On the NVIDIA RTX PRO 6000 Blackwell Workstation Edition, vLLM may perform first-run compilation, CUDA graph capture, and FlashInfer/CUTLASS autotuning.
That initial startup can take substantially longer than later starts because compiled kernels and autotuning results are cached.
During validated startup:
- model weights occupied approximately 65.5 GiB of GPU memory
- vLLM reserved additional VRAM for KV cache and runtime buffers
- total idle allocation was approximately 88 GiB on a ~96 GiB card
- idle GPU utilization remained 0%
High VRAM allocation while idle is therefore expected with --gpu-memory-utilization 0.90; it does not mean the GPU is under heavy compute load.
Context and concurrency
With the validated 131072-token configuration, vLLM reported approximately:
- 16.56 GiB available KV-cache memory
- 822,451 tokens of GPU KV-cache capacity
- about 6.27× maximum concurrency at a full 131072-token context
Actual throughput and concurrency depend on prompt length, output length, multimodal inputs, sampling, and concurrent requests.
Recommended Genesis sampling settings
The upstream Genesis README recommends the following for thinking-mode coding / precise tasks:
temperature = 0.6
top_p = 0.95
top_k = 20
min_p = 0.0
seed = 42
presence_penalty = disabled
repeat_penalty = disabled
Recommended Sampling Settings
The following settings are recommended by the upstream Genesis release.
Thinking mode (coding):
- Coding / precise tasks:
temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, seed=42, presence_penalty=disabled, repeat_penalty=disabled - General:
temperature=0.95, top_p=0.95, top_k=20, min_p=0.0, seed=42, presence_penalty=disabled, repeat_penalty=disabled
temperature = 0.6
top_p = 0.95
top_k = 20
min_p = 0.0
seed = 42
presence_penalty = disabled
repeat_penalty = disabled
Non-thinking mode (creative):
- General:
temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, seed=42, presence_penalty=disabled, repeat_penalty=disabled - General alternative:
temperature=0.95, top_p=disabled, top_k=20, min_p=disabled, seed=42, presence_penalty=disabled, repeat_penalty=disabled
temperature = 0.7
top_p = 0.8
top_k = 20
min_p = 0.0
seed = 42
presence_penalty = disabled
repeat_penalty = disabled
A neutral repeat penalty of 1.0 is equivalent to applying no repetition penalty in runtimes that require a numeric value.
The upstream recommended minimal system prompt is:
You are Qwen, a large language model developed by Alibaba Group's Tongyi Lab. You are a helpful assistant.
The upstream author also recommends keeping at least 128K context for thinking behavior.
OpenCode example
Example provider configuration for an OpenAI-compatible vLLM endpoint:
{
"provider": {
"vllm-genesis": {
"npm": "@ai-sdk/openai-compatible",
"name": "Qwen3.6 Genesis Final BF16 via vLLM",
"options": {
"baseURL": "http://127.0.0.1:10101/v1",
"apiKey": "{file:/home/openrag/.config/opencode/vllm-genesis-api-key}"
},
"models": {
"Qwen3.6-35B-A3B-Uncensored-Genesis-Final": {
"name": "Qwen3.6-35B-A3B Uncensored Genesis Final BF16",
"limit": {
"context": 131072,
"output": 32768
},
"tool_call": true,
"temperature": true,
"reasoning": true,
"attachment": true
}
}
}
},
"model": "vllm-genesis/Qwen3.6-35B-A3B-Uncensored-Genesis-Final"
}
Recommended OpenCode coding-agent sampling profile Thinking mode (coding):
{
"temperature": 0.6,
"top_p": 0.95,
"top_k": 20,
"min_p": 0.0,
"presence_penalty": 0.0,
"repetition_penalty": 1.0,
"seed": 42
}
What this repository is — and is not
This repository is:
- a Transformers-compatible SafeTensors reconstruction
- BF16 storage of tensors reconstructed from the Q8_K_P GGUF lineage
- verified against the source GGUF and compatible Qwen reference layout
- suitable for vLLM inference
- suitable for NVIDIA Blackwell inference
- multimodal-capable through the Hugging Face processor/model layout
This repository is not:
- the original pre-quantization BF16 checkpoint
- a mathematically lossless recovery of weights discarded by Q8_K_P quantization
- a claim that reference-only vision/MTP tensors contain Genesis-specific modifications
- a GGUF checkpoint
- a model that requires a separate
mmproj.ggufwhen used through Transformers/vLLM
Credits
All model-development, fine-tuning/uncensoring, Genesis processing, and upstream release credit belongs to the respective upstream authors.
- Genesis release: LuffyTheFox — if you would like to support his work on the Genesis project, you can find him on Hugging Face and use the donation options listed on his model page.
- Special thanks to HauhauCS for the upstream uncensored model.
- Qwen architecture/reference model: Qwen / Alibaba Cloud
This repository only provides the SafeTensors/BF16 reconstruction and verification described above.
Please refer to the upstream repositories for their full credits, licensing terms, model behavior notes, and recommended usage.