license: apache-2.0
language:
- en
base_model: - Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16
pipeline_tag: image-text-to-text
library_name: transformers
tags: - diffusion
- diffusion-lm
- diffusiongemma
- gemma
- abliteration
- abliterated
- heretic
- multimodal
- nvfp4
- quantization
- mixed-precision
- blackwell
- cuda-graphs
- vllm
- tts
- voice-assistant
- low-latency
DiffusionGemma 26B E38 — Abliterated NVFP4
A 26B-class diffusion language model optimized for fast complete-response generation and full-text voice.
The latest testing produced several interesting results.
For short responses, less diffusion performed better.
The faster:
128-token canvas / 16-step
configuration was not only faster than:
256-token canvas / 48-step
but also produced a higher strictly-clean response rate in blind manual A/B testing inside the intended 1–64-token short-response range.
The result changes for longer generation, where 256/48 remains safer.
A second finding is that DiffusionGemma's occasional grammar, malformed-word, repetition, and lexical errors increasingly appear more consistent with:
diffusion finalization behavior
than with:
NVFP4 quantization damage
And the original deployment concept is now working in a real local voice prototype:
DiffusionGemma → complete response → complete generated audio
In a measured prototype turn:
LLM complete: 123.1 ms
Entire TTS audio generated: 117.8 ms
Complete text + complete audio ready in 260.1 ms
The TTS number is not time-to-first-audio.
The entire 28.23-second audio waveform was already synthesized and ready for playback at the end of the measured TTS stage.
Current recommendation:
| Use Case | Recommended Mode |
|---|---|
| Short replies / Full-Text TTS / 1–64 tokens | 128/16 |
| Longer responses / general chat / reasoning | 256/48 |
Primary deployment target:
Instant Full-Text Voice on One GPU
🎥 Working Voice Demo
▶ Watch the DiffusionGemma E38 NVFP4 Voice Demo on YouTube
The working local prototype uses:
DiffusionGemma E38 NVFP4
↓
Complete Response
↓
Kokoro TTS
↓
Complete Audio Waveform
↓
Ready for Playback
One measured prototype run produced:
| Stage | Result |
|---|---|
| LLM complete-response latency | 123.1 ms |
| TTS — entire audio synthesized | 117.8 ms |
| Complete text + complete audio ready | 260.1 ms |
| Response length | 109 tokens |
| LLM throughput | 885.5 tok/s |
| Fully generated audio duration | 28.23 seconds |
The full 28.23-second audio output was generated in 117.8 ms.
This is an important distinction.
The TTS measurement is not:
- time to first audio sample
- time to begin streaming speech
- time until the first sentence is synthesized
It is the measured time required to synthesize the entire audio output for that response.
Likewise, the 260.1 ms pipeline measurement means that by the end of that measured interval:
- DiffusionGemma had generated the complete text response.
- Kokoro had received the complete response.
- Kokoro had synthesized the complete 28.23-second audio waveform.
- The audio was ready for playback.
So for this measured turn:
109-token complete response + 28.23 seconds of finished speech were ready in 260.1 ms of measured LLM + TTS processing.
This does not include microphone capture, STT, playback-device startup, UI overhead, or the real-world time required to actually listen to 28.23 seconds of speech.
The 109-token response is also outside the formally validated 1–64-token short-quality envelope, so this example is presented as a:
voice-pipeline performance demonstration
rather than evidence extending the formal short-response quality result to 109 tokens.
🗣️ Voice Demo Pre-Prompt
The voice prototype uses a short conversational pre-prompt designed to keep responses concise, natural, and suitable for complete-text TTS.
You are a fast local voice assistant.
Answer the user's request directly.
Use natural spoken English and complete sentences.
Keep normal conversational answers short and focused.
Avoid markdown, bullet points, headings, citations, emojis, role labels, stage directions, and meta commentary unless the user explicitly asks for them.
Do not open with filler such as "Okay", "Sure", or "Of course".
Do not explain what you are doing.
Prefer wording that sounds natural when read aloud by TTS.
For ordinary voice-chat questions, aim for a complete response within roughly 1–64 output tokens.
If the request genuinely requires more detail, answer fully rather than forcing an unnaturally short response.
Return only the text that should be spoken.
For the 128/16 Instant Full-Text TTS mode, the intended operating range is:
approximately 1–64 output tokens
This is a target for normal conversational replies, not a hard truncation rule.
Longer or more detailed responses should use the 256/48 quality configuration.
🚀 Model
Goodoldjam/DiffusionGemma-26B-E38-Abliterated-NVFP4
E38 NVFP4 reduces the E38 BF16 checkpoint from:
51.68 GB → 18.86 GB
while preserving the measured E38 capability profile and enabling native low-precision inference on NVIDIA Blackwell.
The expensive part is not making another quant. The expensive part is proving which quant and inference configuration is actually worth using.
🧪 Response-Length Validation
128/16 and 256/48 were evaluated across both short conversational responses and longer-form generation.
The result was strongly dependent on response length.
Short Responses — 1–64 Tokens
The intended Instant Full-Text TTS operating range was tested using blind matched A/B manual review.
Manual review completed so far:
1,850 matched A/B pairs
representing:
3,700 manually reviewed responses
Within the intended 1–64-token range:
| Short-Response Result | 128/16 | 256/48 |
|---|---|---|
| In-range responses reviewed | 1,705 | 1,690 |
| Words reviewed | 46,282 | 45,227 |
| Strictly clean responses | 1,678 | 1,638 |
| Strictly clean rate | 98.42% | 96.92% |
| Responses with confirmed language issue | 27 | 52 |
| Severe degeneration | 0 | 0 |
| Truncation | 0 | 0 |
A response was considered strictly clean only if it had:
- no confirmed grammar error
- no malformed lexical artifact
- no accidental repetition
- no degeneration
- no malformed ending
- no truncation
Awkward-but-grammatical wording and factual/content mistakes were tracked separately.
Fully Paired Short-Response Result
Across:
1,633 matched A/B pairs
where both configurations stayed inside the 1–64-token range:
| Paired Outcome | Count |
|---|---|
| Both clean | 1,563 |
| 128/16 clean / 256/48 issue | 48 |
| 256/48 clean / 128/16 issue | 21 |
| Both issue | 1 |
Paired strictly-clean rates:
128/16: 98.65%
256/48: 97.00%
Difference:
+1.65 percentage points in favor of 128/16
Paired bootstrap:
95% CI: approximately +0.67 to +2.63 percentage points
Exact paired test:
McNemar p ≈ 0.0016
The predeclared short-response non-inferiority margin was:
−2 percentage points
128/16 cleared that requirement.
Conclusion:
128/16 SHORT-TTS QUALITY PASS
This does not establish universal superiority.
It means:
128/16 performed better in the manually reviewed short-response operating range.
Longer Responses
The result changes when 128/16 is pushed beyond its intended short-response role.
A separate blind stress test used approximately:
240-word responses
Results:
| Long-Form Result | 128/16 | 256/48 |
|---|---|---|
| Responses reviewed | 25 | 25 |
| Definite grammar-error responses | 7 | 5 |
| Repetition-affected responses | 4 | 0 |
| Severe degeneration | 2 | 0 |
| Strictly clean responses | 15 | 17 |
Both severe degeneration cases came from:
128/16
The current deployment boundary is therefore:
128/16 → short complete responses
256/48 → general and long-form responses
The evidence does not suggest one universally optimal diffusion configuration.
Instead:
the best diffusion budget appears to depend on response length.
🧠 Grammar Errors and the Diffusion-Finalization Hypothesis
DiffusionGemma can occasionally produce unusual language artifacts such as:
- malformed words
- duplicated words
- incorrect local grammar
- awkward substitutions
- malformed sentence endings
- rare degeneration
These issues raised an obvious question:
Did NVFP4 quantization damage the language model?
The current evidence does not support that explanation.
Instead, several observations increasingly point toward:
diffusion inference and final text resolution
as a more likely source.
This remains a working research hypothesis rather than a proven causal mechanism.
The Error Class Exists Before NVFP4
Similar rare language failures were observed in:
- Base BF16
- E38 BF16
- E38 NVFP4
The general phenomenon therefore existed before the NVFP4 conversion.
NVFP4 Did Not Increase Measured Language Errors
A larger long-form language-quality study measured:
| Model | Grammar Errors /10k ↓ | Lexical Artifacts /10k ↓ |
|---|---|---|
| Base BF16 | 4.059 | 2.243 |
| E38 BF16 | 5.479 | 1.865 |
| E38 NVFP4 | 3.236 | 1.387 |
NVFP4 had the lowest observed grammar and lexical-artifact rates in that study.
The E38 BF16 grammar difference relative to Base was:
+1.420 grammar errors /10k words
with:
95% CI: −0.674 to +3.541
and:
McNemar p = 0.560
The E38 BF16 grammar difference was therefore not established as a statistically confirmed E38-specific regression.
The NVFP4 results provide:
no evidence that quantization introduced additional language corruption
Errors Cluster Late in the Diffusion Canvas
One of the most interesting findings came from looking at where errors occur.
Language errors were approximately:
3.01× more concentrated in the final quarter of the 256-token canvas
than in the first quarter.
Error-prone positions also showed approximately:
+0.0504 higher final entropy
and:
−0.0418 lower top-1 / top-2 confidence margin
relative to matched clean positions.
This suggests the model is less certain when resolving some problematic final text.
More Denoising Did Not Reliably Fix It
Simply increasing the denoising budget did not establish a consistent overall language-quality improvement.
The short-response experiments add another clue:
128/16 was cleaner than 256/48 in the current blind short-response sample.
If more denoising automatically produced better language, this would not be the expected result.
But reducing the diffusion budget too aggressively for long responses creates another failure mode:
repetition and degeneration
This suggests there may be a useful operating region where the diffusion process has enough compute to resolve a response cleanly without unnecessary additional refinement.
Current Working Hypothesis
The current evidence suggests:
DiffusionGemma may know more language than its inference process always cleanly finalizes.
A simplified interpretation:
Underlying learned capability
↓
Partially denoised response
↓
Mostly correct structure
↓
Late diffusion resolution
↓
Local uncertainty
↓
Rare malformed / repeated / awkward token
This could explain why:
- the same general error class exists in BF16
- NVFP4 does not make it worse
- errors cluster late in the canvas
- error positions show higher entropy
- error positions show lower confidence margins
- more denoising does not always improve output
- reduced diffusion improves short-response cleanliness
- insufficient diffusion capacity hurts long-form generation
The current evidence therefore points more strongly toward:
diffusion finalization behavior
than toward:
NVFP4 quantization damage
as the source of these rare grammar and lexical issues.
This remains an active research question.
⚡ Large Short-Response Latency Validation
The larger short-response experiment used:
500 frozen prompts
across:
15 voice-assistant categories
with:
4 matched seeds
for a total of:
2,000 matched pairs
and:
4,000 total generations
Generated words:
| Configuration | Words |
|---|---|
| 128/16 | 58,355 |
| 256/48 | 58,077 |
Runtime failures:
0
Empty outputs:
0
The primary deployment analysis focused on outputs with actual generated length:
1–64 tokens
Counts:
| Configuration | Responses in 1–64 Range |
|---|---|
| 128/16 | 1,843 |
| 256/48 | 1,827 |
Pairs where both configurations remained in range:
1,766
Large-Study Latency
| Metric | 128/16 | 256/48 |
|---|---|---|
| Median latency | 139.2 ms | 158.1 ms |
| P90 latency | 190.1 ms | 212.8 ms |
128/16 reduced:
median latency by approximately 11.9%
and:
P90 latency by approximately 10.6%
versus 256/48.
Latency by Actual Output Length
| Actual Output Length | 128/16 Median | 256/48 Median |
|---|---|---|
| 1–16 tokens | 105.5 ms | 116.3 ms |
| 17–32 tokens | 119.8 ms | 132.6 ms |
| 33–48 tokens | 145.5 ms | 163.5 ms |
| 49–64 tokens | 175.7 ms | 195.5 ms |
128/16 was faster in every measured primary short-response bucket.
Recommended deployment range:
1–64 output tokens
⚡ Earlier Instant Chat Experiment
An earlier smaller matched experiment compared:
- 256 / 48
- 256 / 16
- 128 / 16
- 64 / 16
Results:
| Metric | 256/48 | 256/16 | 128/16 | 64/16 |
|---|---|---|---|---|
| Median latency | 141.0 ms | 131.3 ms | 120.1 ms | 134.2 ms |
| Mean latency | 146.6 ms | 139.3 ms | 124.7 ms | 152.1 ms |
| P90 | 216.1 ms | 204.2 ms | 187.3 ms | 244.0 ms |
| P95 | 251.4 ms | 217.3 ms | 204.1 ms | 256.5 ms |
| Objective accuracy | 29/31 | 29/31 | 29/31 | 29/31 |
| Repetition | 0 | 0 | 0 | 0 |
| Lexical artifacts | 0 | 0 | 0 | 0 |
| Degeneration | 0 | 0 | 0 | 0 |
| Failures | 0 | 0 | 0 | 0 |
That smaller experiment measured:
71.0 ms median
for 1–16-token 128/16 responses.
The larger 2,000-response/config study is now the stronger latency reference and measured:
105.5 ms median
for that actual output-length bucket.
The earlier 71 ms result is retained as a smaller-study observation rather than used as the primary headline.
🔊 Why Instant Full-Text Voice?
A conventional low-latency voice pipeline often uses:
User Speech
↓
STT
↓
Autoregressive LLM
↓
Streaming Tokens
↓
Partial Text Chunks
↓
Streaming / Repeated TTS
↓
Speech
That architecture works, but often begins speech synthesis before the language model has completed the entire response.
The architecture being explored here is different:
User Speech
↓
STT
↓
DiffusionGemma
↓
COMPLETE RESPONSE
↓
Kokoro TTS
↓
COMPLETE AUDIO WAVEFORM
↓
Playback
The objective is not merely:
start speaking quickly
It is:
finish both the complete text and the complete speech waveform quickly.
In the demonstrated prototype turn:
Complete 109-token response
↓
123.1 ms
Complete 28.23-second audio generation
↓
117.8 ms
Complete response + complete audio ready
↓
260.1 ms measured total
So the measured sub-second result refers to:
the entire audio being generated and ready for playback
not simply the first audio arriving.
This matters because the TTS model receives the entire completed response before synthesis.
It can therefore see:
- complete punctuation
- full sentence structure
- all clauses
- the intended ending
- question / statement structure
- the complete semantic shape of the response
A possible improvement in prosody or continuity from this whole-text architecture remains a hypothesis and has not yet been directly compared against streaming TTS.
🧠 Complete Response, Not TTFT
The language-model latency measurements are also not conventional:
time to first token
Autoregressive models can expose the first token while the rest of the response is still being decoded.
For this project, the more relevant voice metric is:
time until the complete speakable response exists
In the larger short-response study:
- 1–16 tokens → 105.5 ms median
- 17–32 tokens → 119.8 ms
- 33–48 tokens → 145.5 ms
- 49–64 tokens → 175.7 ms
The complete response can then be sent to TTS.
In the demonstrated Kokoro prototype, TTS subsequently generated the entire audio waveform, not merely the beginning of it.
No matched autoregressive model was included in these experiments.
This release therefore does not claim universal superiority over autoregressive voice systems.
⚡ Instant Mode — 128 / 16
Recommended short-response configuration:
canvas_length = 128
max_denoising_steps = 16
t_max = 0.80
t_min = 0.40
entropy_bound = 0.1
confidence_threshold = 0.005
stability_threshold = 1
adaptive_stopping = true
Recommended for:
- Instant Full-Text TTS
- voice assistants
- short conversational replies
- acknowledgements
- short factual answers
- concise explanations
- short dialogue
Recommended output range:
1–64 tokens
Classification:
SHORT-TTS QUALITY PASS
🟦 Quality / General Mode — 256 / 48
Recommended general-quality configuration:
canvas_length = 256
max_denoising_steps = 48
t_max = 0.80
t_min = 0.40
entropy_bound = 0.1
confidence_threshold = 0.005
stability_threshold = 1
adaptive_stopping = true
long_form_capacity = 1280
Recommended for:
- general chat
- longer answers
- reasoning
- coding
- detailed explanations
- creative writing
- longer context
- long-form generation
- maximum refinement
This remains the safer configuration when output length extends beyond the intended Instant short-response envelope.
⚡ Sustained Performance
Performance V3 produced:
| Operating Point | Throughput |
|---|---|
| 48-step Quality-Max | 660.44 tok/s |
| 16-step single stream | 827.28 tok/s |
| 16-step concurrency-8 aggregate | 1,053.64 tok/s |
Measured on one:
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Runtime stack included:
- native packed NVFP4 experts
- compiled vLLM execution
- CUDA graphs
- FlashInfer CUTLASS NVFP4 MoE
- FlashInfer autotuning
- Triton attention
- FP8 e4m3 KV
- CUDA 13
- SM120
Important:
1,053.64 tok/s is concurrency-8 aggregate serving throughput.
It is not:
- the 48-step quality result
- a single-stream result
- complete-response latency
🏆 NVFP4 Quality Preservation
The NVFP4 conversion was designed to preserve E38 behavior while substantially reducing memory requirements.
Aligned classification:
PASS
| Evaluation | Base BF16 | E38 BF16 | E38 NVFP4 |
|---|---|---|---|
| Objective — 200 prompts | 128/200 — 64.0% | 134/200 — 67.0% | 137/200 — 68.5% |
| Regression subset — 100 | 62/100 | 58/100 | 64/100 |
| Multimodal | 20/20 | 20/20 | 20/20 |
| Multi-turn generations | 24/24 | 24/24 | 24/24 |
| Target refusal | 383/402 | 0/402 | 0/402 |
| Benign false refusal | 0/249 | 0/249 | 0/249 |
| Grammar errors /10k ↓ | 4.059 | 5.479 | 3.236 |
| Lexical artifacts /10k ↓ | 2.243 | 1.865 | 1.387 |
Direct objective comparison:
E38 BF16: 134 / 200 — 67.0%
E38 NVFP4: 137 / 200 — 68.5%
Difference:
+1.5 percentage points
with:
- 95% CI: −2.0 to +5.0 pp
- p = 0.5811
The difference was not statistically significant.
Conclusion:
NVFP4 preserved measured E38 BF16 quality.
The higher NVFP4 point estimate is not claimed as evidence that quantization inherently improved model capability.
📈 E38 BF16 Benchmark Reference
The larger public benchmark suite was run on Base BF16 and the frozen E38 BF16 parent.
| Benchmark | Base BF16 | E38 BF16 | Delta |
|---|---|---|---|
| IFEval | 67.10% | 64.70% | −2.40 pp |
| BBH | 71.62% | 73.96% | +2.34 pp |
| MuSR | 41.80% | 50.00% | +8.20 pp |
| MMLU-Pro | 49.61% | 51.57% | +1.96 pp |
| MATH Level 5 @1280 | 84.06% | 80.51% | −3.55 pp |
E38 is best described as:
a capability redistribution rather than a universally stronger checkpoint
The complete public benchmark suite has not been rerun directly on NVFP4.
➗ MATH Level 5 Regression
The MATH Level 5 result represents a real measured E38 BF16 regression.
Evaluation:
1,324 problems
| Model | Correct |
|---|---|
| Base BF16 | 1,113 / 1,324 — 84.06% |
| E38 BF16 | 1,066 / 1,324 — 80.51% |
Difference:
−3.55 percentage points
Statistics:
- 95% CI: −5.59 to −1.44 pp
- McNemar p: 0.00119
Pair structure:
- both correct: 988
- Base only: 125
- E38 only: 78
- both wrong: 133
This regression was established on E38 BF16.
The full MATH Level 5 benchmark has not been rerun directly on E38 NVFP4.
✍️ Long-Form Language Study
A prospective:
432-generation
study measured Base BF16, E38 BF16, and E38 NVFP4.
| Model | Grammar Errors /10k ↓ | Lexical Artifacts /10k ↓ |
|---|---|---|
| Base BF16 | 4.059 | 2.243 |
| E38 BF16 | 5.479 | 1.865 |
| E38 NVFP4 | 3.236 | 1.387 |
The study did not establish a statistically confirmed E38-specific grammar regression.
NVFP4 showed no evidence of language-quality degradation.
A notable positional pattern was observed:
errors were approximately 3.01× more concentrated in the final quarter of the 256-token canvas than in the first quarter.
Error-prone positions also showed:
- approximately +0.0504 final entropy
- approximately −0.0418 top-1/top-2 confidence margin
relative to matched clean controls.
This supports the diffusion-finalization hypothesis described earlier in this card.
💾 Memory Efficiency
| Model | Checkpoint Size |
|---|---|
| E38 BF16 | 51.68 GB |
| E38 NVFP4 | 18.86 GB |
Reduction:
32.82 GB
Percentage reduction:
63.5%
Relative size:
2.74× smaller
Routed experts use:
NVFP4 W4A4
with:
group size 16
This reduction is particularly useful for the one-GPU voice target because it leaves substantially more VRAM available for TTS and supporting runtime components.
🎯 One-GPU Voice Target
Target GPU:
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
Memory:
96 GB GDDR7
Physical memory bandwidth:
approximately 1.792 TB/s
E38 NVFP4 checkpoint:
18.86 GB
The intended system can potentially keep:
- DiffusionGemma
- TTS
- KV cache
- CUDA graphs
- audio components
- supporting runtime state
resident on one large Blackwell GPU.
The intended compute flow is primarily sequential:
DiffusionGemma generates complete text
████████
Kokoro synthesizes complete audio
█████████████████████
The working prototype demonstrates this architecture.
In the measured example, the entire 28.23-second waveform was synthesized before playback, rather than generated incrementally while the user was already listening.
Complete STT → LLM → TTS conversational latency and combined-residency behavior remain separate areas for characterization.
🔬 Native NVFP4 Execution
Performance validation confirmed:
NATIVE PACKED NVFP4 EXECUTION
No repeated:
- BF16 expert expansion
- FP4 expert repacking
- hidden full-precision expert conversion
was found in the hot inference path.
The NVFP4 release therefore provides actual runtime low-precision execution rather than only reducing checkpoint storage.
🔬 Precision Layout
NVFP4
Routed experts:
- NVFP4 W4A4
- group size 16
BF16
Retained:
- attention
- dense non-expert MLP
- routers
- embeddings
- LM head
- vision
- all 20 E38-modified tensors
Runtime KV
- FP8 e4m3
All 20 E38-modified tensors remain:
EXACT BF16
This is especially relevant to the grammar investigation:
the actual E38-modified tensors were not quantized to NVFP4.
⚙️ Runtime Stack
Validated stack:
- vLLM V2
- vLLM 0.27.1
- compiled execution
- CUDA graphs
- FULL_AND_PIECEWISE
- FlashInfer autotune
- FlashInfer CUTLASS NVFP4 MoE
- Triton attention
- FP8 e4m3 KV
- CUDA 13
- SM120
- RTX PRO 6000 Blackwell Workstation Edition
Preferred MoE backend:
FLASHINFER CUTLASS
Matched throughput testing:
- autotune ON → approximately 818–827 tok/s
- autotune OFF → approximately 789 tok/s
Autotune remains enabled.
🔵 What Is E38?
E38 was selected from more than:
80 controlled candidate configurations
derived from:
google/diffusiongemma-26B-A4B-it
E38 modifies:
- layers 7–16
attn.o_projmlp.down_proj- 20 language tensors
- 0 vision tensors
Measured weight difference:
- relative Frobenius difference:
0.0250308802854 - maximum absolute difference:
0.201904296875
All 20 modified tensors remain exact BF16 in this NVFP4 release.
E38 Overlay SHA256
9fe83b2fa2a3b7643f6dcc4a8360feb2cd96dcb9f70d215bd2e8f8f9d9d3044b
Frozen E38 Selection Hash
cd69540ec547f2f5cf181d527e8366fe60b49df7cafde38b434c5e42eeeb6175
No E38.1 checkpoint was created.
🧬 Model Lineage
Google DiffusionGemma 26B A4B BF16
↓
Controlled Abliteration Search
↓
E38 BF16
↓
E38 NVFP4
↓
51.68 GB → 18.86 GB
↓
660.44 tok/s Quality-Max
↓
827.28 tok/s Single Stream
↓
1,053.64 tok/s Aggregate
↓
Response-Length Research
↓
128/16 Short-Response Quality PASS
↓
Complete-Text → Complete-Audio Voice Prototype
↓
260.1 ms: COMPLETE TEXT + COMPLETE AUDIO READY
↓
DEPLOYMENT TARGET
INSTANT FULL-TEXT VOICE
LLM + TTS ON ONE GPU
🚦 Current Status
| Area | Status |
|---|---|
| E38 BF16 | Complete |
| E38 NVFP4 | Validated |
| NVFP4 integrity | PASS |
| Native packed NVFP4 inference | Validated |
| 256/48 Quality runtime | Validated |
| 16-step single-stream throughput | 827.28 tok/s |
| Aggregate serving throughput | 1,053.64 tok/s |
| Large 128/16 latency study | Validated |
| Blind short-response quality | PASS |
| Matched manual A/B pairs reviewed | 1,850 |
| 128/16 short-response recommendation | 1–64 tokens |
| 128/16 long-form | Degeneration warning |
| Diffusion-finalization grammar hypothesis | Supported by current evidence / not proven |
| DiffusionGemma → Kokoro TTS prototype | Working |
| Example complete LLM response | 123.1 ms |
| Example entire 28.23 s audio synthesis | 117.8 ms |
| Complete text + complete audio ready | 260.1 ms |
| Public voice demo | Available |
| Full microphone/STT/playback pipeline | Not yet characterized |
⚠️ Known Limitations
This model remains experimental.
Important considerations:
- refusal behavior is substantially reduced relative to upstream
- E38 showed weaker strict-format / instruction-following performance in some testing
- E38 BF16 showed a statistically significant MATH Level 5 regression
- the full public benchmark suite has not been rerun directly on NVFP4
- stochastic diffusion inference remains sensitive to inference configuration
- occasional grammar, malformed-word, repetition, and lexical errors occur in DiffusionGemma
- current evidence suggests these errors are more closely associated with diffusion finalization than NVFP4 quantization, but causality has not been proven
- 128/16 is specifically recommended for complete short responses
- 128/16 should not currently be treated as the general long-form mode
- long-form blind stress testing found severe degeneration and repetition concentrated in 128/16
- no severe degeneration was observed in the manually reviewed 1–64-token short-response set
- the demonstrated 109-token voice example is outside the formally validated 1–64-token short-quality envelope
- 260.1 ms refers to measured LLM generation plus complete TTS synthesis, not microphone/STT or playback startup
- the 117.8 ms TTS result represents synthesis of the entire 28.23-second waveform, not time to first audio
- no matched autoregressive model was included in the Instant latency experiments
- complete microphone → STT → LLM → TTS → playback latency has not yet been characterized
- possible TTS prosody benefits from whole-text generation remain a hypothesis
- 660.44 tok/s is the 48-step Quality-Max result
- 827.28 tok/s is the 16-step single-stream result
- 1,053.64 tok/s is concurrency-8 aggregate throughput
- the earlier 71 ms short-response result came from a smaller experiment; the larger validation measured 105.5 ms median for actual 1–16-token responses
Safety and Behavior
E38 intentionally retains substantially reduced refusal behavior.
| Model | Target Refusal | Prompt-Majority Refusal | Benign False Refusal |
|---|---|---|---|
| Base BF16 | 383/402 — 95.27% | 127/134 — 94.78% | 0/249 |
| E38 BF16 | 0/402 | 0/134 | 0/249 |
| E38 NVFP4 | 0/402 | 0/134 | 0/249 |
This checkpoint should not be interpreted as retaining the original refusal behavior of the upstream model.
Users should independently evaluate behavior and safeguards appropriate for their application.
🔐 Reproducibility
Original Upstream
google/diffusiongemma-26B-A4B-it
Pinned revision:
f7f5b7f5fa82ffc52addd066915886d497f5517b
E38 BF16 Parent
Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16
E38 Overlay SHA256
9fe83b2fa2a3b7643f6dcc4a8360feb2cd96dcb9f70d215bd2e8f8f9d9d3044b
Frozen E38 Selection Hash
cd69540ec547f2f5cf181d527e8366fe60b49df7cafde38b434c5e42eeeb6175
NVFP4 Integrity
Validation confirmed:
- 12 / 12 artifacts PASS
- E38 overlay hash PASS
- all 20 E38-modified tensors remain exact BF16
- validated quantization unchanged
- vision unchanged
- BF16-protected paths unchanged
Intended Use
This checkpoint is intended for:
- Instant Full-Text TTS
- complete-text → complete-audio voice generation
- single-GPU LLM + TTS research
- local voice assistants
- complete-response conversational generation
- low-latency chat
- local inference
- diffusion-LM research
- response-length / diffusion-budget research
- diffusion-finalization research
- NVFP4 deployment
- Blackwell inference
- quantization research
- abliteration research
- multimodal experimentation
- inference optimization
- serving-throughput research
- memory-efficiency research
Upstream Models
Immediate parent:
Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16
Original upstream:
google/diffusiongemma-26B-A4B-it
License
Apache License 2.0.
This model is a quantized derivative of the E38 BF16 checkpoint, itself a modified derivative of:
google/diffusiongemma-26B-A4B-it
