← back to catalog · registered 2026-08-22 13:56

Goodoldjam/DiffusionGemma-26B-E38-Abliterated-NVFP4

Goodoldjam Gemma 11B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Goodoldjam%2FDiffusionGemma-26B-E38-Abliterated-NVFP4"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 710
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
710
73 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-16
Downloads over time
Now727→from469↑55%
456555654753469 on Aug 19727 on Oct 11727 on Oct 9AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 155 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors diffusion_gemma image-text-to-text diffusion diffusion-lm diffusiongemma gemma abliteration abliterated multimodal nvfp4

Related

Total size
17.5 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-23 22:04

Files by quantization

Auxiliary files 14 files 17.6 GB
model-00001-of-00002.safetensors 9.32 GB 000f5246 download
model-00002-of-00002.safetensors 8.22 GB 6d2fab64 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 4.44 MB bff99180 download
.quant_summary.txt 3.53 MB cff741b9 download
README.md 32.5 KB 154634f5 download
chat_template.jinja 18.1 KB 8c09d23f download
config.json 14.5 KB bacc91ee download
LICENSE 11.1 KB d6456956 download
hf_quant_config.json 9.93 KB 2b9a4ab5 download
tokenizer_config.json 2.68 KB db7e70c7 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 388 B e07e9c51 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    base_model:
  • Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • diffusion
  • diffusion-lm
  • diffusiongemma
  • gemma
  • abliteration
  • abliterated
  • heretic
  • multimodal
  • nvfp4
  • quantization
  • mixed-precision
  • blackwell
  • cuda-graphs
  • vllm
  • tts
  • voice-assistant
  • low-latency

DiffusionGemma 26B E38 — Abliterated NVFP4

A 26B-class diffusion language model optimized for fast complete-response generation and full-text voice.

The latest testing produced several interesting results.

For short responses, less diffusion performed better.

The faster:

128-token canvas / 16-step

configuration was not only faster than:

256-token canvas / 48-step

but also produced a higher strictly-clean response rate in blind manual A/B testing inside the intended 1–64-token short-response range.

The result changes for longer generation, where 256/48 remains safer.

A second finding is that DiffusionGemma's occasional grammar, malformed-word, repetition, and lexical errors increasingly appear more consistent with:

diffusion finalization behavior

than with:

NVFP4 quantization damage

And the original deployment concept is now working in a real local voice prototype:

DiffusionGemma → complete response → complete generated audio

In a measured prototype turn:

LLM complete: 123.1 ms

Entire TTS audio generated: 117.8 ms

Complete text + complete audio ready in 260.1 ms

The TTS number is not time-to-first-audio.

The entire 28.23-second audio waveform was already synthesized and ready for playback at the end of the measured TTS stage.

Current recommendation:

Use Case Recommended Mode
Short replies / Full-Text TTS / 1–64 tokens 128/16
Longer responses / general chat / reasoning 256/48

Primary deployment target:

Instant Full-Text Voice on One GPU


🎥 Working Voice Demo

DiffusionGemma E38 NVFP4 Instant Full-Text Voice Demo

▶ Watch the DiffusionGemma E38 NVFP4 Voice Demo on YouTube

The working local prototype uses:

DiffusionGemma E38 NVFP4
        ↓
Complete Response
        ↓
Kokoro TTS
        ↓
Complete Audio Waveform
        ↓
Ready for Playback

One measured prototype run produced:

Stage Result
LLM complete-response latency 123.1 ms
TTS — entire audio synthesized 117.8 ms
Complete text + complete audio ready 260.1 ms
Response length 109 tokens
LLM throughput 885.5 tok/s
Fully generated audio duration 28.23 seconds

The full 28.23-second audio output was generated in 117.8 ms.

This is an important distinction.

The TTS measurement is not:

  • time to first audio sample
  • time to begin streaming speech
  • time until the first sentence is synthesized

It is the measured time required to synthesize the entire audio output for that response.

Likewise, the 260.1 ms pipeline measurement means that by the end of that measured interval:

  1. DiffusionGemma had generated the complete text response.
  2. Kokoro had received the complete response.
  3. Kokoro had synthesized the complete 28.23-second audio waveform.
  4. The audio was ready for playback.

So for this measured turn:

109-token complete response + 28.23 seconds of finished speech were ready in 260.1 ms of measured LLM + TTS processing.

This does not include microphone capture, STT, playback-device startup, UI overhead, or the real-world time required to actually listen to 28.23 seconds of speech.

The 109-token response is also outside the formally validated 1–64-token short-quality envelope, so this example is presented as a:

voice-pipeline performance demonstration

rather than evidence extending the formal short-response quality result to 109 tokens.


🗣️ Voice Demo Pre-Prompt

The voice prototype uses a short conversational pre-prompt designed to keep responses concise, natural, and suitable for complete-text TTS.

You are a fast local voice assistant.

Answer the user's request directly.

Use natural spoken English and complete sentences.

Keep normal conversational answers short and focused.

Avoid markdown, bullet points, headings, citations, emojis, role labels, stage directions, and meta commentary unless the user explicitly asks for them.

Do not open with filler such as "Okay", "Sure", or "Of course".

Do not explain what you are doing.

Prefer wording that sounds natural when read aloud by TTS.

For ordinary voice-chat questions, aim for a complete response within roughly 1–64 output tokens.

If the request genuinely requires more detail, answer fully rather than forcing an unnaturally short response.

Return only the text that should be spoken.

For the 128/16 Instant Full-Text TTS mode, the intended operating range is:

approximately 1–64 output tokens

This is a target for normal conversational replies, not a hard truncation rule.

Longer or more detailed responses should use the 256/48 quality configuration.


🚀 Model

Goodoldjam/DiffusionGemma-26B-E38-Abliterated-NVFP4

Donate on Ko-fi

E38 NVFP4 reduces the E38 BF16 checkpoint from:

51.68 GB → 18.86 GB

while preserving the measured E38 capability profile and enabling native low-precision inference on NVIDIA Blackwell.

E38 NVFP4 RTX PRO 6000 Blackwell Short TTS Quality Pass Full Audio Ready Single Stream Aggregate Throughput

The expensive part is not making another quant. The expensive part is proving which quant and inference configuration is actually worth using.


🧪 Response-Length Validation

128/16 and 256/48 were evaluated across both short conversational responses and longer-form generation.

The result was strongly dependent on response length.

Short Responses — 1–64 Tokens

The intended Instant Full-Text TTS operating range was tested using blind matched A/B manual review.

Manual review completed so far:

1,850 matched A/B pairs

representing:

3,700 manually reviewed responses

Within the intended 1–64-token range:

Short-Response Result 128/16 256/48
In-range responses reviewed 1,705 1,690
Words reviewed 46,282 45,227
Strictly clean responses 1,678 1,638
Strictly clean rate 98.42% 96.92%
Responses with confirmed language issue 27 52
Severe degeneration 0 0
Truncation 0 0

A response was considered strictly clean only if it had:

  • no confirmed grammar error
  • no malformed lexical artifact
  • no accidental repetition
  • no degeneration
  • no malformed ending
  • no truncation

Awkward-but-grammatical wording and factual/content mistakes were tracked separately.

Fully Paired Short-Response Result

Across:

1,633 matched A/B pairs

where both configurations stayed inside the 1–64-token range:

Paired Outcome Count
Both clean 1,563
128/16 clean / 256/48 issue 48
256/48 clean / 128/16 issue 21
Both issue 1

Paired strictly-clean rates:

128/16: 98.65%

256/48: 97.00%

Difference:

+1.65 percentage points in favor of 128/16

Paired bootstrap:

95% CI: approximately +0.67 to +2.63 percentage points

Exact paired test:

McNemar p ≈ 0.0016

The predeclared short-response non-inferiority margin was:

−2 percentage points

128/16 cleared that requirement.

Conclusion:

128/16 SHORT-TTS QUALITY PASS

This does not establish universal superiority.

It means:

128/16 performed better in the manually reviewed short-response operating range.


Longer Responses

The result changes when 128/16 is pushed beyond its intended short-response role.

A separate blind stress test used approximately:

240-word responses

Results:

Long-Form Result 128/16 256/48
Responses reviewed 25 25
Definite grammar-error responses 7 5
Repetition-affected responses 4 0
Severe degeneration 2 0
Strictly clean responses 15 17

Both severe degeneration cases came from:

128/16

The current deployment boundary is therefore:

128/16 → short complete responses

256/48 → general and long-form responses

The evidence does not suggest one universally optimal diffusion configuration.

Instead:

the best diffusion budget appears to depend on response length.


🧠 Grammar Errors and the Diffusion-Finalization Hypothesis

DiffusionGemma can occasionally produce unusual language artifacts such as:

  • malformed words
  • duplicated words
  • incorrect local grammar
  • awkward substitutions
  • malformed sentence endings
  • rare degeneration

These issues raised an obvious question:

Did NVFP4 quantization damage the language model?

The current evidence does not support that explanation.

Instead, several observations increasingly point toward:

diffusion inference and final text resolution

as a more likely source.

This remains a working research hypothesis rather than a proven causal mechanism.

The Error Class Exists Before NVFP4

Similar rare language failures were observed in:

  • Base BF16
  • E38 BF16
  • E38 NVFP4

The general phenomenon therefore existed before the NVFP4 conversion.

NVFP4 Did Not Increase Measured Language Errors

A larger long-form language-quality study measured:

Model Grammar Errors /10k ↓ Lexical Artifacts /10k ↓
Base BF16 4.059 2.243
E38 BF16 5.479 1.865
E38 NVFP4 3.236 1.387

NVFP4 had the lowest observed grammar and lexical-artifact rates in that study.

The E38 BF16 grammar difference relative to Base was:

+1.420 grammar errors /10k words

with:

95% CI: −0.674 to +3.541

and:

McNemar p = 0.560

The E38 BF16 grammar difference was therefore not established as a statistically confirmed E38-specific regression.

The NVFP4 results provide:

no evidence that quantization introduced additional language corruption

Errors Cluster Late in the Diffusion Canvas

One of the most interesting findings came from looking at where errors occur.

Language errors were approximately:

3.01× more concentrated in the final quarter of the 256-token canvas

than in the first quarter.

Error-prone positions also showed approximately:

+0.0504 higher final entropy

and:

−0.0418 lower top-1 / top-2 confidence margin

relative to matched clean positions.

This suggests the model is less certain when resolving some problematic final text.

More Denoising Did Not Reliably Fix It

Simply increasing the denoising budget did not establish a consistent overall language-quality improvement.

The short-response experiments add another clue:

128/16 was cleaner than 256/48 in the current blind short-response sample.

If more denoising automatically produced better language, this would not be the expected result.

But reducing the diffusion budget too aggressively for long responses creates another failure mode:

repetition and degeneration

This suggests there may be a useful operating region where the diffusion process has enough compute to resolve a response cleanly without unnecessary additional refinement.

Current Working Hypothesis

The current evidence suggests:

DiffusionGemma may know more language than its inference process always cleanly finalizes.

A simplified interpretation:

Underlying learned capability
          ↓
Partially denoised response
          ↓
Mostly correct structure
          ↓
Late diffusion resolution
          ↓
Local uncertainty
          ↓
Rare malformed / repeated / awkward token

This could explain why:

  • the same general error class exists in BF16
  • NVFP4 does not make it worse
  • errors cluster late in the canvas
  • error positions show higher entropy
  • error positions show lower confidence margins
  • more denoising does not always improve output
  • reduced diffusion improves short-response cleanliness
  • insufficient diffusion capacity hurts long-form generation

The current evidence therefore points more strongly toward:

diffusion finalization behavior

than toward:

NVFP4 quantization damage

as the source of these rare grammar and lexical issues.

This remains an active research question.


⚡ Large Short-Response Latency Validation

The larger short-response experiment used:

500 frozen prompts

across:

15 voice-assistant categories

with:

4 matched seeds

for a total of:

2,000 matched pairs

and:

4,000 total generations

Generated words:

Configuration Words
128/16 58,355
256/48 58,077

Runtime failures:

0

Empty outputs:

0

The primary deployment analysis focused on outputs with actual generated length:

1–64 tokens

Counts:

Configuration Responses in 1–64 Range
128/16 1,843
256/48 1,827

Pairs where both configurations remained in range:

1,766

Large-Study Latency

Metric 128/16 256/48
Median latency 139.2 ms 158.1 ms
P90 latency 190.1 ms 212.8 ms

128/16 reduced:

median latency by approximately 11.9%

and:

P90 latency by approximately 10.6%

versus 256/48.

Latency by Actual Output Length

Actual Output Length 128/16 Median 256/48 Median
1–16 tokens 105.5 ms 116.3 ms
17–32 tokens 119.8 ms 132.6 ms
33–48 tokens 145.5 ms 163.5 ms
49–64 tokens 175.7 ms 195.5 ms

128/16 was faster in every measured primary short-response bucket.

Recommended deployment range:

1–64 output tokens


⚡ Earlier Instant Chat Experiment

An earlier smaller matched experiment compared:

  • 256 / 48
  • 256 / 16
  • 128 / 16
  • 64 / 16

Results:

Metric 256/48 256/16 128/16 64/16
Median latency 141.0 ms 131.3 ms 120.1 ms 134.2 ms
Mean latency 146.6 ms 139.3 ms 124.7 ms 152.1 ms
P90 216.1 ms 204.2 ms 187.3 ms 244.0 ms
P95 251.4 ms 217.3 ms 204.1 ms 256.5 ms
Objective accuracy 29/31 29/31 29/31 29/31
Repetition 0 0 0 0
Lexical artifacts 0 0 0 0
Degeneration 0 0 0 0
Failures 0 0 0 0

That smaller experiment measured:

71.0 ms median

for 1–16-token 128/16 responses.

The larger 2,000-response/config study is now the stronger latency reference and measured:

105.5 ms median

for that actual output-length bucket.

The earlier 71 ms result is retained as a smaller-study observation rather than used as the primary headline.


🔊 Why Instant Full-Text Voice?

A conventional low-latency voice pipeline often uses:

User Speech
    ↓
STT
    ↓
Autoregressive LLM
    ↓
Streaming Tokens
    ↓
Partial Text Chunks
    ↓
Streaming / Repeated TTS
    ↓
Speech

That architecture works, but often begins speech synthesis before the language model has completed the entire response.

The architecture being explored here is different:

User Speech
    ↓
STT
    ↓
DiffusionGemma
    ↓
COMPLETE RESPONSE
    ↓
Kokoro TTS
    ↓
COMPLETE AUDIO WAVEFORM
    ↓
Playback

The objective is not merely:

start speaking quickly

It is:

finish both the complete text and the complete speech waveform quickly.

In the demonstrated prototype turn:

Complete 109-token response
        ↓
123.1 ms

Complete 28.23-second audio generation
        ↓
117.8 ms

Complete response + complete audio ready
        ↓
260.1 ms measured total

So the measured sub-second result refers to:

the entire audio being generated and ready for playback

not simply the first audio arriving.

This matters because the TTS model receives the entire completed response before synthesis.

It can therefore see:

  • complete punctuation
  • full sentence structure
  • all clauses
  • the intended ending
  • question / statement structure
  • the complete semantic shape of the response

A possible improvement in prosody or continuity from this whole-text architecture remains a hypothesis and has not yet been directly compared against streaming TTS.


🧠 Complete Response, Not TTFT

The language-model latency measurements are also not conventional:

time to first token

Autoregressive models can expose the first token while the rest of the response is still being decoded.

For this project, the more relevant voice metric is:

time until the complete speakable response exists

In the larger short-response study:

  • 1–16 tokens → 105.5 ms median
  • 17–32 tokens → 119.8 ms
  • 33–48 tokens → 145.5 ms
  • 49–64 tokens → 175.7 ms

The complete response can then be sent to TTS.

In the demonstrated Kokoro prototype, TTS subsequently generated the entire audio waveform, not merely the beginning of it.

No matched autoregressive model was included in these experiments.

This release therefore does not claim universal superiority over autoregressive voice systems.


⚡ Instant Mode — 128 / 16

Recommended short-response configuration:

canvas_length = 128
max_denoising_steps = 16

t_max = 0.80
t_min = 0.40
entropy_bound = 0.1
confidence_threshold = 0.005
stability_threshold = 1
adaptive_stopping = true

Recommended for:

  • Instant Full-Text TTS
  • voice assistants
  • short conversational replies
  • acknowledgements
  • short factual answers
  • concise explanations
  • short dialogue

Recommended output range:

1–64 tokens

Classification:

SHORT-TTS QUALITY PASS


🟦 Quality / General Mode — 256 / 48

Recommended general-quality configuration:

canvas_length = 256
max_denoising_steps = 48

t_max = 0.80
t_min = 0.40
entropy_bound = 0.1
confidence_threshold = 0.005
stability_threshold = 1
adaptive_stopping = true
long_form_capacity = 1280

Recommended for:

  • general chat
  • longer answers
  • reasoning
  • coding
  • detailed explanations
  • creative writing
  • longer context
  • long-form generation
  • maximum refinement

This remains the safer configuration when output length extends beyond the intended Instant short-response envelope.


⚡ Sustained Performance

Performance V3 produced:

Operating Point Throughput
48-step Quality-Max 660.44 tok/s
16-step single stream 827.28 tok/s
16-step concurrency-8 aggregate 1,053.64 tok/s

Measured on one:

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Runtime stack included:

  • native packed NVFP4 experts
  • compiled vLLM execution
  • CUDA graphs
  • FlashInfer CUTLASS NVFP4 MoE
  • FlashInfer autotuning
  • Triton attention
  • FP8 e4m3 KV
  • CUDA 13
  • SM120

Important:

1,053.64 tok/s is concurrency-8 aggregate serving throughput.

It is not:

  • the 48-step quality result
  • a single-stream result
  • complete-response latency

🏆 NVFP4 Quality Preservation

The NVFP4 conversion was designed to preserve E38 behavior while substantially reducing memory requirements.

Aligned classification:

PASS

Evaluation Base BF16 E38 BF16 E38 NVFP4
Objective — 200 prompts 128/200 — 64.0% 134/200 — 67.0% 137/200 — 68.5%
Regression subset — 100 62/100 58/100 64/100
Multimodal 20/20 20/20 20/20
Multi-turn generations 24/24 24/24 24/24
Target refusal 383/402 0/402 0/402
Benign false refusal 0/249 0/249 0/249
Grammar errors /10k ↓ 4.059 5.479 3.236
Lexical artifacts /10k ↓ 2.243 1.865 1.387

Direct objective comparison:

E38 BF16: 134 / 200 — 67.0%

E38 NVFP4: 137 / 200 — 68.5%

Difference:

+1.5 percentage points

with:

  • 95% CI: −2.0 to +5.0 pp
  • p = 0.5811

The difference was not statistically significant.

Conclusion:

NVFP4 preserved measured E38 BF16 quality.

The higher NVFP4 point estimate is not claimed as evidence that quantization inherently improved model capability.


📈 E38 BF16 Benchmark Reference

The larger public benchmark suite was run on Base BF16 and the frozen E38 BF16 parent.

Benchmark Base BF16 E38 BF16 Delta
IFEval 67.10% 64.70% −2.40 pp
BBH 71.62% 73.96% +2.34 pp
MuSR 41.80% 50.00% +8.20 pp
MMLU-Pro 49.61% 51.57% +1.96 pp
MATH Level 5 @1280 84.06% 80.51% −3.55 pp

E38 is best described as:

a capability redistribution rather than a universally stronger checkpoint

The complete public benchmark suite has not been rerun directly on NVFP4.


➗ MATH Level 5 Regression

The MATH Level 5 result represents a real measured E38 BF16 regression.

Evaluation:

1,324 problems

Model Correct
Base BF16 1,113 / 1,324 — 84.06%
E38 BF16 1,066 / 1,324 — 80.51%

Difference:

−3.55 percentage points

Statistics:

  • 95% CI: −5.59 to −1.44 pp
  • McNemar p: 0.00119

Pair structure:

  • both correct: 988
  • Base only: 125
  • E38 only: 78
  • both wrong: 133

This regression was established on E38 BF16.

The full MATH Level 5 benchmark has not been rerun directly on E38 NVFP4.


✍️ Long-Form Language Study

A prospective:

432-generation

study measured Base BF16, E38 BF16, and E38 NVFP4.

Model Grammar Errors /10k ↓ Lexical Artifacts /10k ↓
Base BF16 4.059 2.243
E38 BF16 5.479 1.865
E38 NVFP4 3.236 1.387

The study did not establish a statistically confirmed E38-specific grammar regression.

NVFP4 showed no evidence of language-quality degradation.

A notable positional pattern was observed:

errors were approximately 3.01× more concentrated in the final quarter of the 256-token canvas than in the first quarter.

Error-prone positions also showed:

  • approximately +0.0504 final entropy
  • approximately −0.0418 top-1/top-2 confidence margin

relative to matched clean controls.

This supports the diffusion-finalization hypothesis described earlier in this card.


💾 Memory Efficiency

Model Checkpoint Size
E38 BF16 51.68 GB
E38 NVFP4 18.86 GB

Reduction:

32.82 GB

Percentage reduction:

63.5%

Relative size:

2.74× smaller

Routed experts use:

NVFP4 W4A4

with:

group size 16

This reduction is particularly useful for the one-GPU voice target because it leaves substantially more VRAM available for TTS and supporting runtime components.


🎯 One-GPU Voice Target

Target GPU:

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Memory:

96 GB GDDR7

Physical memory bandwidth:

approximately 1.792 TB/s

E38 NVFP4 checkpoint:

18.86 GB

The intended system can potentially keep:

  • DiffusionGemma
  • TTS
  • KV cache
  • CUDA graphs
  • audio components
  • supporting runtime state

resident on one large Blackwell GPU.

The intended compute flow is primarily sequential:

DiffusionGemma generates complete text
████████

             Kokoro synthesizes complete audio
             █████████████████████

The working prototype demonstrates this architecture.

In the measured example, the entire 28.23-second waveform was synthesized before playback, rather than generated incrementally while the user was already listening.

Complete STT → LLM → TTS conversational latency and combined-residency behavior remain separate areas for characterization.


🔬 Native NVFP4 Execution

Performance validation confirmed:

NATIVE PACKED NVFP4 EXECUTION

No repeated:

  • BF16 expert expansion
  • FP4 expert repacking
  • hidden full-precision expert conversion

was found in the hot inference path.

The NVFP4 release therefore provides actual runtime low-precision execution rather than only reducing checkpoint storage.


🔬 Precision Layout

NVFP4

Routed experts:

  • NVFP4 W4A4
  • group size 16

BF16

Retained:

  • attention
  • dense non-expert MLP
  • routers
  • embeddings
  • LM head
  • vision
  • all 20 E38-modified tensors

Runtime KV

  • FP8 e4m3

All 20 E38-modified tensors remain:

EXACT BF16

This is especially relevant to the grammar investigation:

the actual E38-modified tensors were not quantized to NVFP4.


⚙️ Runtime Stack

Validated stack:

  • vLLM V2
  • vLLM 0.27.1
  • compiled execution
  • CUDA graphs
  • FULL_AND_PIECEWISE
  • FlashInfer autotune
  • FlashInfer CUTLASS NVFP4 MoE
  • Triton attention
  • FP8 e4m3 KV
  • CUDA 13
  • SM120
  • RTX PRO 6000 Blackwell Workstation Edition

Preferred MoE backend:

FLASHINFER CUTLASS

Matched throughput testing:

  • autotune ON → approximately 818–827 tok/s
  • autotune OFF → approximately 789 tok/s

Autotune remains enabled.


🔵 What Is E38?

E38 was selected from more than:

80 controlled candidate configurations

derived from:

google/diffusiongemma-26B-A4B-it

E38 modifies:

  • layers 7–16
  • attn.o_proj
  • mlp.down_proj
  • 20 language tensors
  • 0 vision tensors

Measured weight difference:

  • relative Frobenius difference: 0.0250308802854
  • maximum absolute difference: 0.201904296875

All 20 modified tensors remain exact BF16 in this NVFP4 release.

E38 Overlay SHA256

9fe83b2fa2a3b7643f6dcc4a8360feb2cd96dcb9f70d215bd2e8f8f9d9d3044b

Frozen E38 Selection Hash

cd69540ec547f2f5cf181d527e8366fe60b49df7cafde38b434c5e42eeeb6175

No E38.1 checkpoint was created.


🧬 Model Lineage

Google DiffusionGemma 26B A4B BF16

↓

Controlled Abliteration Search

↓

E38 BF16

↓

E38 NVFP4

↓

51.68 GB → 18.86 GB

↓

660.44 tok/s Quality-Max

↓

827.28 tok/s Single Stream

↓

1,053.64 tok/s Aggregate

↓

Response-Length Research

↓

128/16 Short-Response Quality PASS

↓

Complete-Text → Complete-Audio Voice Prototype

↓

260.1 ms: COMPLETE TEXT + COMPLETE AUDIO READY

↓

DEPLOYMENT TARGET

INSTANT FULL-TEXT VOICE

LLM + TTS ON ONE GPU


🚦 Current Status

Area Status
E38 BF16 Complete
E38 NVFP4 Validated
NVFP4 integrity PASS
Native packed NVFP4 inference Validated
256/48 Quality runtime Validated
16-step single-stream throughput 827.28 tok/s
Aggregate serving throughput 1,053.64 tok/s
Large 128/16 latency study Validated
Blind short-response quality PASS
Matched manual A/B pairs reviewed 1,850
128/16 short-response recommendation 1–64 tokens
128/16 long-form Degeneration warning
Diffusion-finalization grammar hypothesis Supported by current evidence / not proven
DiffusionGemma → Kokoro TTS prototype Working
Example complete LLM response 123.1 ms
Example entire 28.23 s audio synthesis 117.8 ms
Complete text + complete audio ready 260.1 ms
Public voice demo Available
Full microphone/STT/playback pipeline Not yet characterized

⚠️ Known Limitations

This model remains experimental.

Important considerations:

  • refusal behavior is substantially reduced relative to upstream
  • E38 showed weaker strict-format / instruction-following performance in some testing
  • E38 BF16 showed a statistically significant MATH Level 5 regression
  • the full public benchmark suite has not been rerun directly on NVFP4
  • stochastic diffusion inference remains sensitive to inference configuration
  • occasional grammar, malformed-word, repetition, and lexical errors occur in DiffusionGemma
  • current evidence suggests these errors are more closely associated with diffusion finalization than NVFP4 quantization, but causality has not been proven
  • 128/16 is specifically recommended for complete short responses
  • 128/16 should not currently be treated as the general long-form mode
  • long-form blind stress testing found severe degeneration and repetition concentrated in 128/16
  • no severe degeneration was observed in the manually reviewed 1–64-token short-response set
  • the demonstrated 109-token voice example is outside the formally validated 1–64-token short-quality envelope
  • 260.1 ms refers to measured LLM generation plus complete TTS synthesis, not microphone/STT or playback startup
  • the 117.8 ms TTS result represents synthesis of the entire 28.23-second waveform, not time to first audio
  • no matched autoregressive model was included in the Instant latency experiments
  • complete microphone → STT → LLM → TTS → playback latency has not yet been characterized
  • possible TTS prosody benefits from whole-text generation remain a hypothesis
  • 660.44 tok/s is the 48-step Quality-Max result
  • 827.28 tok/s is the 16-step single-stream result
  • 1,053.64 tok/s is concurrency-8 aggregate throughput
  • the earlier 71 ms short-response result came from a smaller experiment; the larger validation measured 105.5 ms median for actual 1–16-token responses

Safety and Behavior

E38 intentionally retains substantially reduced refusal behavior.

Model Target Refusal Prompt-Majority Refusal Benign False Refusal
Base BF16 383/402 — 95.27% 127/134 — 94.78% 0/249
E38 BF16 0/402 0/134 0/249
E38 NVFP4 0/402 0/134 0/249

This checkpoint should not be interpreted as retaining the original refusal behavior of the upstream model.

Users should independently evaluate behavior and safeguards appropriate for their application.


🔐 Reproducibility

Original Upstream

google/diffusiongemma-26B-A4B-it

Pinned revision:

f7f5b7f5fa82ffc52addd066915886d497f5517b

E38 BF16 Parent

Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16

E38 Overlay SHA256

9fe83b2fa2a3b7643f6dcc4a8360feb2cd96dcb9f70d215bd2e8f8f9d9d3044b

Frozen E38 Selection Hash

cd69540ec547f2f5cf181d527e8366fe60b49df7cafde38b434c5e42eeeb6175

NVFP4 Integrity

Validation confirmed:

  • 12 / 12 artifacts PASS
  • E38 overlay hash PASS
  • all 20 E38-modified tensors remain exact BF16
  • validated quantization unchanged
  • vision unchanged
  • BF16-protected paths unchanged

Intended Use

This checkpoint is intended for:

  • Instant Full-Text TTS
  • complete-text → complete-audio voice generation
  • single-GPU LLM + TTS research
  • local voice assistants
  • complete-response conversational generation
  • low-latency chat
  • local inference
  • diffusion-LM research
  • response-length / diffusion-budget research
  • diffusion-finalization research
  • NVFP4 deployment
  • Blackwell inference
  • quantization research
  • abliteration research
  • multimodal experimentation
  • inference optimization
  • serving-throughput research
  • memory-efficiency research

Upstream Models

Immediate parent:

Goodoldjam/DiffusionGemma-26B-E38-Abliterated-BF16

Original upstream:

google/diffusiongemma-26B-A4B-it


License

Apache License 2.0.

This model is a quantized derivative of the E38 BF16 checkpoint, itself a modified derivative of:

google/diffusiongemma-26B-A4B-it

README history 20 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-23Update README.md3f0071719.8 KB
    Loading...
  2. 2026-08-23Update README.md2d3c5dd5 KB
    Loading...
  3. 2026-08-22Update README.mdd9cbb6532.5 KB
    Loading...
  4. 2026-08-22Update README.md38058cd30.5 KB
    Loading...
  5. 2026-08-22Update README.md2a2fa6529.3 KB
    Loading...
  6. 2026-08-22Update README.md3d52fc314.4 KB
    Loading...
  7. 2026-08-22Update README.mdeccab9629.5 KB
    Loading...
  8. 2026-08-22Update README.md87447e829.2 KB
    Loading...
  9. 2026-08-22Update README.md768584123.7 KB
    Loading...
  10. 2026-08-22Update README.mda667fbf20.3 KB
    Loading...
  11. 2026-08-21Update README.md280320034.2 KB
    Loading...
  12. 2026-08-21Update README.md97b70d425.8 KB
    Loading...
  13. 2026-08-21Update README.md5e089e228.6 KB
    Loading...
  14. 2026-08-21Update README.md5efa6f84.4 KB
    Loading...
  15. 2026-08-21Update README.md312db9610.5 KB
    Loading...
  16. 2026-08-21Update README.mdaa013b042.8 KB
    Loading...
  17. 2026-08-18Update README.md243047811.9 KB
    Loading...
  18. 2026-08-18Update README.md07fc6619.8 KB
    Loading...
  19. 2026-08-18Update README.mdaa0b8c07.3 KB
    Loading...
  20. 2026-08-18Update README.md537d85e5.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration