Google and NVIDIA Open-Source DiffusionGemma, Speeds Up Single-Card Inference by 4x
On June 10, 2026, Google launched DiffusionGemma, an experimental open-source language model that breaks from the traditional autoregressive paradigm of generating text one word at a time. It pioneers the use of diffusion mechanisms from image AI in text generation, starting from random noise and refining iteratively to output 256 token blocks in parallel per step.

In terms of hardware performance, NVIDIA's deep optimizations enable the model to run nearly four times faster than comparable traditional models in single‑GPU single‑user mode. On an H100 GPU, it achieves output speeds of up to 1,000 tokens per second for a single request, and even on high‑end consumer GPUs like the RTX 5090 it can exceed 700 tokens per second.
DiffusionGemma has 26 billion parameters and uses a mixture‑of‑experts (MoE) architecture, activating only 3.8 billion parameters per step. While its text generation quality and accuracy slightly trail traditional Gemma4 series models on standard benchmarks, its unique “full‑block awareness” overcomes the backward‑only limitation of autoregressive models. Because all tokens can reference each other during generation, it excels in tasks involving nonlinear and structured data, such as text completion, code filling, Sudoku solving, and amino acid sequence processing.

The model weights are now open‑sourced on Hugging Face under the Apache 2.0 license and work seamlessly with mainstream inference frameworks like vLLM and MLX. This exploration not only removes memory bandwidth constraints on GPU compute power but also opens a new technical path for future AI applications in complex logic and nonlinear text generation.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
On June 10, 2026, Google launched DiffusionGemma, an experimental open-source language model that breaks from the traditional autoregressive paradigm of generating text one word at a time. It pioneers the use of diffusion mechanisms from image AI in text generation, starting from random noise and refining iteratively to output 256 token blocks in parallel per step.

In terms of hardware performance, NVIDIA's deep optimizations enable the model to run nearly four times faster than comparable traditional models in single‑GPU single‑user mode. On an H100 GPU, it achieves output speeds of up to 1,000 tokens per second for a single request, and even on high‑end consumer GPUs like the RTX 5090 it can exceed 700 tokens per second.
DiffusionGemma has 26 billion parameters and uses a mixture‑of‑experts (MoE) architecture, activating only 3.8 billion parameters per step. While its text generation quality and accuracy slightly trail traditional Gemma4 series models on standard benchmarks, its unique “full‑block awareness” overcomes the backward‑only limitation of autoregressive models. Because all tokens can reference each other during generation, it excels in tasks involving nonlinear and structured data, such as text completion, code filling, Sudoku solving, and amino acid sequence processing.

The model weights are now open‑sourced on Hugging Face under the Apache 2.0 license and work seamlessly with mainstream inference frameworks like vLLM and MLX. This exploration not only removes memory bandwidth constraints on GPU compute power but also opens a new technical path for future AI applications in complex logic and nonlinear text generation.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






