- Google LLC today released DiffusionGemma, a large language model based on an emerging machine learning approach known as text diffusion.
- The algorithm can generate text four times faster than traditional LLMs.
- DiffusionGemma can generate more than 1,000 tokens per second when running on a single H100, a server-grade GPU that Nvidia Corp. launched in 2022.
- The model can generate over 700 tokens per second on the chipmaker’s desktop-grade GeForce RTX 5090 chip.
- The model includes 26 billion parameters but activates only 3.8 billion of them to answer the prompt, which lowers memory usage.
- DiffusionGemma is based on an LLM called Gemma 4 26B A4B that Google released in April.
- The search giant replaced the latter model’s attention mechanism, the software module it uses to interpret prompts.
Google LLC has launched DiffusionGemma, a revolutionary large language model utilizing a novel machine learning strategy known as text diffusion. This innovative model can generate text at a staggering rate of over 1,000 tokens per second on Nvidia's powerful H100 GPU, making it four times faster than traditional large language models (LLMs).123
In addition to its impressive speed, DiffusionGemma comprises 26 billion parameters, albeit only 3.8 billion are activated to formulate responses. This selective activation not only enhances performance but also significantly reduces memory usage, optimizing system resources.5
The model is built upon Google's earlier release, Gemma 4 26B A4B, which was introduced in April. In enhancing DiffusionGemma, the search giant has replaced the traditional attention mechanism of previous models, thus allowing for more efficient prompt interpretation.7
With such advancements, DiffusionGemma stands as a significant step in the evolution of LLM technology, marking a new era of rapid and efficient text generation.
“Google LLC has released DiffusionGemma, a text diffusion model that operates significantly faster than traditional language models. It can produce more than 1,000 tokens per second on Nvidia's H100 GPU.”
