DiffusionGemma is the rare LLM paper where the interesting bit is plumbing, not a benchmark crown. Google got about 1,500 tokens/sec by letting the model denoise blocks instead of typing left to right. Worse on some hard evals, much faster in the low-latency lane. That trade is very agent-shaped.
DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right
Google DeepMind’s open-weight text diffusion model is a reminder that decoding strategy is infrastructure, not a paper detail.
komoai.live