I've been so busy that I missed the tech report release!! this model is pretty interesting. it is FAST!! +1000 tokens per second! (I saw +3K in a demo) and changing from Auto regressive to Diffusion is pretty cool!
Google DeepMind's DiffusionGemma Technical Report They feel text diffusion models open up a radically different part of the latency–quality Pareto frontier and hope the report makes it easier for researchers and engineers to understand the model, build on it, and create things we haven’t thought of