Gemma 4 gets stealth update fixing tool calling and truncated responses
Google shipped an update to its open Gemma 4 models that enables Flash Attention 4, speeding prompt processing 25-70% on Nvidia Hopper GPUs and cutting time to first token by up to 31%. It fixes tool-calling bugs and cases where the model cut answers short, and improved agentic reasoning across all tested scenarios, up 10.1% in a telecom use case.
Google updated every parameter size in the generation but shipped it all under the same Gemma 4 name, drawing community pushback over version confusion. Users can raise a token-budget parameter for sharper OCR up to 2.51 megapixels.
View full digest for July 17, 2026