---
format: "aidr-story-markdown/v1"
id: "ffa877d2473d00cb2178255e98beefd79948a317ee9a893d66dc210551c16ea5"
canonical_url: "https://aidr.today/ffa877d2?lang=en"
title: "Google DeepMind Recirculation Lifts Gemma 3 GSM8k Accuracy 21% Without Retraining"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-23T16:40:33.000Z"
category: "Research"
topics: ["google","gemma","llm","reasoning","inference"]
source_urls: ["https://huggingnews.com/ai/google-deepmind-recirculation-lifts-gemma-3-gsm8k-accuracy-21percent-wit-a8dcdf18","https://x.com/omarsar0/status/2091548272968245466"]
summary: "The Recirculation method allows transformers to feed activations back through themselves during prefill, enabling models to act as dynamical systems for tracking belief states. When applied to the Gemma 3 family, this adaptive variant reduced perplexity by 23% and increased GSM8k accuracy by 21% using frozen weights and light hyperparameter tuning. Standard feedforward transformers are limited to as many internal state updates as they have layers, often forcing models to use chain-of-thought for state tracking in text. Recirculation implements recurrence at inference time to evolve the architecture without retraining, while shifting the additional serial computational work to the prefill stage to keep generation costs flat."
---

# Google DeepMind Recirculation Lifts Gemma 3 GSM8k Accuracy 21% Without Retraining

> [Open the canonical story](<https://aidr.today/ffa877d2?lang=en>)

**Published:** 2026-08-23T16:40:33.000Z
**Category:** Research
**Topics:** google, gemma, llm, reasoning, inference

## Summary

The Recirculation method allows transformers to feed activations back through themselves during prefill, enabling models to act as dynamical systems for tracking belief states\. When applied to the Gemma 3 family, this adaptive variant reduced perplexity by 23% and increased GSM8k accuracy by 21% using frozen weights and light hyperparameter tuning\. Standard feedforward transformers are limited to as many internal state updates as they have layers, often forcing models to use chain\-of\-thought for state tracking in text\. Recirculation implements recurrence at inference time to evolve the architecture without retraining, while shifting the additional serial computational work to the prefill stage to keep generation costs flat\.

## Sources

- [Story source](<https://huggingnews.com/ai/google-deepmind-recirculation-lifts-gemma-3-gsm8k-accuracy-21percent-wit-a8dcdf18>)
- [Supporting source](<https://x.com/omarsar0/status/2091548272968245466>)

