Google researchers find a way to keep self-improving AI agents from memorizing their tests
Google researchers find a way to keep self-improving AI agents from memorizing their tests AI agents that keep optimizing their own working environment quickly tend to overspecialize on their test tasks. A new method from Google Cloud AI Research and several universities aims to prevent that while also cutting compute costs.

- AI agents that keep optimizing their own working environment quickly tend to overspecialize on their test tasks.
- A new method from Google Cloud AI Research and several universities aims to prevent that while also cutting compute costs.
- Modern AI agents wrap a fixed language model in a so-called harness, a framework of prompts, workflows, tools, memory, and logic that controls what the model sees at each step.
AI agents that keep optimizing their own working environment quickly tend to overspecialize on their test tasks. A new method from Google Cloud AI Research and several universities aims to prevent that while also cutting compute costs. Modern AI agents wrap a fixed language model in a so-called harness, a framework of prompts, workflows, tools, memory, and logic that controls what the model sees at each step. The harness decides whether an agent reads the right file before changing it, whether it recovers from a mistake, and whether it delivers its results cleanly. According to a new research paper, much of the recent progress in agents comes from work on the harness, not from new models. Until recently, this was done by hand. People reviewed failed runs and patched the harness manually. Newer methods automate the loop by having a language model rewrite the harness itself, again and again, based on feedback from the test tasks. The researchers call this a practical form of recursive self-improvement. The system produces feedback that it uses to optimize the harness, which in turn controls the system's own behavior. Self-optimization leads agents to memorize their test tasks The paper shows that this self-optimization comes with a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorizing them. Its scores on the training tasks go up, while gains on new, unseen tasks shrink or disappear entirely. The researchers say this happens in several ways. The search memorizes patterns that only fit one particular benchmark, favors candidates that score well purely by chance, and piles on unnecessary complexity that raises the test score without making the agent any better. Other methods mostly improve on the training tasks, but little of that carries over to unseen benchmarks. RRSI raises scores there in all three domains. | Image: Google Shrinking edit budgets and a strict critic keep the harness general RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) works on both ends of the optimization loop while leaving the harness fully editable. When the system proposes new changes, a budget caps how many independent edits a candidate can bundle at once. That budget shrinks over time. Early rounds allow larger rewrites, while later rounds only permit small changes that can be clearly traced to a result. The system also keeps track of earlier attempts so it doesn't keep chasing the same failed ideas. When progress stalls, it deliberately experiments with parts of the harness it hasn't touched yet. RRSI reins in self-optimization at two points, when proposing new changes and when deciding which of them become a permanent part of the harness. | Image: Google When it comes to picking changes, a critic reviews every proposal and throws out any that hardcode task names, solutions, or other benchmark-specific tricks. Another rule only accepts higher compute costs if they come with a measurable performance gain. Components that no longer help get removed. Giving up training gains pays off on new tasks The researchers tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design. The underlying model, Claude Opus 4.8, stayed frozen throughout. The team compared RRSI with the unmodified baseline harness and four recent optimization methods. According to the paper, RRSI gains up to 14.1 points on the tasks it was trained on and up to 4.7 points on five benchmarks it never saw. It also uses about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any of the unseen benchmarks, which typically happens with a harness that has memorized its tasks. The RRSI harness also improves on every task outside the training set, with the biggest gain of 4.7 points on JobBench. | Image: Google Every method did well on the training tasks, but the results flipped on new ones. Two methods even ended up below the baseline harness. RRSI posted the smallest training gain of all the variants and was the only method that landed well above the baseline on unseen tasks. The guardrails are meant to produce exactly this tradeoff. Among the optimized harnesses, RRSI needs the fewest tokens and steps and performs best on new tasks, though the unmodified baseline harness is even leaner. | Image: Google Harnesses optimized on one model also help weaker ones A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications. The mechanisms the system found don't depend on the capability of the model used to discover them. The authors note that their study only covers harnesses built around frozen models and doesn't address cases where the model weights change. They conclude that self-improvement only makes AI agents reliably more capable when repeated feedback gets turned into lasting changes. The code is available on GitHub. Manually designed harnesses often don't generalize to new tasks, as tests on ARC-AGI-3 have shown. With a purpose-built harness, Opus 4.6 scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Nvidia recently presented SoL-Pi, a related method in which a research agent automatically rebuilds the harness of coding agents. It cuts token use by up to 49 percent without a noticeable drop in performance. Shortly before that, Google had agents "dream" about past search runs to improve their search strategy. That work also leaves the model itself unchanged. AI News Without the Hype – Curated by Humans Full access to every article on THE DECODER Join the comments and community discussions 6x/year: "AI Radar" — deep dives on the AI topics that matter most.
Sources
Related stories

Vulnerabilities in AI Agents Expose Trust Gap Flaw
Google and other organizations have acknowledged vulnerabilities in their AI agents, which exploit trust gaps in the Model Context Protocol.

All the AI agents that can live in your text messages
All the AI agents that can live in your text messages | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

Meta wants your next gadget to be Muse-infused
Meta wants your next gadget to be Muse-infused | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

Apple says it’s tightening macOS ‘Full Disk Access’ controls due to new risks from AI agents
Apple says it s tightening macOS Full Disk Access controls due to new risks from AI agents | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.