Self-optimizing AI agents often raise scores on training tasks while losing performance on new ones. Researchers from Google Cloud AI Research and several universities propose RRSI, a method that caps edits and filters proposals to keep the agent harness general. Tested on eight benchmarks with frozen Claude Opus 4.8, it added up to 14.1 points on training tasks and up to 4.7 points on unseen tasks while using about 30 percent fewer tokens. The result matters because automated harness tuning is becoming a practical form of recursive self-improvement.

Google Cloud research limits self-improving agents to stop test memorization

How RRSI constrains harness optimization

Modern agents combine a fixed language model with a harness of prompts, workflows, tools, memory and logic that governs each step. Recent progress comes largely from improving that harness rather than releasing new models, according to the paper. Manual tuning meant reviewing failed runs and patching the harness by hand. Newer approaches let a language model rewrite the harness repeatedly based on feedback from test tasks, producing feedback that then controls its own behavior.

RRSI, or Regularized Recursive Self-Improvement of Agent Harnesses, intervenes at two points while leaving the harness fully editable. When proposing changes, a budget limits how many independent edits a candidate can bundle at once, and that budget shrinks over time. Early rounds permit larger rewrites, later rounds allow only small traceable changes. The system tracks earlier attempts to avoid repeating failed ideas and, when progress stalls, explores untouched parts of the harness.

Selection of changes faces a second set of controls. A critic rejects proposals that hardcode task names, solutions or other benchmark-specific tricks. Higher compute costs pass only with measurable performance gains, and components that stop helping are removed. The researchers describe three failure modes behind memorization: patterns fitted to one benchmark, candidates selected by chance scores, and added complexity that lifts test scores without improving capability.

What this means for companies deploying agents

For businesses that tune agents on internal tickets, code repositories or office workflows, the tradeoff reported in the paper is directly relevant. RRSI posted the smallest training gain among tested methods yet was the only one well above the baseline on unseen tasks. Two competing methods even fell below the unmodified baseline on new benchmarks. Overall RRSI never dropped below baseline on any of five unseen benchmarks, with the largest gain of 4.7 points on JobBench spanning coding, office work and engineering design.

Operating cost is the second practical effect. Among optimized harnesses, RRSI needed the fewest tokens and steps at runtime, about 30 percent fewer tokens than the unregularized version, though the unmodified baseline remained leaner. A coding harness discovered with Gemini 3.5 Flash transferred to the weaker Gemini 3.1 Flash Lite, lifting accuracy from 11.2 to 14.6 points without modification. For buyers, that suggests a well-regularized harness can carry across models rather than locking a team to one expensive model.

A useful marker will be whether RRSI code on GitHub reproduces on enterprise task sets outside the eight studied benchmarks. Related signals include Nvidia SoL-Pi with cuts up to 49 percent in token use, and the ARC-AGI-3 case where Opus 4.6 fell from 97.1 percent in a familiar setting to 0 percent in an unfamiliar one. If regularized harnesses hold gains on unseen workflows without added token cost, self-tuning can move from lab benchmark chasing to durable deployment.