# RRSI Raises AI Agent Scores While Keeping Models Fixed

By Simon Yoon

Canonical URL: https://www.tokenpost.com/news/technology/26124
Published: 2026-10-01T09:52:56.000Z
Updated: 2026-10-01T09:52:56.000Z
Section: Technology

> The framework revises prompts, tools, control flow and memory. Gains also appeared across six held-out splits.

Researchers from Google Cloud AI Research and academic collaborators introduced a framework that improves AI agent benchmark scores by changing the software around a model while keeping the model fixed.

Called Regularized Recursive Self-Improvement of Agent Harnesses, or RRSI, the framework repeatedly edits an agent’s “harness,” including its prompts, tool interfaces, control flow, memory and context management. The model itself remains unchanged during this process.

RRSI constrains how it proposes and selects changes. Its controls include gradually shrinking the edit budget and screening candidates for benchmark-specific logic, evaluation noise and inference cost.

With Claude Opus 4.8 held fixed as the policy model, Terminal-Bench 2.1 scores rose from 74.2 to 80.2. SWE-bench Verified scores increased from 82.0 to 83.8, while EngDesign scores climbed from 50.0 to 54.9.

In a separate run using Gemini 3.5 Flash, Terminal-Bench 2.1 scores increased from 64.6 to 78.7, and SWE-bench Verified scores rose from 76.8 to 79.0.

The framework is designed to limit overfitting, which can occur when a system is repeatedly tuned against a finite set of tasks. Improvements also appeared on all six held-out splits. The results remain tied to the benchmark suites and experimental setup; they do not establish that the method will generalize to other models or settings or improve a model’s underlying capabilities.

The project code is available under the Apache 2.0 license. The project is not an officially supported Google product.
