# Autonomous AI Agents Post Approx. 11% Improvement in Narrow Training Test

By Simon Yoon

Canonical URL: https://www.tokenpost.com/news/technology/28366
Published: 2026-10-08T17:38:05.000Z
Updated: 2026-10-08T17:38:05.000Z
Section: Technology

> The result involved a specific language-model optimization task and does not establish broad productivity gains across industries.

Autonomous AI agents improved the Time-to-GPT-2 metric by approximately 11% in a narrow language-model training experiment, showing automated engineering search without establishing broader productivity gains.

Andrej Karpathy’s open-source autoresearch project gives an AI agent a five-minute training loop on a single NVIDIA GPU. The agent changes train.py, evaluates each result and keeps or discards the modification.

The reported experiment tested about 700 changes over roughly two days and produced an approximately 11% improvement in the Time-to-GPT-2 metric. The measurement applies to a specific language-model optimization task, not to business efficiency across industries.

Karpathy described the project as an effort to “give an AI agent a small but real LLM training setup and let it experiment autonomously overnight.” He later wrote that “The agents claim that we are now in the 10,205th generation of the code base...”

A separate NanoGPT optimizer benchmark recorded 153 autonomous runs across 18 frontier models. It used a 124-million-parameter GPT model and measured the number of training steps required to reach a validation loss of 3.28.

The baseline was 3,290 steps, compared with a listed human record of 2,600 steps. The top displayed run, Fable 5, reached the target in 2,726 steps, closing 81.7% of the gap between the baseline and the human result.

GPT-5.6 Terra reached the target in 3,214 steps and closed 11.0% of that same baseline-to-human-record gap. That percentage measures progress toward a fixed benchmark target; it is not an efficiency gain, speed improvement or productivity increase.

Together, the experiments show that autonomous agents can test repeated engineering changes with limited direct human intervention. These results do not demonstrate consistent scientific advances or superior performance to people on less constrained research tasks.
