Enterprise Willingness to Let AI Agents Make Production Changes Falls to 56%
The August survey found 42% of organizations deploying autonomous agents expected to retain human review, while evaluation-platform adoption continued to rise.

The share of respondents at organizations deploying autonomous AI agents that allowed or were preparing to allow production changes without human approval fell to 56% in August from 75% in July, signaling continued caution around operational control.
The August survey included 140 respondents, 53% of whom were final AI purchasing decision-makers. Among those decision-makers at organizations deploying autonomous agents, the comparable share fell to 61% from 88% in July.
The 56% figure combines organizations that already allowed selected production changes based solely on automated evaluations with those engineering toward that capability within 12 months. Among organizations deploying autonomous agents, 32% already permitted unreviewed changes for specific low-risk agents or changes, while 24% were building toward it.
Human oversight remained common. Forty-two percent of respondents at organizations deploying autonomous agents expected to retain human review, up from 20% in July.
Across all respondents, 27% already allowed some low-risk agents to deploy changes without human validation, and 20% were engineering toward that capability within 12 months. Another 35% ruled it out for the foreseeable future, 16% did not deploy autonomous agents and 2% were unsure.
Reliability concerns also persisted after pre-deployment testing. Among organizations conducting those evaluations, 61% reported at least one case in the previous 12 months in which an AI agent or large language model feature passed internal tests but later caused a customer-facing failure. Twenty-two percent reported more than one such incident. The August figure is not directly comparable with the previously reported July measure because the respondent calculations differed, and the increase was not statistically significant.
Monitoring approaches varied. Real-time automated checks on live outputs were the primary approach for 29% of respondents at organizations deploying autonomous agents, while 36% primarily used transaction trace logging.
Human-review workflows were the reliability or evaluation investment most likely to grow over the next year, selected by 30% of respondents. Production observability tools followed at 26%, and automated evaluation pipelines at 21%.
OpenAI’s native evaluation tools appeared in 59% of respondents’ technology stacks in August, up from 31% in July. Overall, 62% planned to adopt, add or replace an evaluation platform within 12 months, including 35% planning to do so within three months.
The August findings differ from earlier 2026 survey results, but the separate, self-selected July and August groups do not establish a market-wide shift in enterprise behavior. The results instead show organizations continuing to invest in evaluation and monitoring tools while 42% of organizations deploying autonomous agents expected to retain human review.