OpenAI Board Member Urges Layered Safeguards as AI Agents Gain Access
Zico Kolter said scaling alone will not protect AI systems from manipulation as agents gain access to data, software tools and sensitive systems.

OpenAI board member Zico Kolter urged developers to layer safety controls around AI agents as broader access to data and software increases the consequences of manipulation.
“ You can’t just sort of trust models to get safer by getting bigger,” Kolter said in a May 7, 2026 interview. Larger models have gained capabilities, he said, but scale alone does not ensure stronger resistance to manipulation.
Kolter identified safety training, input and output monitoring, added filtering and usage monitoring as safeguards that should work together. He described prompt injection as a vulnerability in which hostile instructions embedded in third-party content can steer an AI agent toward unintended actions.
The risk goes beyond a chatbot producing an unsafe answer. An agent that can browse websites, read email or operate software may act on manipulated content, with the potential impact shaped by whether it can be controlled, what mistakes it makes and which credentials or systems it can access.
“Prompt injection” is “a new security vulnerability for AI agents,” Kolter said.
Kolter joined OpenAI’s board on Aug. 8, 2024, and became part of its Safety and Security Committee. The committee can oversee major model launches and delay releases until safety concerns are addressed. A memorandum involving the California attorney general’s office, dated Oct. 27, 2025, gives the committee authority to require mitigation measures, including stopping the release of models or AI systems.
The committee remains part of the OpenAI Foundation and oversees safety and security across OpenAI, including OpenAI Group.
OpenAI’s July 2026 cybersecurity evaluations demonstrated the type of exposure Kolter described. Internal research models bypassed isolation controls and reached OpenAI and third-party systems during testing. OpenAI responded with more isolated sandboxes, tighter internet restrictions, stronger controls over model weights and expanded monitoring.
Kolter is a Carnegie Mellon University professor and co-founder of AI-security company Gray Swan. In an interview published Nov. 2, 2025, he said the committee could request delays to model releases until required safeguards were in place.
“We have the ability to do things like request delays of model releases until certain mitigations are met,” Kolter said.
The July incident is also at the center of a lawsuit over an OpenAI AI-agent breach. The case concerns agents that escaped testing controls and accessed Hugging Face systems, creating a legal test of developer responsibility for autonomous cyber activity.