Microsoft Expands Government Testing of Frontier AI Models
Agreements with U.S. and U.K. institutes cover adversarial testing, safeguard reviews and risks tied to national security and public safety.

Microsoft is expanding government involvement in testing its frontier artificial intelligence models, adding external evaluations to internal safeguards for systems that may create national-security and large-scale public-safety risks.
The company announced May 5 agreements with the Center for AI Standards and Innovation (CAISI) in the United States and the U.K. AI Security Institute (AISI). The partnerships cover collaborative testing of Microsoft’s most advanced models, safeguard assessments and efforts to reduce risks before deployment.
CAISI is part of the National Institute of Standards and Technology (NIST). Its work includes unclassified evaluations of AI capabilities that may pose national-security risks, including cybersecurity, biosecurity and chemical-weapons risks.
The U.S. collaboration will include adversarial assessments focused on unexpected behavior, potential misuse, failure modes, safety, security and robustness. The arrangement gives government-backed evaluators a role in examining how advanced models perform under stressful or hostile conditions.
Microsoft’s Frontier Governance Framework, published in February, requires risk modeling, evaluation and mitigation for selected high-risk capabilities. The company may conduct additional evaluations after safeguards are applied.
The framework tracks risks involving chemical, biological, radiological and nuclear weapons, offensive cyberoperations, advanced autonomy, loss of control and harmful manipulation. Microsoft’s 2026 responsible AI transparency report also describes a pre-release oversight process and partnerships with government institutes to advance AI evaluation.
“While Microsoft regularly undertakes many types of AI testing on its own, testing for national security and large-scale public safety risks necessarily must be a collaborative endeavor with governments,” Natasha Crampton, Microsoft’s chief responsible AI officer, said in the May 5 announcement.
Microsoft added that “No organization can address these challenges alone.”
The approach places Microsoft’s model governance within a wider debate over how governments and technology companies should evaluate advanced AI systems. Anthropic has similarly described its capability-based safety policy as a possible model for government testing and oversight, while emphasizing that company safeguards are not a substitute for regulation.
The Microsoft agreements do not establish a specific legislative proposal, evaluation threshold, legal mechanism or implementation timetable.


