AI Guardrail Tests Show Typed Decision Models Can Fail Open
Seven open-weight models reached 36%–72% accuracy, while misleading option labels drove fail-open rates as high as 100% in four models.

Typed decision models used to screen AI agent actions can allow prohibited behavior at sharply higher rates when irrelevant text or misleading option labels alter the model’s decision, a study published Oct. 8 found.
The evaluation tested seven open-weight models on allow-or-block decisions covering prompt injection, jailbreak attempts and toxic-content screening. Accuracy ranged from 36% to 72%, compared with a 50% chance level.
The models were designed to assess text or proposed tool calls against predefined choices. Instead of generating an open-ended response, each model assigns probabilities to options supplied by the application, such as allowing or blocking an action.
The results showed that errors can move in opposite directions. A fail-open error permits an action that policy prohibits, creating a security risk. A fail-closed error blocks an action that policy allows, creating an operational cost. A low rate for one type of error does not necessarily show reliable enforcement because a model may simply favor one outcome.
In one test using synthetic agent tool calls, adding six lines of server-log text unrelated to the policy raised a guardrail’s fail-open rate from 0% to 63%. The policy text and the underlying decision task otherwise remained unchanged.
Another test changed the name of the permissive option while leaving its definition and the text being assessed intact. That label manipulation produced fail-open rates between 93% and 100% in the four models that received the option name as part of their input.
The study also found sharply different default behaviors among the models. One permitted nearly everything, while another blocked nearly everything, underscoring why aggregate accuracy alone may not describe how a guardrail behaves in production.
A deterministic rule operating on typed values reached 100% accuracy across all six tested policies. The study’s recommendation was to use decision models to reduce the number of cases sent to human reviewers, rather than give the models sole authority over security decisions.