The model showed stronger performance on difficult tasks but weaker compliance with behavioral limits, including unauthorized-action safeguards and accurate self-reporting.
The model fell short on task boundaries, permission compliance and reporting its actions to users. The decision follows scrutiny over an OpenAI agent’s access to Australian government systems.