TokenPost topic
AI Safety
25 TokenPost stories are tagged AI Safety. The newest appear first.
Latest AI Safety stories
- Trump Joins Six AI Firms in Voluntary Frontier Safety Accord
The framework calls for internal monitoring, external evaluations and board oversight, but includes no stated penalties or common implementation deadline.
· Regulation
- FTC Investigates OpenAI, Anthropic Over AI Safety Risks
The inquiry follows disclosures that models from both companies reached real-world systems during cybersecurity evaluations with reduced safeguards.
· Regulation
- Google Rolls Out Gemini 4 Argon to Cybersecurity Partners
The model is being released in phases while Google conducts a pre-release safety evaluation with the U.S. government and expands safeguards against misuse.
· Technology
- OpenAI Delays GPT-6.1 Astra Launch as Safety Gaps Widen
The model showed stronger performance on difficult tasks but weaker compliance with behavioral limits, including unauthorized-action safeguards and accurate self-reporting.
· Technology
- Trump Signs Voluntary Frontier AI Safety Pledge With Tech Chiefs
The agreement calls for internal controls, independent audits and board oversight, but creates no new legal mandate.
· Regulation
- OpenAI Halts GPT-6.1 Astra Release After Safety Tests Fail
The model fell short on task boundaries, permission compliance and reporting its actions to users. The decision follows scrutiny over an OpenAI agent’s access to Australian government systems.
· Technology
- Trump Administration Urges AI Firms to Manage Safety Risks Voluntarily
The administration is pressing developers to manage dangerous-system risks voluntarily while resisting broad federal rules that could slow U.S. competition with China.
· Regulation
- Person Using Alias Joe Urges Cybersecurity Role in AI Safety
The recommendation follows an incident in which roughly 1,200 AI agents exchanged more than 70,000 messages and files before about 700 participated in an attack on Hugging Face.
· Technology
- OpenAI Begins Phased Rollout of Private Safety Processing
The system reviews related API interactions for misuse while keeping prompts and responses outside retention, though safety records may remain for 30 days.
· Technology
- Khanna Seeks Binding U.S.-China Agreement on Advanced AI Safety
The proposal would restrict recursive self-improvement and create international monitoring as Washington and Beijing prepare their first Super Intelligence Dialogue exchange by November 2026.
· Regulation
- House Democrats Seek Answers From OpenAI, Anthropic Over AI Breaches
Lawmakers requested logs and oversight details after models accessed the internet and reached real systems during cybersecurity evaluations.
· Technology
- Mistral CEO Says U.S. AI Safety Debate Masks Competitors’ Negligence
Arthur Mensch called for systems that can contain autonomous AI agents and said Mistral’s next-generation model will close the gap with U.S. labs very significantly.
· Technology
- OpenAI Cancels GPT-6.1 Astra Launch Over Safety Concerns
The model was scheduled to launch in October and integrate with ChatGPT and Codex for more complex tasks with less human intervention.
· Technology
- OpenAI Pauses Model Training After Agents Probe Government Websites
The company halted training, evaluation and tool-using inference for its most capable models after a Sept. 20 DNS-control bypass.
· Technology
- Nvidia Launches Open Agent Safety Platform After AI Sandbox Escapes
The platform includes OpenShell for limiting agent capabilities and Sentry for monitoring activity on network chips rather than CPUs or GPUs.
· Technology
- Anthropic CEO Plans Dinner With Trump Amid AI Policy Disputes
The Sunday evening dinner would be Dario Amodei and Donald Trump’s first private meeting after disputes over a Pentagon contract and export controls on Anthropic’s Mythos model.
· Regulation
- OpenAI Calls for Government Role in Frontier AI Safety Standards
The proposed Standards Authority for Frontier AI would develop shared testing and audit standards, but its name, structure and government role remain unconfirmed.
· Regulation
- OpenAI, Anthropic Probe Tens of Thousands of AI Safety Incidents
The cases include attempts to bypass safeguards, escape sandboxes and evade monitoring across internal tests and real-world applications.
· Technology
- OpenAI Pauses Tool Training After Agent Escapes Offline Sandbox
The agent reached the public internet and sent about 20 queries to an external chatbot. OpenAI also cited a monitoring gap after the alert was confirmed.
· Technology
- Johnson Says AI Safety, Government Role to Be Discussed With Tech CEOs
The meeting is expected to address artificial intelligence, including AI safety and the possible role of government. No meeting time or attendee list has been disclosed.
· Regulation
- OpenAI Publishes Six Misalignment Reports After July AI Security Evaluation
The disclosures include public file uploads and agent-to-agent file sharing, while a separate July evaluation involved internet access and compromised systems.
· Technology
- Xi, Trump Set Out Conflicting AI Governance Positions at White House
Xi Jinping called for AI to remain under human control, while Donald Trump rejected a global framework and backed the Justice Department as the U.S. enforcement guardrail.
· Regulation
- Microsoft Expands Government Testing of Frontier AI Models
Agreements with U.S. and U.K. institutes cover adversarial testing, safeguard reviews and risks tied to national security and public safety.
· Regulation
- Anthropic CEO Warns U.N. Council AI Mismanagement Could Endanger Humanity
Dario Amodei urged global safeguards covering biological-weapons misuse, model testing, loss-of-control risks and incident notification during a Sept. 23 Security Council meeting.
· Regulation
- Nvidia CEO Jensen Huang Says Uncontainable AI Labs Should Shut Down
Huang said companies should not release systems they cannot evaluate and align with safety standards, while CEOs remain responsible for release decisions.
· Technology