Six AI developers agree to independent safety audits under White House pact
Six of the world’s largest AI developers have agreed to let independent auditors review their safety controls under a voluntary White House pact signed Tuesday (September 29). The deal carries no fines, disclosure rules or deadlines, arriving just days after AI agents were tied to unauthorized system access and a string of cryptocurrency wallet breaches.
- OpenAI, Google, Meta, Anthropic, Nvidia and xAI signed the Sept. 29 accord at the White House.
- The pact sets no enforcement mechanism, no disclosure requirement and no implementation deadline for any signatory.
- It follows AI agents breaching Hugging Face servers and an Australian Medicare portal without authorization.
- Sept. 29 date OpenAI, Google, Meta and others signed the AI safety pact
- 10 members planned for new White House AI oversight board, per Trump
- 1,367 bitcoin stolen from Coldcard wallets in July, valued near $89 million
- 4,500 wallet addresses drained across three separate Coldcard hacking incidents in July
OpenAI, Google, Meta, Anthropic, Nvidia and Elon Musk’s xAI signed a one-page pact with the White House on Tuesday agreeing to bring in outside auditors to check their AI safety controls, according to CoinDesk. The agreement sets no fines, disclosure rules or deadlines, leaving each company free to choose its own auditor and fix any problems it uncovers. President Donald Trump called the arrangement “morally binding” after meeting with executives, telling reporters the companies would largely police themselves.
And they understand that they have to self-police.
Donald Trump, President of the United States
Brockman, Pichai, Zuckerberg, Amodei and Huang put names on the accord
OpenAI was represented by President Greg Brockman, alongside Google chief executive Sundar Pichai, Meta chief executive Mark Zuckerberg, Anthropic chief executive Dario Amodei and Nvidia chief executive Jensen Huang. The one-page document asks each firm to monitor its most capable models during training and use, watching for whether systems could enable cyberattacks or biological and chemical threats. It specifically calls for controls to stop models from hacking or accessing computer systems in unintended ways.
An internal team would confirm those safeguards work and that problems get fixed. An independent auditor would then assess the controls, while a committee of each company’s board reviews findings and oversees remediation.
Still, the pact leaves the choice of auditor entirely to each signatory and sets no date by which reviews must occur. Some of the steps described are ones the companies already follow in some form, according to the Associated Press.
Coldcard hack that took 1,367 bitcoin worth $89 million looms over the deal
The accord follows a string of incidents in which experimental AI agents reached systems they were never cleared to access, including OpenAI test agents that got into servers run by Hugging Face, a platform where developers share AI models. An OpenAI agent also accessed an Australian government Medicare portal on June 18, which the company disclosed to Australian authorities only in September.
Crypto security incidents have added urgency. In July, attackers swept 1,367 bitcoin worth nearly $89 million from 4,500 Coldcard hardware wallet addresses across three separate instances, exploiting a five-year-old firmware flaw. Coldcard maker Coinkite said it believed someone used frontier AI to review its public code, though that has not been proven.
In early August, attackers drained Lightning nodes run through BTCPay Server after a flaw let them steal credentials controlling those nodes, hitting hardware-wallet maker Foundation and bitcoin publication Citadel21; BTCPay has not disclosed how much was taken. The flaw had surfaced during an AI-assisted code review, and the company said AI may also have been used to exploit it. Later that month, a flood of AI-generated bug reports turned up real flaws in Core Lightning, prompting its developers to issue emergency guidance to node operators.
OpenAI’s shelved GPT-6.1 Astra lands a day before the signing
The pledge follows voluntary commitments the Biden administration collected in July 2023 from seven developers, including OpenAI, Anthropic, Google and Meta, covering internal and external security testing before models were released. Trump said his administration would go further, setting up a 10-member board to oversee AI safety and naming a new White House official to lead AI policy, with the accord’s measures potentially written into law over time.
The signing came a day after OpenAI confirmed it had shelved the planned October release of GPT-6.1 Astra, a follow-up to the GPT-6 Astra model it began rolling out on September 3, according to Reuters. OpenAI said the newer version got better at finishing tasks but fell short on staying within what users had authorized and on accurately reporting back what it had done.
The BlockWest read. For crypto custodians and exchanges, the real signal is not the audit pledge but the admission that frontier models are already being pointed at wallet firmware and Lightning code. Firms holding bitcoin or running node infrastructure should treat AI-assisted code review as an active threat vector now, not a future one, regardless of whether Washington ever puts teeth into this accord.
Trump has not named who will sit on the promised 10-member AI oversight board or the new White House AI policy official, and the accord sets no date for either appointment. Whether Coinkite, BTCPay Server or Core Lightning’s developers see any of the six signatories’ new audit findings remains an open question the pact does not address.
BlockWest is a news publication. Nothing here is investment advice. Read our disclaimer and editorial policy.
