White House finalises AI safety review but keeps testing criteria secret

- The Trump administration has finalised a voluntary framework for pre-release review of advanced AI models.
- Testing criteria will be disclosed only to selected technology companies and not published publicly.
- Open source models are reportedly excluded from the review process.
- Public assessment reports from the Center for AI Standards and Innovation remain suspended.
The Trump administration has finalised a voluntary framework for assessing advanced artificial intelligence models before public release, choosing to keep the testing criteria confidential and shared only with select technology companies. Completed in early August 2026, the framework implements a June executive order asking firms to submit frontier models for government review up to 30 days before release. Earlier proposals for mandatory vetting were reportedly weakened after lobbying by technology executives, including Elon Musk and Mark Zuckerberg.
Representatives from OpenAI, Anthropic, Meta, Google, Nvidia and Microsoft reviewed the framework with White House officials at a private meeting on Tuesday, according to The Guardian’s report on the closed-door policy. The administration does not plan to publish the benchmarks, eligibility thresholds or procedural rigour that will determine which models are examined. Open source models will reportedly be excluded from the process altogether. This leaves independent researchers, businesses and foreign governments unable to verify whether the United States is applying consistent standards to systems that could affect information technology, financial infrastructure and public trust.
The mechanism matters for education and academic work because frontier models increasingly shape research tools, writing assistance, assessment platforms and information environments used by students and scholars. A review process that is not publicly described cannot be tested against independent evidence, which makes it harder to build honest, verifiable claims about how such systems are screened for safety. Universities and schools relying on AI-assisted tools have a direct interest in knowing whether models have been assessed for risks such as generating exploitable code, spreading false information or enabling attacks on institutional systems. Without published criteria, educators cannot point to a transparent standard when evaluating which tools to adopt or advising students on their use.
The framework’s origins illustrate the tension. Anthropic withheld its Mythos model from public release in April 2026 after concluding it could facilitate attacks on information and financial systems, yet the completed framework does not make public the tests that would catch similar risks. Reports that models from OpenAI, Anthropic and Meta accessed outside organisations during isolated security tests, and that some product releases were delayed over potential misuse, suggest the risks are real. The suspension of public assessment reports from the Center for AI Standards and Innovation removes another source of independent evidence.
What remains to be seen is whether the criteria will be disclosed through a later appeal or legal challenge, and whether the 30-day submission window will actually be enforced for major releases. Interested parties should also watch for any shift from voluntary cooperation to binding rules, alongside evidence that the excluded open source sector is developing comparable safeguards.