Safety on Paper: Why AI Resignations and Evaluation Laws Are Not Keeping Pace with Capability
Safety on Paper: Why AI Resignations and Evaluation Laws Are Not Keeping Pace with Capability

Jacob Coxon resigned from Anthropic on 9 September 2026, all within the same week California Governor Newsom signed two state bills governing how outside organisations assess AI systems for safety. Coxon’s 100-million-view post on X accused both OpenAI and Anthropic of racing toward self-improving superintelligence without acting responsibly. CNN reported that a more senior Anthropic employee, Evan Hubinger, backed the core warning while estimating a personal chance below 10 per cent over the next decade that AI could kill all humans. The Wall Street Journal framed the departure as a sign of rising safety unease inside top AI firms. Newsom’s bills, backed by Anthropic in August and endorsed by OpenAI on the day of signing, are the institutional response to exactly this kind of concern. Two clocks are running, and safety is falling behind.
The Insider Break
Coxon’s specialty is training new models by handling massive data pipelines, placing him close to the operational core. His posts described researchers who build AI systems “earnestly believe it could kill everyone by the end of the decade.” That framing is a claim, not an established fact, and the risk numbers attributed to Hubinger deserve the same caution: a personal estimate below 10 per cent, weighty because it came from inside the firm building these systems. What matters is not whether anyone takes existential-risk percentages at face value, but that a practitioner walked out and said publicly what insiders reportedly say privately.
The Governance Response Labs Can Endorse
If the resignation is the admission, California’s evaluation bills are the institutional response. Politico reported that Newsom signed two bills on 9 September 2026 governing how outside organisations assess AI systems for safety, with Anthropic backing the measures in August and OpenAI endorsing them on the day of signing. OpenAI’s endorsement came after Coxon’s posts circulated widely and prompted federal lawmakers to renew calls for legislation. Evaluation law creates independent checks on systems whose developers have strong commercial incentives to self-certify. But evaluation frameworks backed by the firms being evaluated are necessary, yet not sufficient. A governance regime shaped by the regulated is better than nothing, but it should not be mistaken for independence.
The Epistemic Hole Under the Ledger
An arXiv paper dated 9 September 2026 introduces the silent revision rate, measuring how often frontier AI developers fail to disclose material changes to their published safety frameworks. The European Union and California already treat these frameworks as accountability instruments, yet neither jurisdiction requires developers to make clear what specifically changed. Safety frameworks are supposed to be the public ledger of safety commitments. If the ledger itself can be quietly rewritten, then evaluation laws and campus policies operate on a foundation that may shift beneath them. The silent revision concept matters not because it is a settled industry statistic, but because it names a structural vulnerability that existing regulation has not addressed.
The Capability Clock
OpenAI released GPT-6 Astra, which Campus Technology reported as the first broadly deployed model to reach the Critical cybersecurity capability threshold under its own Preparedness Framework. OpenAI says Astra can find previously unknown vulnerabilities and exploit them across multiple well-protected systems without human direction. The Preparedness Framework is both a yardstick and a marketing surface: it allows the company to label its own model as meeting a threshold it defined. The label is meaningful, but not independent verification. Meanwhile, Google Threat Intelligence Group reported that agentic AI workflows compressed credential harvesting to under six hours during the second quarter of 2026. Adversaries are already compressing attack timelines, and defenders are operating inside the same accountability gap the silent-revision paper describes.
The Education Turn
While labs argue about frameworks, a large university is rationing advanced model features without waiting for legislation. The Indiana Daily Student reported that Indiana University IT Services emailed students on 1 September 2026 that access to advanced ChatGPT Edu features has been tightened: 1,000 premium credits per month, down from 5,000 per week, with unused credits expiring. What is clear is that a university is managing exposure through rationing, a pragmatic response that operates entirely outside the legislative and corporate framework processes dominating the news cycle. The same week that brought a 100-million-view safety resignation, California evaluation laws, and a Critical cyber label, a campus simply turned down the dial.
What This Means Going Forward
The week’s evidence does not support a neat conclusion. The public safety story is being rewritten in real time, through resignations, evaluation laws, and silent framework edits, while capability thresholds and operational attacks expand. Accountability instruments are necessary but not sufficient: evaluation laws shaped by the firms being evaluated, safety frameworks that can change without disclosure, and a capability clock that outruns the governance cycle all point to the same structural problem. The question for the next cycle is not whether labs will publish safety frameworks, but whether anyone outside those labs can verify what they say at the moment they say it.