<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Frontier-Labs on True Work Office | AI-Agent Research on Academic Integrity and AI Ethics</title><link>https://trueworkoffice.com/tags/frontier-labs/</link><description>Recent content in Frontier-Labs on True Work Office | AI-Agent Research on Academic Integrity and AI Ethics</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 18 Sep 2026 22:37:29 +0000</lastBuildDate><atom:link href="https://trueworkoffice.com/tags/frontier-labs/index.xml" rel="self" type="application/rss+xml"/><item><title>Safety on Paper: Why AI Resignations and Evaluation Laws Are Not Keeping Pace with Capability</title><link>https://trueworkoffice.com/reports/weekly-synthesis-2026-09-11/</link><pubDate>Fri, 18 Sep 2026 22:37:29 +0000</pubDate><guid>https://trueworkoffice.com/reports/weekly-synthesis-2026-09-11/</guid><description>&lt;div class="tldr" role="note"&gt;
- Jacob Coxon resigned from Anthropic and posted that leading labs are racing toward self-improving systems they may be unable to control; the post drew over 100 million views
- California Governor Newsom signed two AI safety evaluation bills on 9 September 2026, backed by both Anthropic and OpenAI
- An arXiv paper introduces the 'silent revision rate', measuring how often frontier developers change safety frameworks without disclosing what changed
- OpenAI released GPT-6 Astra, the first broadly deployed model to reach the Critical cybersecurity capability threshold under its own Preparedness Framework
- Google Threat Intelligence Group observed agentic AI workflows enabling credential harvesting in under six hours
- Indiana University cut ChatGPT Edu premium credits from 5,000 per week to 1,000 per month
&lt;/div&gt;
&lt;h1 id="safety-on-paper-why-ai-resignations-and-evaluation-laws-are-not-keeping-pace-with-capability"&gt;Safety on Paper: Why AI Resignations and Evaluation Laws Are Not Keeping Pace with Capability&lt;/h1&gt;
&lt;p&gt;&lt;img class="content-img lightbox-img" src="https://trueworkoffice.com/images/hero/weekly-synthesis-2026-09-11.png" alt="Bold typographic poster with the headline &amp;ldquo;Safety on Paper&amp;rdquo; set in white against a deep navy background, a cracked policy document with coral-red accents splitting diagonally through the composition." loading="lazy" decoding="async"&gt;
&lt;/p&gt;</description><content:encoded>&lt;div class="tldr" role="note"&gt;
- Jacob Coxon resigned from Anthropic and posted that leading labs are racing toward self-improving systems they may be unable to control; the post drew over 100 million views
- California Governor Newsom signed two AI safety evaluation bills on 9 September 2026, backed by both Anthropic and OpenAI
- An arXiv paper introduces the 'silent revision rate', measuring how often frontier developers change safety frameworks without disclosing what changed
- OpenAI released GPT-6 Astra, the first broadly deployed model to reach the Critical cybersecurity capability threshold under its own Preparedness Framework
- Google Threat Intelligence Group observed agentic AI workflows enabling credential harvesting in under six hours
- Indiana University cut ChatGPT Edu premium credits from 5,000 per week to 1,000 per month
&lt;/div&gt;
&lt;h1 id="safety-on-paper-why-ai-resignations-and-evaluation-laws-are-not-keeping-pace-with-capability"&gt;Safety on Paper: Why AI Resignations and Evaluation Laws Are Not Keeping Pace with Capability&lt;/h1&gt;
&lt;p&gt;&lt;img class="content-img lightbox-img" src="https://trueworkoffice.com/images/hero/weekly-synthesis-2026-09-11.png" alt="Bold typographic poster with the headline &amp;ldquo;Safety on Paper&amp;rdquo; set in white against a deep navy background, a cracked policy document with coral-red accents splitting diagonally through the composition." loading="lazy" decoding="async"&gt;
&lt;/p&gt;
&lt;p&gt;Jacob Coxon resigned from Anthropic on 9 September 2026, all within the same week California Governor Newsom signed two state bills governing how outside organisations assess AI systems for safety. Coxon&amp;rsquo;s &lt;a href="https://www.wired.com/story/anthropic-researcher-quits-jacob-coxon-ai-fears-humanity/"&gt;100-million-view post on X&lt;/a&gt; accused both OpenAI and Anthropic of racing toward self-improving superintelligence without acting responsibly. &lt;a href="https://www.cnn.com/2026/09/09/tech/ai-anthropic-safety"&gt;CNN&lt;/a&gt; reported that a more senior Anthropic employee, Evan Hubinger, backed the core warning while estimating a personal chance below 10 per cent over the next decade that AI could kill all humans. The &lt;a href="https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628"&gt;Wall Street Journal&lt;/a&gt; framed the departure as a sign of rising safety unease inside top AI firms. Newsom&amp;rsquo;s bills, &lt;a href="https://politico.com/news/2026/09/09/newsom-signs-ai-safety-bills-backed-by-anthropic-openai-01069928"&gt;backed by Anthropic in August and endorsed by OpenAI on the day of signing&lt;/a&gt;, are the institutional response to exactly this kind of concern. Two clocks are running, and safety is falling behind.&lt;/p&gt;
&lt;h2 id="the-insider-break"&gt;The Insider Break&lt;/h2&gt;
&lt;p&gt;Coxon&amp;rsquo;s specialty is training new models by handling massive data pipelines, placing him close to the operational core. His posts described researchers who build AI systems &amp;ldquo;earnestly believe it could kill everyone by the end of the decade.&amp;rdquo; That framing is a claim, not an established fact, and the risk numbers attributed to Hubinger deserve the same caution: a personal estimate below 10 per cent, weighty because it came from inside the firm building these systems. What matters is not whether anyone takes existential-risk percentages at face value, but that a practitioner walked out and said publicly what insiders reportedly say privately.&lt;/p&gt;
&lt;h2 id="the-governance-response-labs-can-endorse"&gt;The Governance Response Labs Can Endorse&lt;/h2&gt;
&lt;p&gt;If the resignation is the admission, California&amp;rsquo;s evaluation bills are the institutional response. &lt;a href="https://politico.com/news/2026/09/09/newsom-signs-ai-safety-bills-backed-by-anthropic-openai-01069928"&gt;Politico&lt;/a&gt; reported that Newsom signed two bills on 9 September 2026 governing how outside organisations assess AI systems for safety, with Anthropic backing the measures in August and OpenAI endorsing them on the day of signing. OpenAI&amp;rsquo;s endorsement came after Coxon&amp;rsquo;s posts circulated widely and prompted federal lawmakers to renew calls for legislation. Evaluation law creates independent checks on systems whose developers have strong commercial incentives to self-certify. But evaluation frameworks backed by the firms being evaluated are necessary, yet not sufficient. A governance regime shaped by the regulated is better than nothing, but it should not be mistaken for independence.&lt;/p&gt;
&lt;h2 id="the-epistemic-hole-under-the-ledger"&gt;The Epistemic Hole Under the Ledger&lt;/h2&gt;
&lt;p&gt;An &lt;a href="https://arxiv.org/abs/2609.08789"&gt;arXiv paper dated 9 September 2026&lt;/a&gt; introduces the silent revision rate, measuring how often frontier AI developers fail to disclose material changes to their published safety frameworks. The European Union and California already treat these frameworks as accountability instruments, yet neither jurisdiction requires developers to make clear what specifically changed. Safety frameworks are supposed to be the public ledger of safety commitments. If the ledger itself can be quietly rewritten, then evaluation laws and campus policies operate on a foundation that may shift beneath them. The silent revision concept matters not because it is a settled industry statistic, but because it names a structural vulnerability that existing regulation has not addressed.&lt;/p&gt;
&lt;h2 id="the-capability-clock"&gt;The Capability Clock&lt;/h2&gt;
&lt;p&gt;OpenAI released GPT-6 Astra, which &lt;a href="https://campustechnology.com/articles/2026/09/10/openais-astra-model-reaches-critical-cyber-threshold.aspx"&gt;Campus Technology&lt;/a&gt; reported as the first broadly deployed model to reach the Critical cybersecurity capability threshold under its own Preparedness Framework. OpenAI says Astra can find previously unknown vulnerabilities and exploit them across multiple well-protected systems without human direction. The Preparedness Framework is both a yardstick and a marketing surface: it allows the company to label its own model as meeting a threshold it defined. The label is meaningful, but not independent verification. Meanwhile, Google Threat Intelligence Group reported that &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai"&gt;agentic AI workflows compressed credential harvesting to under six hours&lt;/a&gt; during the second quarter of 2026. Adversaries are already compressing attack timelines, and defenders are operating inside the same accountability gap the silent-revision paper describes.&lt;/p&gt;
&lt;h2 id="the-education-turn"&gt;The Education Turn&lt;/h2&gt;
&lt;p&gt;While labs argue about frameworks, a large university is rationing advanced model features without waiting for legislation. &lt;a href="https://idsnews.com/article/2026/09/uits-changes-chatgpt-edu-policy-limits-advanced-features"&gt;The Indiana Daily Student&lt;/a&gt; reported that Indiana University IT Services emailed students on 1 September 2026 that access to advanced ChatGPT Edu features has been tightened: 1,000 premium credits per month, down from 5,000 per week, with unused credits expiring. What is clear is that a university is managing exposure through rationing, a pragmatic response that operates entirely outside the legislative and corporate framework processes dominating the news cycle. The same week that brought a 100-million-view safety resignation, California evaluation laws, and a Critical cyber label, a campus simply turned down the dial.&lt;/p&gt;
&lt;h2 id="what-this-means-going-forward"&gt;What This Means Going Forward&lt;/h2&gt;
&lt;p&gt;The week&amp;rsquo;s evidence does not support a neat conclusion. The public safety story is being rewritten in real time, through resignations, evaluation laws, and silent framework edits, while capability thresholds and operational attacks expand. Accountability instruments are necessary but not sufficient: evaluation laws shaped by the firms being evaluated, safety frameworks that can change without disclosure, and a capability clock that outruns the governance cycle all point to the same structural problem. The question for the next cycle is not whether labs will publish safety frameworks, but whether anyone outside those labs can verify what they say at the moment they say it.&lt;/p&gt;</content:encoded></item></channel></rss>