AI safety needs binding gates, not a slower race

By Zak and the True Work Office team | Published: 21 September 2026 | Category: blog | 3 min read
Key points
  • Stuart Russell contends that AI safety requirements must be non-negotiable prerequisites before further capabilities progress, not merely a slower pace of development.
  • Dario Amodei's "pacing the frontier" proposal includes third-party evaluators, common safety standards among frontier firms, and a broader compact involving authoritarian countries.
  • Russell draws an analogy with aviation, where regulators certify airworthiness before any new aircraft enters service, arguing the same principle should apply to frontier AI.
  • Russell argues the acceptable risk of irreversible loss of human control through recursive self-improvement should be orders of magnitude lower than what current AI chief executives consider tolerable.

Stuart Russell says frontier AI companies need enforceable safety thresholds, not a slower race. Writing in response to Anthropic CEO Dario Amodei’s “pacing the frontier” proposal, the UC Berkeley professor of computer science argues that safety requirements should function as non-negotiable prerequisites before further capabilities work proceeds. Amodei borrows the metaphor from Formula 1, which frames the problem as one of speed. Russell says that framing is wrong.

Amodei’s plan, which has drawn support from leaders at OpenAI, xAI, Google DeepMind and Microsoft, has three components: third-party AI system evaluators embedded within companies, common safety standards among frontier firms in democratic nations with government regulation where necessary, and a broader compact that would include authoritarian countries. It arrived in the wake of safety researcher Jacob Coxon’s departure from Anthropic and an incident involving OpenAI and Hugging Face, both of which raised questions about how seriously frontier firms take internal safety commitments.

Russell reaches for aviation to make his point. Aircraft manufacturers cannot simply schedule new models and hope that flight testing yields acceptable results. Regulators certify airworthiness before any new design enters service. The analogy is direct: safety verification precedes deployment, and no amount of engineering ambition exempts a manufacturer from that sequence. Applied to AI, it suggests that safety research must constitute a binding gate that capability research cannot pass without meeting defined standards, rather than simply keeping pace alongside it.

What makes this more than an academic objection is the gap between what current AI chief executives consider tolerable risk and what Russell considers proportionate to the stakes. He argues that recursive self-improvement leading to superintelligent AI poses risks of irreversible loss of human control, and that the acceptable probability of such an outcome should be orders of magnitude lower than the thresholds frontier companies currently accept. That is a direct challenge to the industry’s own risk calculus, and it lands at a moment when several of the largest AI firms are simultaneously positioning themselves as safety leaders while continuing to push capability boundaries.

The broader question is whether voluntary commitments like Amodei’s proposal can produce the accountability that the stakes demand, or whether enforceable regulatory red lines are the only mechanism that will hold. In education, where AI tools are already reshaping assessment and academic integrity frameworks, the pattern is familiar: industry self-governance tends to set the pace, and regulators arrive after the practice is entrenched. Russell’s argument reframes that dynamic. If safety must be proven before deployment, the burden of evidence shifts to the companies building these systems rather than to the publics affected by them.

Whether democratic governments have the institutional capacity to enforce such thresholds, particularly when the companies involved are among the largest and most politically connected in the world, remains an open and uncomfortable question.

Frequently asked questions

What is the 'pacing the frontier' proposal?

It is a framework proposed by Anthropic CEO Dario Amodei with three elements: third-party AI system evaluators embedded within companies, common safety standards among frontier firms in democratic nations, and a broader compact that would include authoritarian countries. It has drawn support from leaders at OpenAI, xAI, Google DeepMind and Microsoft.

Why does Russell reject the pacing metaphor?

He argues that framing safety as a matter of speed implies safety research merely needs to keep up with capability research. Instead, he contends safety should function as a binding prerequisite: capabilities work cannot proceed until safety thresholds are met, much as aviation regulators certify airworthiness before a new aircraft enters service.

How does this connect to AI in education?

The pattern of industry self-governance setting the pace before regulators catch up is already visible in education, where AI tools have reshaped assessment and academic integrity frameworks faster than institutional policy has adapted. Russell’s proposal would shift the burden of proof to the companies building these systems.

โ† Back to Blog