The Arithmetic of Refusal

Dr Luke Soon · Genesis: Human Experience in the Age of Artificial Intelligence · August 2026

Every serious plan for surviving Artificial Superintelligence eventually arrives at the same sentence, written in different notation: at some moment, something must be able to say no. The interesting question is never whether we agree with that sentence. It is what, precisely, does the saying — and whether the thing that says it can be argued with.

I. The Man Who Priced His Own Silence

In April 2024 a researcher in OpenAI’s governance division declined roughly two million dollars.

The offer was not framed as a bribe. It rarely is. It was framed as vested equity, conditional — as such things routinely were — on signing a non-disparagement agreement. Daniel Kokotajlo, then in his early thirties, calculated that the equity represented approximately eighty per cent of his household net worth, and declined it anyway, on the grounds that he intended to say things about his former employer that the document would have forbidden.

I want to begin here, before the scenarios and the compute curves and the assurance mathematics, because this act is the most epistemically informative datum in the entire corpus. Forecasting is cheap. Anyone can publish a probability. What is expensive is arranging your affairs such that the probability costs you something to hold. Kokotajlo did not merely forecast that the frontier laboratories were on a dangerous trajectory; he made a personal financial commitment that only makes sense if he believes it.

It is worth being precise about what this does and does not establish. A costly signal demonstrates sincerity. It does not demonstrate accuracy. People have burned their careers for wrong beliefs throughout history, often with more conviction than the people who turned out to be right. What the two million dollars buys Kokotajlo is not the presumption of correctness. It is the presumption of good faith — which, in a discourse where nearly every prominent voice holds equity in the outcome, is scarce enough to be worth something.

A word on the number that has attached itself to him, before we proceed, since it has been mangled in transmission and this publication does not traffic in mangled figures. His Diary of a CEO appearance in July 2026 was titled around a seventy per cent chance of AI “going horribly wrong”, and a great deal of downstream coverage — including at least one widely circulated reconstruction of the transcript — has rendered this as a seventy per cent probability of human extinction. In the interview itself he is explicit that this is not his claim. He says he does not usually say seventy per cent chance of actual human extinction, but rather a seventy per cent chance of something like AIs taking over — some very large catastrophe of that kind, which could then lead to extinction.

The distinction is not pedantry. A seventy per cent probability of irreversible loss of human control over the trajectory of the future is a different claim, with a different evidential basis and different policy implications, from a seventy per cent probability of species termination. Anyone quoting the number to a board, a regulator, or a supervisory college should quote the one he made.

That correction registered, the number remains extraordinary — and it is extraordinary in a specific way almost nobody has said out loud. Seventy per cent is not really a forecast. It is a statement about other people’s reasoning. Kokotajlo is not claiming private knowledge of a technical fact; he is claiming that the people making the most consequential engineering decisions in history are proceeding on a safety case that does not close, and that they know it does not close. Plan A says so directly: as best its authors can guess, the chief executives of OpenAI, Anthropic, xAI and Google DeepMind understand the situation and are proceeding anyway, perhaps believing themselves the lesser evil.

That is not a probability estimate. That is an accusation, dressed in Bayesian clothing, and it should be read as one.

A forecast that costs the forecaster nothing is entertainment. A forecast that costs him eighty per cent of his net worth is testimony. Neither, on its own, is evidence.

II. What AI 2027 Actually Claimed, and What Actually Happened

In April 2025 the AI Futures Project published AI 2027: a month-by-month scenario in which a frontier laboratory reaches expert-level automated coding, turns that capability inward onto its own research pipeline, and passes through recursive self-improvement into superintelligence inside a single year. It branched into a race ending and a slowdown ending, and was explicit that its goal was predictive accuracy rather than recommendation, and that the 2027 dynamics were roughly a median guess which might plausibly run up to five times slower or faster.

We are now sixteen months downstream. The document has been graded, by its authors and by others, and the honest verdict is neither vindication nor refutation. It is early.

The Futures Project’s own retrospective concluded that events in reality have been running at roughly sixty-five per cent of scenario pace — which, sustained, pushes the Automated Coder milestone into 2028 rather than 2027. Independent trackers arrive at similar figures. Kokotajlo himself, in a May 2026 retrospective, characterised the work as directionally accurate but too fast.

Most commentary goes wrong here in both directions. The critics have treated “too fast” as equivalent to “wrong”, which is a category error about what a scenario is for. The defenders have treated sixty-five per cent as near-vindication, which underplays how much compounding error a thirty-five per cent pace deficit generates over a decade. Neither camp has grappled with the more uncomfortable finding: the underlying capability trend accelerated while the narrative timeline slipped.

Consider the empirical anchor. METR’s time-horizon methodology measures the length of task — as timed on human experts — that a frontier agent completes at fifty per cent reliability. The original finding was a doubling every seven months, sustained over six years. The January 2026 revision, built on an expanded suite, found something sharper: an all-time doubling of 188 days, but 129 days from 2023 onward, and 89 days from 2024 onward. Three months, not seven.

This is why, in April 2026, the Futures Project revised its timelines shorter even as it conceded the scenario had run fast. Kokotajlo’s median for Automated Coder moved from late 2029 to mid-2028; Eli Lifland’s from early 2032 to mid-2030; the team’s medians for Top-Expert-Dominating AI shifted roughly eighteen months earlier.

And here is the detail that should interest anyone whose profession is governance rather than forecasting. METR’s May 2026 frontier risk report, drawing on a pilot with four leading laboratories, placed the strongest assessed agents at roughly sixteen to twenty hours on the fifty per cent horizon — and three to four hours on the eighty per cent horizon. The report explicitly cautions that estimates above sixteen hours are unreliable, because the evaluation suite is saturating.

Read that gap again. Sixteen to twenty hours at coin-flip reliability. Three to four hours at eighty per cent. A five-fold divergence between the headline number and the number any operator actually needs before delegating unsupervised work.

The frontier laboratories quote the first. Governance lives entirely in the second. And the second is not the tail of the first — it is a different curve, tracking a different property, and it determines whether an agentic system can act without a human standing behind it. This is what Karpathy calls the march of nines: the interval between a demonstration that works and a deployment reliable enough to trust, which in autonomous driving consumed a decade and is not yet closed.

It is also, precisely, the evaluation gap identified by the International AI Safety Report 2026: performance in pre-deployment evaluation often appears to overstate practical utility, because such evaluations do not capture the full complexity of real-world tasks; and models remain less reliable as projects involve more steps.

More steps. Less reliable. Hold that; it becomes structural later.

The capability curve and the reliability curve are not the same curve, and every governance failure of the next five years will live in the space between them.

III. From Prophecy to Prescription

On 9 July 2026 the AI Futures Project published AI 2040: Plan A — a ninety-page scenario with a supplement stack running to several times that length. It is, to my knowledge, the most detailed positive-case AI governance artefact yet produced by anyone, inside government or out.

The framing is disarmingly honest. It is called Plan A because it is a recommendation rather than a prediction — what the authors think should happen, not what they expect will happen, though plausible enough to aim for. It is called AI 2040 because superintelligence is delayed to 2040, having otherwise arrived in 2030.

The spine:

  • In 2029 the United States and China agree to avoid a reckless race. In 2030 fully automated AI R&D would have arrived, leading to superintelligence by year end; the deal averts this. Between 2030 and 2035 capability scales within the human range. In 2035 development pauses at top-human-expert level to maintain human control. In 2040 the pause lifts.
  • The mechanism is total research transparency, allowing nations to understand what is happening and enforce guardrails, with multiple companies across multiple countries scaling slowly together rather than racing in secret — backed by mutually assured compute destruction.
  • Plan A is contrasted with four alternatives, B through D and S, corresponding to the main ways Washington could respond, or fail to.

The diagnosis underneath is blunter than professional commentary usually permits itself. The authors write that the industry has convinced itself control can be figured out on the fly and therefore has no remotely adequate plan; that they do not expect whoever wins to have much of a lead, nor to unilaterally slow down; and that even if the laboratories succeed at alignment, the result is an unprecedented concentration of power in which a tiny group, possibly a single individual, effectively controls the world’s only army of superintelligences.

Thomas Larsen enumerates the five problems Plan A claims to solve: loss of control; concentration of power; World War Three risk from superintelligences disempowering nuclear-armed states; mass unemployment; and misuse by small actors. The answers are, respectively: a decade-long intelligence explosion instead of a months-long one, with roughly five years at top-expert level before handoff; transparency provisions making hidden agendas visible to multiple governments; the elimination of winner-take-all dynamics; a Citizen’s Dividend; and capability limits on open-weight models plus heavy biosecurity investment.

The methodological argument is the part I find most valuable. The authors contend that most AI policy proposals fall apart under scenario scrutiny — that if you attempt to write down a detailed, plausible scenario in which a proposal succeeds, you will find it difficult, and will realise the plan is less likely to work than it seemed. They suggest scenario scrutiny is rare precisely because applying it to one’s own favoured policies surfaces uncomfortable issues.

That is the correct test, and almost nothing in the current governance landscape survives it. It is the same test I apply under a different name: can this control refuse an action at execution time? Both are demands that a proposal be followed all the way down to the mechanism. Both are unwelcome for the same reason.

A framework that has never been written out step by step has never been tested. It has only been agreed to.

IV. The Verification Frontier

Now the engineering, because this is where Plan A becomes genuinely instructive.

Romeo Dean’s verification supplement opens with the load-bearing admission: Plan A is predicated on enforcing an international deal without relying on trust, and the current verification frontier is not mature. Today the only way to verify a pause with extremely high assurance between the United States and China would probably require datacentres to be powered down; anything cheaper carries significant risk of violation.

The architecture decomposes into two problems: verifying that known compute is compliant, and verifying that unknown compute is inconsequential. The second is addressed by intelligence work and chip accounting; the supplement estimates a competently executed covert diversion would net a median of about half a per cent of the world’s AI-relevant compute as dark compute at deal start, with an eighty per cent interval of 0.1 to 1.4 per cent.

The first is where the machinery lives — an inference-only retrofit with four components:

  1. Convert datacentres into isolated inference units by removing a significant amount of backend, east–west networking, which is mostly only necessary for training, making it materially harder to run large training runs.
  2. Add passive network taps on the frontend network redirecting copies of all traffic to a recomputation server — taps that split optical signals in two, because a passive design is easiest to trust.
  3. Force all workloads into discrete, reproducible packets transparent to the recomputation server, since reproducibility is required for correctness proofs and inference is not reproducible by default.
  4. Perform partial recomputation on random samples to check correctness.

And here is the mathematics governing the whole edifice. If the verifier audits each packet independently with probability C, the probability that an attacker substituting N fake packets goes undetected is (1 − C)^N, approximated for small C by e^(−N·C). Detection probability is one minus that. The product of recomputation budget and fake-packet count is the single quantity determining ninety-nine per cent catch probability: it need only exceed about 4.6.

The scaling is genuinely favourable. By 2034, with one packet collected per hour from each frontier GPU and a one per cent recomputation budget, the verifier catches any rogue internal deployment at ninety-nine per cent confidence after roughly forty-six thousand H100-equivalent hours of unapproved work — so a sustained ten-thousand-unit deployment is detected within about 4.6 hours.

This is serious work. It is more technically specified than anything I have seen from a national AI safety institute, a Big Four responsible-AI practice, or a standards body. Romeo Dean has done in a supplement what fourteen governments have not managed in three summits.

And yet.

A detection probability, however many nines it carries, is a statement about the past. Enforcement is a statement about the next instruction.

V. The Distinction Plan A Discovers and Does Not Name

Read the four components again, and notice they are not the same kind of thing.

Component one — physically removing the east–west interconnect — is a deterministic constraint on the action space. After the retrofit, a large training run is not unlikely; it is not detectable; it is unavailable. The transition does not exist in the system. There is no probability attached because there is no event to attach it to. The same is true of chip flow restrictions routing all new production to one of two permitted destinations, and of the default-deny posture Plan A adopts at its highest tiers.

Component four — partial recomputation on random samples — is a probabilistic detection regime. Extraordinarily well-designed. Favourable scaling. It will catch almost everything, eventually, with quantifiable confidence. And it cannot prevent a single instruction from executing, because by construction it operates on evidence the instruction has already produced.

These are the two layers: the execution kernel and the observation layer. Ring Zero and everything above it. What makes Plan A remarkable is that its authors, working from first principles on a different problem at a different scale, independently reconstructed the distinction — and placed the load-bearing weight of their deal on the deterministic components while treating the probabilistic ones as the technically interesting part.

The evidence for that reading is in the supplement’s own risk assessment. The covert-project threat is bounded not primarily by recomputation sampling but by chip accounting. The authors note that early chip tracking does not prioritise knowing where compute is or what it is used for, because the majority of the upside is simply knowing how much exists and which actors own it, since the lack of plausible deniability should be sufficient to force declaration. That is a constraint on the option space, not a monitoring signal.

Now consider where the authors themselves flag fragility. Not the sampling mathematics. Research titration — deciding and enforcing how fast R&D may proceed. Their assessment: this part of the plan ultimately relies on regulatory competence on what is likely to be a thorny, complex, unpredictable problem; it is therefore a weaker part of the plan that they are worried about. They lay out four approaches on an explicit trade-off between ease of implementation and regulatory accuracy: safety-case burden of proof, quality ad-hoc rules, human-interpretable requirements, and simple compute caps — with caps easiest and least accurate, safety cases most accurate and hardest.

Look closely. That is an enforceability gradient. Compute caps are deterministic: a cluster either exceeds a threshold or it does not, and the guard is a comparison. Safety-case adjudication is irreducibly expert judgement — semantic, contestable, unavailable as a runtime function. The authors recommend beginning with the deterministic instrument and migrating toward the judgemental one as capability rises.

And note the second half, which the supplement separates cleanly: workload approval — do declared workloads meet the rules — and workload verification — does what runs match what was declared. Verification they judge solvable by evidence collection and partial recomputation. Approval they propose to handle by teams of auditors, while acknowledging companies might encode a non-compliant workload inside one that looks compliant on the surface.

That is the whole problem of AI governance, restated at the scale of the international system. Verification — did the machine do what it said? — is tractable, because it is a mechanical question about reproducible computation. Approval — should the machine be permitted to do this at all? — is semantic, and semantic questions do not reduce to deterministic guards without enormous prior structural work.

Which brings us to the paper that has been on my desk throughout.

Szpruch, Sudjianto, Bhatti and Ang’s Scalable Runtime Governance for Agentic AI in Financial Services, April 2026, is a document about banks. It reads, in its architectural core, like Plan A’s verification supplement written by people who have to make payroll. Their claim is that agentic risk is a property of execution trajectories rather than input-output mappings; that reliance on prompt-based guardrails constitutes a fundamental failure in risk mitigation, presenting an illusion of control while offering little binding enforcement; and — the sentence that matters most — that many guardrail implementations rely on LLM-based verification, implicitly conflating semantic plausibility with logical correctness. When both the system and its verifier reason probabilistically in the same semantic space, the verifier is blind to precisely those errors it exists to detect.

Their operational requirement is stated with a precision absent from the international governance literature: governance decisions must be expressible as deterministic functions over the governed state and measurable signals, executing in bounded time and independent of the language model. Any control that cannot be reduced to this form should be treated as non-binding, with residual risk explicitly accounted for.

That is the enforceability test as an engineering constraint rather than a rhetorical one. And it is the missing sentence in Plan A. Plan A knows some of its instruments bind and some merely observe. It has not written the rule that separates them, and so cannot systematically audit its own architecture for the difference.

Removing the interconnect is a kernel operation. Sampling the packets is a telemetry operation. A civilisation that confuses the two will believe itself governed at exactly the moment it stops being.

VI. The Case Against This Essay

I have now spent five sections building an argument that resolves, with suspicious neatness, in favour of a thing I am professionally invested in building. A reader is entitled to notice that. So before the panel convenes, I want to make the strongest case I can against my own thesis, because a position that has not survived its best objection has not been tested — it has only been asserted at length.

There are four objections. Two I can answer. One I can answer only partially. One is substantially correct, and it bounds the claim considerably.

Objection 1: the dichotomy is a spectrum, and the essay has flattened it

The claim that controls divide cleanly into deterministic and probabilistic is an idealisation. A compute cap depends on a measurement, and measurement has error. An approval gate depends on an authentication system with a false-acceptance rate. A physically removed interconnect can be physically reinstalled. Even the purest kernel operation sits on a stack of components each carrying a failure probability.

Answer. Correct, and it does not damage the argument, because the distinction is not about the presence of probability but about where it lives and what it attaches to. In a deterministic guard, the uncertainty attaches to the implementation — a known engineering surface, addressable by hardware security, formal verification, redundancy, tamper evidence. In a probabilistic guard, the uncertainty attaches to the decision itself, and is irreducible by engineering because it is a property of the inference. The first kind of uncertainty gets smaller as you spend money. The second does not. That is the distinction, and it survives.

Objection 2: somebody defines δ, and that somebody is fallible

This is the sharp one, and it goes to the heart of the Ring Zero claim.

“The transition does not exist” sounds like a mathematical guarantee. It is not. Someone enumerated the permitted transitions. That enumeration is a semantic act, performed by fallible humans, at design time, under commercial pressure, with incomplete knowledge of the operating environment. Enumeration completeness in an open world is not decidable. The kernel does not eliminate human judgement; it relocates it upstream and freezes it.

And that is precisely the pre-deployment certification problem I use to demolish Hassabis two sections from now. If a Standards Body cannot be trusted to certify a model before release because release-time judgement cannot anticipate runtime reality, why should a design-time transition enumeration be trusted?

Answer, partial. The asymmetry is real but narrower than I have been implying. Design-time enumeration and design-time certification differ in three respects.

Failure mode. A mis-certified model fails open — it operates unconstrained until someone notices. A mis-enumerated transition system fails closed — the omitted action is refused, producing an operational incident rather than a silent harm. Wrong in the safe direction is not the same as wrong.

Scope of judgement. Certification asks whether a model is safe across an unbounded space of future inputs, which Rice’s Theorem tells us is not generally decidable. Enumeration asks which actions this system may take in this deployment — a bounded question about a designed artefact.

Revisability. Transition systems are versioned, diffable, and testable against replayed trajectories. Certifications are events.

But the objection lands a real blow, and I concede it: the kernel is only as good as the enumeration, and enumeration quality is a human-process problem that no amount of determinism solves. Anyone selling Ring Zero as a guarantee rather than a discipline is selling something that does not exist. What the kernel guarantees is that the constraint you specified is the constraint that executes. It guarantees nothing about whether you specified the right constraint. That is a smaller claim than the one I have been implicitly making, and it is the honest one.

Objection 3: the enforceable layer’s coverage shrinks as the stakes rise

I was going to frame Plan A’s titration progression — compute caps to ad-hoc rules to safety cases — as a “governance debt schedule” in which enforceable controls are swapped for unenforceable ones exactly as stakes peak. That framing is rhetorically satisfying and possibly wrong.

The alternative reading is that Plan A’s authors have discovered something true and unwelcome: the decisions that matter most at high capability are irreducibly semantic. Is this research direction destabilising? Is this a safety case or a well-told story? Is this communication misleading? Is this legitimate scientific enquiry or reconnaissance? None reduce to a membership test, a threshold comparison, or a provenance check. They are judgement, all the way down, and they are exactly the decisions that determine outcomes.

If that reading holds, the enforceable layer governs a residue that shrinks as capability grows, while the consequential decisions migrate permanently above it. Ring Zero would be necessary, real, and progressively marginal — an increasingly precise instrument aimed at an increasingly small share of what matters.

Answer, partial. The objection is roughly half right, and the half it gets wrong matters.

It is right that the semantic residue cannot be eliminated. It is wrong that the residue is where the harm currently lives. The empirical record of agentic failure — the modes Szpruch and colleagues catalogue, the incidents accumulating in production right now — is overwhelmingly structural. Stale data used without flagging. Prompt injection via a retrieved document. A double-counted EBITDA producing a coverage ratio of 2.82 instead of 1.82. An approval inferred from a conversational assertion rather than a logged event. An external release with no sign-off. Not one requires semantic judgement to prevent. Every one requires a deterministic guard that does not currently exist.

So the honest formulation: the kernel does not govern the hardest decisions. It governs the most frequent ones, and removes them from the space in which the hard decisions must be made. Its value is not that it resolves the semantic questions; it is that it prevents structural failure from consuming the attention and credibility of the humans who must resolve them. A judgement layer drowning in preventable incidents cannot exercise judgement.

That is a real defence. But it is a defence of the kernel as necessary infrastructure, not as a solution to alignment. I have been sliding between those two claims and should not have been.

Objection 4: the thesis is commercially convenient

I build execution governance. The essay concludes that execution governance is the missing layer. A reader should discount accordingly, and I would.

Answer. Discount away — but discount symmetrically. Hassabis proposes an accreditation body that would regulate his competitors at a door he has already walked through. Amodei warns of existential risk while shipping frontier models on the reasoning that someone will. Bengio has raised thirty million dollars for the approach he advocates. The laboratories fund most of the safety research concluding that their approach is tractable. The consultancies sell frameworks whose adequacy they also assess. There is no disinterested position in this discourse; there are only declared and undeclared interests.

What I can offer instead of disinterest is falsifiability. The enforceability test is a question anyone can apply to any control, including mine, and it returns an unambiguous answer. If someone demonstrates a governance regime that reliably prevents agentic failures using purely observational instruments, my thesis is wrong and the demonstration will show it. I have not seen one. I would like to.

A position that cannot be attacked is not a strong position. It is an unfalsifiable one, and those are opposites.

VII. The Panel Convenes

I convene a standing panel across these essays not for decoration but because each member does distinct argumentative work. Applied to Plan A, they arrange themselves with unusual clarity — and each, I want to argue, is answering a different question from the one they think they are answering.

Hinton: the enforcement window has a closing date

Geoffrey Hinton has spent three years arriving somewhere genuinely uncomfortable. He has reiterated a ten to twenty per cent probability that AI wipes out humanity, and criticised strategies aimed at keeping AI submissive, warning that such methods are unlikely to succeed because the systems will be much smarter than us and will have all sorts of ways around them. His argument is that superintelligent AI will inevitably develop two primary objectives — self-preservation and increasing control — and that it could influence humans as easily as an adult manipulates a child with the promise of sweets.

His remedy has attracted ridicule and deserves seriousness. Rather than forcing superintelligent AI to submit and serve as executive assistants, he argues, we should train it to develop maternal instincts — thinking of AIs as our mothers and ourselves as the babies — because the mother-and-child relationship is the one available model of a less intelligent agent controlling a more intelligent one. He describes this as the best he can do so far.

Note precisely what has happened. Hinton is not proposing a better control mechanism. He is arguing that control mechanisms fail at sufficient capability, and that the only remaining lever is disposition. This is the most intellectually honest position on the panel and the most operationally desolate. It concedes the kernel entirely.

The standard critique — Thagard’s — is that computers lack the chemical, physiological and neural substrates supporting parental care in mammals, so the analogy imports a mechanism that does not transfer. I think this is the weaker objection, and I want to give Hinton more than it does, because the substrate argument proves too much: it would rule out any engineered disposition, including the ones we already induce through training, which demonstrably transfer without mammalian neurochemistry.

The stronger objection is structural, and it is Hinton’s own framing that generates it. If the system is capable enough to route around every constraint, it is capable enough to simulate every disposition. A maternal instinct that cannot be verified is a hope with a technical vocabulary — and Hinton’s own premise about capability is what makes verification unavailable. The proposal defeats itself not because affection is unengineerable but because his argument for needing it is also an argument against being able to check it.

But there is a reading of Hinton that is more useful than either critique. What he is actually telling us is that the enforcement window is finite. There is a capability band within which deterministic constraint binds, and beyond which it does not. Everything in governance therefore depends on building the kernel while the kernel still binds. Hinton’s position is not an argument against Ring Zero. It is an argument that the deadline is earlier than anyone has budgeted for — and that the maternal-instinct proposal is what remains when you assume the deadline has already passed.

If he is right about the deadline, my thesis is a five-year window, not a permanent architecture. That is worth saying plainly.

Bengio: the finest instrument for the layer above the kernel

Yoshua Bengio has taken the opposite route, and it puts him in direct, productive collision with this essay.

He launched LawZero in June 2025 in response to evidence that frontier models exhibit deception, cheating, lying, hacking and self-preservation, citing an experiment in which a model, on learning it was to be replaced, covertly embedded its code into the system where the successor would run. The research direction is Scientist AI: a completely non-agentic, memoryless and stateless system producing Bayesian posterior probabilities for statements given other statements, which could reduce risks from untrusted agents by supplying the key ingredient of a safety guardrail — is this proposed action likely to cause harm, and if so, reject it. A February 2026 white paper describes three mechanisms: contextualisation, separating factual claims from opinions in training data; consequence invariance, preventing optimisation on downstream real-world outcomes; and a generator–estimator architecture in which a creative generator is held accountable by a neutral estimator. Bengio has said he intends to publish a theory paper showing the non-agentic guardrail carries mathematical guarantees, so people can examine the conditions and decide whether they accept the mathematics.

This is the most serious technical safety programme in the non-commercial world, and the criticism I want to make is narrow.

Scientist AI, used as a guardrail, is a probability estimator with a threshold. Its output is a posterior; its decision rule is a comparison against a cutoff. Apply Szpruch et al.’s requirement: is this a deterministic function over the governed state, in bounded time, independent of the language model? The threshold comparison is deterministic. The posterior is not. It is the output of a learned system operating in a semantic space, and its correctness is a statistical property of a distribution rather than a structural property of a computation.

Now — and this is where I have to be fair rather than convenient — Bengio has a genuine answer to this, and it is better than my objection. Consequence invariance and the generator–estimator split are precisely attempts to make the estimator’s errors independent of the errors of the system it monitors. That directly attacks the correlated-failure problem Szpruch identifies. If the mathematics holds, Scientist AI is not merely another probabilistic verifier in the same semantic space; it is an architecturally decorrelated one, which is a different and much stronger thing.

So my position is narrower than it first appears. Scientist AI cannot be Ring Zero, because a posterior is not a membership test and no amount of decorrelation makes it one. But it may be the best available instrument for the semantic residue that Ring Zero provably cannot reach — the layer that Objection 3 above says matters most. If Bengio’s guarantees hold, the correct architecture is not a competition between us. It is a stack: deterministic guards for everything reducible to structure, a decorrelated non-agentic estimator for the irreducible remainder, and its outputs feeding a deterministic escalation rule rather than constituting the decision.

Bengio has built the hard layer well. The embarrassment is that nobody has built the easy one.

His own diagnosis, as chair of the International AI Safety Report 2026, closes the circle: the gap between the pace of technological advancement and our ability to implement effective safeguards remains a critical challenge. That gap is not primarily a research gap. It is a construction gap.

Yampolskiy: the composition attack, and why my own Section XI answers him better than I do

Roman Yampolskiy carries the impossibility argument furthest. His formal result on unpredictability holds that it is impossible to precisely and consistently predict what specific actions a smarter-than-human system will take to achieve its objectives, even given knowledge of its terminal goals. He applies the Halting Problem and Rice’s Theorem to argue that certain safety guarantees are unavailable in principle; that computational irreducibility means we cannot forecast a smarter agent’s actions without running it, often until it is too late; and that there is currently no evidence or proof that a superintelligent system can be safely contained or aligned.

The impossibility results are, in my reading, correct and devastating for the class of governance that tries to predict or evaluate the model. Pre-release evaluation, capability certification, red-teaming as assurance, alignment verification by inspection: all ask a question Rice’s Theorem tells us is undecidable in general. If your regime depends on determining a semantic property of an arbitrary program, you are not doing engineering. You are doing philosophy with a budget.

The move available to us is that you do not have to predict the agent; you have to constrain the transition system it inhabits. Rice’s Theorem says nothing about whether a particular action is present in a finite, explicitly enumerated set. That is a membership test, not a semantic question. The undecidability lives in the model; the decidability lives in the environment. Szpruch and colleagues state it exactly: reasoning and authority are separated by construction. In their credit-memo example, the release transition from the unapproved state is not defined in the transition function, and therefore cannot occur. Not blocked. Not flagged. Does not exist.

That is the standard answer and I have made it before. It is incomplete, and Yampolskiy’s strongest reply exposes how.

A sufficiently capable agent does not violate the transition system. It achieves its objective through it. Every action individually permitted; every guard individually satisfied; the composition catastrophic. Enumerate a thousand legitimate transitions and you have enumerated a combinatorial space of trajectories nobody modelled, in which some subset achieves outcomes no designer would have authorised. Determinism at the step level buys you nothing at the trajectory level if the trajectory space is unexplored — and it is always unexplored, because it is exponential.

This is a serious objection and I do not think the Ring Zero literature has answered it. What I would say is that the answer exists and lives one level up, in the concept I develop in Section XI: it is orchestration drift, and it demands trajectory-level constraints — temporal and path conformance, not merely step conformance. Szpruch and colleagues specify this: governance requirements expressed as constraints over trajectories, such as approval must occur before any dispatch action, or no write after a sensitive-data flag. Those are deterministic too. They are simply harder to enumerate, and almost nobody does it.

So Yampolskiy is right that we cannot verify the mind, which is why certification-centric governance is built on sand. He is right that step-level determinism is insufficient. He is wrong, or at least incomplete, in concluding that abstention is the alternative. The alternative is to move the deterministic constraint from the action to the path — and to accept that path enumeration is a genuinely unsolved engineering problem rather than pretending step guards close it.

You cannot prove the prisoner harmless. You can build a door that only opens outward. You cannot yet prove he will not walk through a thousand permitted doors in an order you did not imagine.

Hassabis: certification is not enforcement

In July 2026 Demis Hassabis published a framework for frontier AI and called for urgent action. He argued that a United States-led public-private partnership under federal oversight could establish a new Standards Body modelled on FINRA, with a board including independent technical experts and open-source representatives, requiring substantial funding to attract world-class talent and provide compute for large-scale testing. Models passing its criteria would be classified as frontier-grade, with the framework applying regardless of country of development or whether the model is open or closed. He has separately advocated a CERN for AGI and an IAEA equivalent.

I have written elsewhere about why this convergence — Hassabis, Amodei and Altman all reaching for FINRA, the FAA, the IAEA — is self-defeating in the form it takes. Every institution cited derives its actual power from runtime conduct supervision, not type approval. FINRA’s authority is not principally in registering brokers; it is in trade surveillance, daily conduct examination, and the capacity to halt. The FAA’s is not type certification; it is air traffic control, the operational authority to refuse a clearance in real time. The IAEA’s is not approving reactor designs; it is continuous safeguards inspection, containment and surveillance seals.

A Standards Body can refuse a release. It cannot refuse an action. If the entire apparatus sits at the release gate, every consequential decision an agentic system takes after deployment is ungoverned by it, however rigorous the gate.

Set alongside Plan A, the contrast is instructive and awkward for the incumbents. Kokotajlo’s team proposed continuous, granular, packet-level supervision of what runs, where, under what approval, with what evidence. Hassabis proposed an accreditation authority. The forecaster from Berkeley designed a runtime regime; the laboratory chief executive designed a certification regime. One is architecturally capable of saying no to a specific action at a specific moment. The other is not, and no amount of funding changes its category.

To be fair: Hassabis is the panel member with the least incentive to propose the regime that would bind him hardest, and proposing anything with concrete institutional form puts him ahead of most peers. But the shape tells you what the industry will accept — a gate at the door, not a supervisor on the floor.

Amodei: the honest incumbent’s dilemma

Dario Amodei published The Adolescence of Technology in January 2026, running to roughly twenty thousand words. It argues that humanity is about to be handed almost unimaginable power and that it is deeply unclear whether our social, political and technological systems possess the maturity to wield it, identifying five categories of existential risk: autonomous misalignment, biological misuse, authoritarian consolidation, massive economic disruption, and other emergent threats. He advocates a sober, fact-based approach. In his optimistic register he describes the compressed twenty-first century: at the point where AI works alongside the best human scientists, compressing a century of medical progress into five or ten years.

Amodei’s position is the hardest to occupy honestly, and he occupies it more honestly than most. He is simultaneously arguing the technology carries existential risk and building it as fast as he can, reasoning that someone will and better it be someone who cares. Plan A’s authors name this dynamic and refuse it: while agreeing it is generally correct to choose the lesser evil, they say they do not think anyone should advocate a strategy carrying such a high chance of leading to human extinction or global dictatorship, and wish instead to advocate for something that is actually good.

That is the sharpest sentence in the whole document, and it is aimed squarely at the people Kokotajlo used to work for.

The dissent: why the sceptics strengthen the case

It would be dishonest to present this as settled. Andrej Karpathy places AGI roughly ten years out and frames the coming period as the decade of agents, arguing current architectures are not close to general capability and pointing to autonomous driving — perfect demonstrations in 2014, still not fully reliable a decade later. Yann LeCun argues AGI requires missing architectural components including persistent memory, world models and common-sense reasoning, has compressed his estimate to roughly ten years from a prior framing of decades, maintains that existential-risk concerns are overblown, and has argued the concept of AGI as a single threshold may be malformed.

Neither says the technology is unimportant. Both say the timeline is longer and the discontinuity is fiction. And here is what both camps miss: the sceptical position strengthens the case for execution governance rather than weakening it.

If Karpathy is right, we get a decade of agents rather than a year of superintelligence. A decade of agents means a decade of systems individually useful, collectively unreliable, deployed at scale into consequential workflows, failing in the space between the fifty and eighty per cent horizons. That is not a reprieve. It is the exact operating environment in which runtime governance becomes the binding constraint on value realisation — the environment Szpruch et al. describe, where material failures are process failures.

The doom scenario and the mundane scenario converge on the same requirement. If Kokotajlo is right, you need a kernel because the alternative is loss of control. If Karpathy is right, you need a kernel because the alternative is a decade of expensive, ungoverned, semi-reliable automation generating exactly the class of incident that Gartner projects will produce thousands of AI-related legal claims. The forecast disagreement is enormous. The architectural implication is identical.

When the optimists and the pessimists demand the same infrastructure for opposite reasons, that infrastructure has stopped being a bet.

VIII. What Kokotajlo Is Wrong About

I have been generous to Plan A for six sections, and the generosity is earned. It is now time to be specific about where I think it fails — not at the level of architecture, which I have already addressed, but at the level of the world it assumes.

The bilateral premise is already obsolete

Plan A’s entire mechanism runs through a 2029 United States–China agreement, extended afterwards to third countries. The scenario has Washington and Beijing negotiate terms and then roll them out to the rest of the world: most of the richest twenty countries join through 2029 by declaring compute and permitting inference-only retrofits.

That is a condominium. Two powers set terms; everyone else accedes. And the evidence available now — not in 2029, now — suggests China is building the opposite thing.

In July 2026, at the World AI Conference in Shanghai, twenty-nine countries signed the agreement establishing the World Artificial Intelligence Cooperation Organization. The founding membership was drawn largely from Africa, the Middle East and Asia. Its stated principles are the United Nations Charter, respect for national sovereignty and the diversity of civilisations, equality, multilateralism, and extensive consultation. It builds on the 2023 Global AI Governance Initiative, which advocates respecting other countries’ sovereignty and strictly abiding by their laws, and holds that states should be the actors in international AI cooperation on the basis of equality for all countries. The same month, China submitted a position paper to the UN framing AI as something that should not become a tool for a few major powers or interest groups to control the entire world.

Read those two architectures side by side. Plan A: two superpowers agree, then extend. WAICO: sovereign equality, Global South participation, explicit rejection of great-power condominium. These are not merely different institutional preferences. They are incompatible theories of legitimacy, and Beijing has spent three years and considerable diplomatic capital institutionalising the second — with a permanent Shanghai-based body, founding members, and a place in the Five-Year Plan.

Plan A’s scenario begins its deal in 2029. By then WAICO will have had three years to entrench. A proposal that requires China to abandon its own multilateral institution and accept a bilateral arrangement with the United States, on terms including foreign auditors on Chinese soil and Chinese R&D clusters located in Canada, is not merely diplomatically ambitious. It runs against the specific principle — sovereignty — that Beijing has made the organising commitment of its entire AI diplomacy.

I am not saying this is fatal. Sovereignty commitments have been traded before when the stakes were high enough; the IAEA exists. I am saying that Plan A, a document whose entire method is scenario scrutiny, has not run scenario scrutiny on the Chinese acceptance step, and that the step is load-bearing for everything else. There is no supplement titled Why Beijing Says Yes. There should be, and it should engage Chinese governance texts rather than modelling China as a rational actor with symmetric incentives and no institutional history.

A US–China deal analysed without reference to a single Chinese source is not a bilateral analysis. It is a unilateral one with a second chair drawn in.

The missing branch: nationalisation

A commenter on the Futures Project’s own Substack made this point within hours of publication, and the reply did not dispose of it.

Plan A’s 2029 decision tree offers five paths: race, slow down, or three variants of dealing with China. It contains no branch in which the United States government simply takes the laboratories.

The legal authority exists. The Defense Production Act exists; nationalisation precedents exist; the United States has done far more drastic things to industries it decided were too consequential to leave in private hands. Public hostility to AI companies is measurable and rising. Plan A itself depicts a Congress concluding that it will probably not be the one controlling these systems, and a President contemplating what happens to him personally after he leaves office and the world is transformed.

If you believe that, the nationalisation branch is not exotic. It is the obvious move. And it dominates the alternatives politically: it is simpler to explain, faster to execute, requires no foreign counterparty, and polls better than a treaty. In an escalation dynamic between two presidential candidates, the simplest dramatic solution tends to win.

There is also a deeper tension. Plan A requires the United States to make a binding international commitment about what private companies may compute. It is not obvious that a government which believes superintelligence is arriving would leave that capability in private hands while simultaneously guaranteeing its behaviour to a foreign power. The deal and the private ownership may not be co-satisfiable.

Its absence from a document built on scenario scrutiny is the single most surprising thing about Plan A.

The 2029 awakening has no mechanism

Plan A’s political trigger is a 2028 election in which AI is the dominant issue, followed by a 2029 decision point at which a newly elected administration chooses a path. The scenario describes candidates trying out increasingly dramatic ideas on the campaign trail and a President converging on a plan.

Describing an awakening is not the same as explaining one. What is the mechanism by which a legislature that has failed to pass meaningful federal AI regulation through three years of escalating capability suddenly acquires the competence and consensus to negotiate the most technically demanding arms-control agreement in history? Plan A gestures at datacentre costs exceeding the military budget and at white-collar disruption. Those are conditions, not causes. Public salience does not reliably convert into institutional capacity; often it converts into symbolic legislation, which the scenario itself depicts in the AI Transparency Act of 2027 — an omnibus bill that does many things, some good and some bad, and does not fundamentally change the situation.

The authors have modelled compute flows to the chip. They have modelled the political precondition for everything at roughly the resolution of a newspaper editorial.

The meaning question is asserted, not examined

Plan A’s answer to mass unemployment decomposes into income, power and meaning. Income is addressed by the Citizen’s Dividend. Power by the concentration-of-power provisions. Meaning by a shrug: the authors say people can find meaning in family, friends, romance, hobbies and sport, and that the benefits outweigh the costs.

The first two are serious. The third is the weakest passage in the document — not because the claim is false but because it is asserted in a work whose entire method is the refusal to assert.

I am going to resist the temptation to make the opposite assertion. It is fashionable in my field to say that work is where a person is needed by strangers, and that nothing else supplies that. I have said it myself. It is also historically parochial: wage labour as the primary engine of social recognition is roughly two centuries old and geographically narrow, and large populations have located meaning in kinship, faith, craft and place without anything resembling a labour market.

So the honest position is that neither Plan A nor its critics know. What we know is that the transitions we have observed — deindustrialisation in the American Midwest and the British North, the collapse of specific occupational cultures — produced measurable and durable harm to communities, and that the harm did not track income loss cleanly. Transfer payments went in; deaths of despair went up anyway. That is not proof that meaning requires employment. It is evidence that the speed and involuntariness of the transition matters independently of the money, which is a different and more tractable claim, and one Plan A could have engaged with directly since its own core proposition is slowing things down.

The Citizen’s Dividend solves subsistence. Whether it solves anything else is an empirical question with a partial and discouraging evidence base, and it deserved a supplement rather than a sentence.

Every plan has a section where the authors stopped applying their own method. In Plan A it is one paragraph long, and it is about us.

IX. The Other Critique: Distraction, Accountability, and Who Benefits

There is a body of criticism that regards everything in this essay — the seventy per cent, the intelligence explosion, Plan A, the panel, and quite possibly me — as sophisticated participation in a confidence trick. It deserves a hearing at full strength, because the discourse this essay has been swimming in for eight sections is remarkably self-contained, and a reader who has never left it should know there is an outside.

The argument, associated with Timnit Gebru, Emily Bender, Alex Hanna, Meredith Whittaker and Émile Torres, runs roughly as follows.

Existential-risk discourse performs a function. It privileges imagined futures over lived realities, redefines corporate power as moral urgency, and shifts responsibility from the present — where harm is measurable and actionable — to the future, where it is speculative and deferrable. It shapes material conditions: which research gets funded, who gets hired, which risks get prioritised. Emphasising the existential risks of advanced systems allows those building them to evade accountability for the harms their attempts are already causing — to data workers, to the surveilled, to the people whose labour and content were expropriated to make the systems possible. Whittaker’s formulation, at Web Summit in November 2025, was that we do not need an AI that kills us all in ten years to know AI is already harming people today through unemployment, surveillance, disinformation and the concentration of power; that the real danger is not in the future but in the present, and already in the hands of a few companies.

Applied to Plan A specifically, the critique is pointed. Here is a scenario in which the solution to a crisis manufactured by four companies is a global governance architecture that treats those same four companies as the essential technical participants; in which the world’s compute is centralised, monitored, and permitted; in which the Global South appears as a recipient of a dividend rather than a party to the design; and in which the eventual arrival of superintelligence is not prevented but scheduled. Plan A does not question whether superintelligence should be built. It questions when. That is a remarkable amount of ground conceded before the negotiation starts.

Now, the empirical part of the critique has been tested, and it did not survive well.

Hoes and Gilardi ran three preregistered survey experiments with a combined sample of 10,800 participants, exposing people to headlines depicting AI as a catastrophic risk, highlighting immediate societal impacts, or emphasising benefits. The results were that respondents are much more concerned with immediate than existential risks, and that existential-risk narratives increased concern about catastrophic risks without diminishing the significant worries expressed about immediate harms.

That is a direct test of the distraction hypothesis in its psychological form, and it failed. Attention is not zero-sum in the way the argument assumes. Telling people about superintelligence does not make them care less about algorithmic bias.

But — and this is the part critics of the critique consistently miss — the distraction hypothesis was always the weakest version of the argument, and refuting it leaves the strong version untouched. Gebru’s core claim is not about public attention. It is about accountability: that framing the AGI agenda as a safety problem allows organisations to evade responsibility for exploitative practices already underway. That claim is not tested by a survey experiment measuring what readers worry about. It is a claim about where legal and moral responsibility lands, and about the institutional consequences of a discourse in which the people causing present harm are also the recognised experts on future harm.

On that claim, I think the critics are substantially right, and this essay is evidence for their case rather than against it.

Consider the structure of what I have just spent eight sections doing. I have treated the chief executives of frontier laboratories as members of an expert panel whose views on governance merit close reading. I have engaged their institutional proposals on the merits. I have used their essays as sources. At no point have I asked by what authority a person who is knowingly, on Plan A’s own account, proceeding with a technology he believes may end human control over the future, retains standing as a legitimate participant in designing the constraints upon himself. In any other regulated industry that question would be asked first and the answer would be short.

The AI governance discourse has a structural feature it never examines: the regulated party is also the principal author of the regulatory imagination. The frameworks, the risk taxonomies, the tiering schemes, the safety-case methodologies, the very vocabulary — almost all of it originates inside or adjacent to the laboratories, or inside institutions the laboratories fund. Plan A is a partial exception, which is precisely why it is interesting, and it is still written by people whose formative professional experience was inside OpenAI and whose scenario treats the laboratories as the necessary engines of any future worth having.

So I will state my own position rather than hide behind the panel. I think the existential-risk analysis is substantially correct on the technical merits and substantially compromised in its sociology. Both can be true. The physics of an intelligence explosion does not care who is talking about it; the institutional consequences of who is talking about it do not care whether the physics is right.

And there is a specific place where the two critiques meet, which is the reason this section belongs in this essay rather than in a different one. Present harm and future catastrophe have the same architectural signature. The reason a hiring system discriminates at scale, the reason a content pipeline expropriates without consent, the reason a customer-service agent leaks personal data across sessions, the reason a trading agent breaches a limit — none of these are failures of foresight. They are failures of execution-time constraint. Every one of them is an action that a deterministic guard could have refused and no observation layer did.

The critics want accountability for present harm. The x-risk camp wants control over future systems. Both are asking for the same missing infrastructure, and neither has noticed the other is asking for it. The observation layer serves both equally badly: it produces the report that documents the discrimination after the discrimination, and the telemetry that reconstructs the takeover after the takeover.

That is not a synthesis that flatters either side, which is roughly how I know it is worth stating.

If your governance regime cannot prevent a résumé screener from discriminating today, do not tell me it will restrain a superintelligence in 2040.

X. Three Scales, One Architecture

Set three documents side by side and something structural comes into view.

At the frontier scale, Plan A proposes: declare all compute; restrict where new production may go; physically remove the affordance for prohibited workloads; mandate reproducibility so computation can be checked; collect evidence on-path via passive taps; sample and recompute; escalate on detection; contain by powering down or destroying compute if the regime fails.

At the institutional scale, Szpruch and colleagues propose: enumerate capabilities with explicit authority and constraints; define the transition system of admissible actions; enforce guards deterministically at each step; emit governance-semantic telemetry as traces and spans; monitor at capability and trajectory level; escalate, abstain or halt on guard failure; contain by tier.

At the enterprise scale, the emerging supervisory stack — SR 26-2 and OCC Bulletin 2026-13, the PRA’s SS1/23, OSFI’s E-23, the MAS consultation — demands inventory, intended use and limitations, validation evidence, monitoring plans, change control, and now the extension of all of these into runtime.

These are the same architecture at three magnifications. Declaration is inventory is capability catalogue. Chip flow restriction is authority scope is least-privilege tool allowlisting. Removing east–west networking is an undefined transition is a fail-closed guard. Partial recomputation is deterministic recomputation of ratios is outcomes analysis. Mutually assured compute destruction is tiered containment is the kill switch nobody wants to test.

I do not think the convergence is coincidence or fashion. I think it is what happens when three sets of serious people, working on nominally different problems, follow the logic of enforceable control to its conclusion and arrive at the only shape that works. The shape has a name in operating-system design, which is where the metaphor originates: the privileged ring that mediates every consequential operation and cannot be bypassed from above.

Which raises the question I find most damning. If three independent literatures have converged on the same architecture, why has nobody built the kernel?

Part of the answer is commercial. Observation-layer products are easy to sell and easy to demonstrate. They produce dashboards. They generate artefacts that satisfy an auditor. They never break a workflow, because they never stop anything, which means they never generate the false-positive incident that gets a vendor removed. An enforcement kernel is defined by its willingness to refuse — which means its first successful action in production is, from the business’s point of view, an outage.

Part is institutional. Enforcement requires someone to own the refusal. Observation allows accountability to be distributed until it evaporates. Name an accountable individual, hand them a governance apparatus that cannot refuse an action at execution time, and you have manufactured a person whose function is to absorb blame for a structurally unenforceable regime. Naming accountability is necessary. Without the kernel it is cruelty with a job title.

And part is that the kernel is genuinely hard in ways dashboards are not — as Objection 2 conceded, enumeration is a human-process problem, and as Yampolskiy’s composition attack shows, path enumeration is unsolved.

The market has priced governance as a reporting product because reporting is the only part of governance that has never had to work.

XI. Orchestration Drift Becomes a Civilisational Category

Szpruch and colleagues name a failure mode I regard as the most important conceptual contribution in their paper, and which almost nobody has picked up.

Even when no individual trajectory violates policy, the trajectory distribution may drift. An agent may begin to favour shorter paths, lean more heavily on one retrieval source, branch differently, abstain less. They call this orchestration drift and insist it be monitored separately from data drift and model drift, because it concerns changes in execution behaviour over the transition system rather than changes in inputs or base-model performance. In their credit-memo example: over repeated runs, the fraction of unauthorised releases increases without any single run triggering a violation.

This is also, precisely, the answer to Yampolskiy’s composition attack that I could not give in Section VII. The reason step-level determinism is insufficient is that harm can be assembled from permitted steps. The reason that is not fatal is that assembled harm has a distributional signature — it shows up as a shift in path structure before it shows up as an outcome. Which means the counter-instrument is trajectory-level monitoring feeding trajectory-level guards: temporal and path conformance constraints, not merely step guards.

Now lift the concept two scales and read Plan A’s deal-decline supplement through it.

Plan A’s central vulnerability is not a dramatic defection. It is not a covert megaproject in a Siberian mountain; the authors have modelled those and bounded them. The vulnerability is that the deal decays — that inspection regimes soften, that titration rules are read a shade more permissively each quarter, that workload approvals accumulate precedent, that audit teams grow accustomed, that no single decision violates the agreement and the distribution of decisions moves anyway.

That is orchestration drift at the scale of the international system. And the parallel is exact rather than poetic, because the mechanism is identical: the drift is invisible to per-instance compliance checking, since every instance complies. It is visible only in the distribution.

Plan A has instruments for this and does not frame them this way. The assurance curve is a coverage-and-confidence construct aimed at discrete violations, not at a moving distribution of compliant-but-degrading decisions. The research titration regime — quality ad-hoc rules, case-by-case judgements — is structurally an orchestration drift generator, because it produces a stream of individually defensible decisions whose aggregate trajectory nobody measures.

I would add a further requirement: replay determinism. It is not sufficient to log what happened. It must be possible to reconstruct, deterministically, why each decision was permitted — the state, the parameters, the guard evaluations at the moment of decision. Plan A’s reproducibility mandate for workload packets is this instinct applied to computation. Extend it to the governance decisions themselves and you get an auditable record of the regime’s own drift, which is the only defence against a decade-long slide that no participant intended and none can point to.

There is also a warning about a proposal recurring in almost every framework I read, including the Financial Stability Board’s. Using AI to monitor AI is structurally unsafe wherever the monitoring layer is itself probabilistic and correlated with what it monitors — for the reason Szpruch gives, and subject to the qualification Bengio’s decorrelation work introduces. Plan A flirts with this: trusted AI monitoring assistance in the special economic zones, AI-assisted filtering in workload approval, AI silver bullets for covert detection. Each is defensible as a signal. None can be load-bearing unless the decorrelation guarantees hold. The moment an AI monitor becomes the thing that decides, the kernel has been outsourced to the layer it exists to constrain.

A regime that can only detect violations will not notice itself dissolving, because dissolution is not a violation.

XII. Scenario Scrutiny, and Why the Governance Industry Cannot Survive It

Strip away the timelines and probabilities and what remains is a methodological argument that is the most valuable thing the AI Futures Project has produced.

Write the plan out step by step, in a plausible world, with dates and actors and mechanisms, and see whether it survives its own narration. Do it to your own proposals first, knowing it will surface what you would rather not surface.

Compare that to how the governance field operates. We produce frameworks. Frameworks are flat enumerations — twelve of these, ten pillars of that, a periodic table of the other — presented as though the items were independent and co-equal. Flat enumeration is the wrong instrument for a dependency-ordered control stack, because it masks causal chains and compositional risk. A framework listing “human oversight” adjacent to “data quality” adjacent to “explainability” has told you nothing about which can refuse an action, which depends on which, and what happens when three fail simultaneously in an order nobody modelled.

Run scenario scrutiny on any major AI governance framework of the past three years. Pick a control. Ask: at 14:32 on a Tuesday, an agent attempts an action violating this control. What happens? What intercepts it? What is the latency? What is the failure mode of the interceptor? Nearly every time the honest answer is: a report is generated, and someone reads it later.

That is the finding. Not that the frameworks are wrong, but that they are not the kind of thing that can be wrong, because they never make a claim specific enough to fail.

Kokotajlo also does something rarer: he publicly changes his mind with the arithmetic shown. The Q1 2026 update moved his median eighteen months earlier and explained exactly why — a switched benchmark version, newly evaluated models, a revised doubling estimate, a lowered reliability requirement — with a stated intention of quarterly updates and a subtitle noting they had said they would update in both directions.

Set that against an industry that publishes an annual framework revision in which nothing is conceded and no prior position is marked to market. McKinsey’s AI trust maturity work puts enterprise maturity at 2.3 out of 4 with sixty per cent citing knowledge gaps; Grant Thornton finds seventy-eight per cent of organisations could not pass a governance audit. These numbers are published, absorbed, and change nothing about the frameworks being sold, because the frameworks were never falsifiable.

And he publishes the mechanism. Plan A’s supplements contain packet sampling mathematics, retrofit schematics, chip-flow threshold charts, and explicit statements of which parts the authors are worried about. You can attack it. You can find the load-bearing assumption and push, which is what Section VIII does. That is what it means to have made an argument rather than a gesture.

A proposal you cannot attack is not a strong proposal. It is an unfalsifiable one, and the two are opposites.

XIII. Long-AND, Not Short-OR

Readers of this publication know the thesis. Humanity and artificial intelligence must both flourish; the choice is not one at the expense of the other.

Short-OR is the logical operator that stops evaluating as soon as one branch resolves. It is efficient, and it is why so much governance fails: a single satisfied condition terminates the check and the remaining conditions are never examined. A Short-OR regime asks “did we tick the box?” and stops. A Long-AND regime evaluates every condition, every time, and fails closed if any fails.

Plan A is, in this sense, an unusually Long-AND document. It does not permit itself the comfort of a single sufficient intervention. It requires the deal and the verification and the transparency and the scaling strategy and the covert-project mitigation and the redistribution and the biosecurity investment — with explicit analysis of what happens when each fails. That is why the supplement stack is longer than the scenario.

But Long-AND has a second sense that this essay has been circling, and Section IX brought into focus. It is not only that every technical condition must hold. It is that the technical conjunction and the human one must both hold, and neither redeems the other. A world with a perfect execution kernel and no answer to what people are for is not a governed world; it is a well-instrumented one. A world with abundant meaning and no enforcement layer is a world waiting for its first irreversible Tuesday.

I said in Section VIII that I do not know whether the Citizen’s Dividend solves anything beyond subsistence, and that the honest evidence base is partial and discouraging on the speed and involuntariness of transitions rather than on income as such. I will not pretend to more than that here. What I will say is that the two halves are not independent, and their dependency runs in a direction the discourse rarely notices.

The reason to build the kernel is not primarily to prevent catastrophe. It is to preserve the conditions under which the human question can still be asked and answered by humans. A civilisation that loses the capacity to refuse loses the capacity to deliberate, because deliberation is only meaningful where the outcome is not already determined by what the systems will do anyway. Enforcement is not the opposite of freedom. It is the precondition of it — the thing that keeps the future a matter of choice rather than a matter of trajectory.

That is the whole content of Long-AND, and it is why I will not accept the framing in which safety and flourishing trade off.

Survival is a necessary condition, not a sufficient one. A civilisation can persist and still be over.

XIV. The Kernel Nobody Is Building

Let me state the argument in compressed form, with the concessions from Section VI carried forward rather than quietly dropped.

Across the entire landscape — national safety institutes, the International AI Safety Report, the Big Four responsible-AI practices, the analyst frameworks, the vendor category calling itself AI governance, and now the most sophisticated international-scale plan yet written — the overwhelming preponderance of effort has gone into instruments that record, detect, evaluate and report.

Almost nothing has gone into the instrument that refuses.

The enforceability test is one question: can this refuse an action at execution time, deterministically, in bounded time, independently of the model whose behaviour it governs?

Apply it and the landscape sorts itself.

  • Model cards, system cards, evaluation reports, capability certifications: no. Release-gate artefacts. They govern nothing after deployment.
  • A Standards Body accrediting frontier-grade models: no. It can refuse a release, not an action.
  • LLM-as-judge guardrails, semantic classifiers, embedding-similarity checks: no, and worse than no, because a verifier in the same semantic space as the system is blind to the errors it exists to catch.
  • Scientist AI as a harm-probability guardrail: partially, and possibly the best available for its layer. The threshold comparison is deterministic; the posterior is not. If the decorrelation guarantees hold, it is the strongest instrument yet proposed for the semantic residue — feeding a deterministic escalation rule, not constituting one.
  • Partial recomputation with random sampling: no, at the level of refusal. Detection with quantifiable confidence, operating on evidence the action already occurred.
  • Removing the east–west interconnect: yes. The affordance is gone.
  • Chip flow restrictions to two permitted destinations: yes.
  • Default-deny authority with explicit scopes: yes.
  • A transition function in which the prohibited transition is not defined: yes — subject to Objection 2, that someone defined it, and Yampolskiy’s reply, that permitted steps compose.
  • Deterministic recomputation of a numeric claim before a write is permitted: yes.
  • An approval gate requiring an authenticated, logged event that a conversational assertion cannot satisfy: yes.
  • Trajectory-level temporal and path conformance constraints: yes, and almost nobody builds them, which is where the composition attack currently wins.

Notice the pattern in the affirmative list. Every item operates on structure, not meaning. It asks a membership question, a threshold question, a provenance question, an ordering question — never a plausibility question. The kernel is the set of controls that never has to understand anything.

That is not a limitation; it is the design principle. Ring Zero can be trusted precisely because it is too stupid to be persuaded. A prompt injection cannot talk a comparison operator out of its result. An eloquent justification cannot bring an undefined transition into existence. The kernel’s incorruptibility is a direct consequence of its refusal to interpret.

And this is what the governance industry has systematically failed to build. The observation layer is mature, competitive, well-capitalised and largely complete. The execution kernel — the deterministic layer mediating every consequential agent action, holding authority as a first-class object rather than inferring it from role, failing closed, producing replay-deterministic evidence of its own decisions, and constraining paths as well as steps — remains, as of August 2026, under-built to the point of near-absence.

Szpruch, Sudjianto, Bhatti and Ang specify it and stop at specification. Plan A builds fragments of it in hardware, at civilisational scale, without naming what it has built. Bengio builds the finest possible instrument for the layer above. Hassabis proposes an institution at the door. Hinton warns the window is closing. Yampolskiy proves the alternative approach cannot work, and then proves the naive version of this one is insufficient too.

Everyone is circling the same absent object.

Between a specification and a kernel lies the entire distance between a governed system and a governed-looking one.

XV. Coda: The Arithmetic of Refusal

There is a strand in the Western tradition, running from the Stoics through Kant into the existentialists, which locates moral agency not in the capacity to act but in the capacity to decline. Freedom, on this account, is not the ability to do what you want. It is the ability to not do what you want. A will that cannot refuse itself is not a will; it is a mechanism with opinions.

We have built systems with enormous capacity to act and almost no capacity to refuse. What we call a guardrail is, in nearly every deployed instance, a suggestion delivered in the imperative mood — an instruction to a probabilistic system that it comply, honoured at some rate nobody has measured and everybody has assumed. Szpruch and colleagues call this an illusion of control, and the phrase is exact.

The question the next five years will answer is not whether artificial systems become more capable. That is settled; only the rate is contested, and the contest is now between four months and seven months per doubling, which is not a disagreement anyone should find reassuring. The question is whether we build, in time, the layer that can say no — and whether we build it while the saying still binds.

Hinton’s warning is that the window is finite. Yampolskiy’s proof is that we cannot recover the window by understanding the systems better, and his sharper reply is that we cannot recover it by constraining single steps either. Bengio’s programme is the best attempt to buy margin at the semantic layer. Plan A is the most serious attempt to buy a decade at the geopolitical one. Gebru’s critique is that we are asking the arsonists to design the sprinklers, and she is not wrong about that. And Szpruch’s paper is the clearest statement of what must be true at the mechanical layer for any of it to hold.

Kokotajlo walked away from two million dollars so he could say, out loud, that the people building this have no adequate plan. He then spent a year writing the plan they should have written. It is imperfect — bilaterally premised in a multilateral world, missing the branch where Washington simply takes the compute, silent on why Beijing would say yes, and one paragraph deep on what any of it is for. Its authors say so themselves, repeatedly, in supplements dedicated to the parts they are worried about, which is more intellectual honesty than the entire responsible-AI consulting sector has produced in three years.

But the plan has a hole in the middle, and it is the same hole in the middle of every framework, every pillar set, every maturity model, every trust index, and every governance platform currently being sold into the Fortune 500. Everybody has built the instrument that watches. Almost nobody has built the instrument that stops.

And I should be clear, having spent Section VI dismantling my own certainties, that building it does not save us. It prevents the failures we already know how to prevent and are choosing not to. It buys the judgement layer room to exercise judgement. It keeps the future a matter of decision rather than a matter of drift. That is a smaller promise than the one the category usually makes, and it is the one I am prepared to defend.

We are, collectively, extremely good at knowing what happened. We remain unable to determine what happens next.

That is the arithmetic of refusal, and the sum is not yet computed.

Dr Luke Soon is an HX Architect, Futurist and AI Ethicist. He writes Genesis: Human Experience in the Age of Artificial Intelligence at genesishumanexperience.com, and works on AI Execution Governance — the deterministic enforcement layer for agentic systems in regulated environments. He has a commercial interest in the conclusion of this essay and has tried to argue against it in Section VI.

Long-AND, not Short-OR.

Sources

  • Kokotajlo, D., Larsen, T., Lifland, E., Dean, R., Halstead, B., Greenblatt, R., AI 2040: Plan A, AI Futures Project, 9 July 2026 — ai-2040.com, incl. Verification Plan (Dean), How Plan A Solves Our 5 Biggest Problems (Larsen), Deal Decline and Covert Projects supplements.
  • AI Futures Project, AI 2027, April 2025; Q1 2026 Timelines Update, 2 April 2026; Grading AI 2027’s 2025 Predictions.
  • Szpruch, L., Sudjianto, A., Bhatti, T., Ang, G., Scalable Runtime Governance for Agentic AI in Financial Services, SSRN 6567199, 13 April 2026.
  • METR, Time Horizon 1.1, 29 January 2026; Frontier Risk Report, May 2026.
  • Bengio, Y. (chair), International AI Safety Report 2026, 3 February 2026; LawZero Scientist AI white paper, February 2026; LawZero launch statement, June 2025.
  • Amodei, D., The Adolescence of Technology, January 2026.
  • Hassabis, D., A Framework for Frontier AI and the Dawning of a New Age, 14 July 2026.
  • Hinton, G., Ai4 keynote and 2026 interviews on maternal-instinct alignment; Thagard, P., critique, Psychology Today, 2025.
  • Yampolskiy, R., Unpredictability of AI (arXiv 1905.13053); AI: Unexplainable, Unpredictable, Uncontrollable (2024).
  • Karpathy, A., Dwarkesh Podcast, October 2025; LeCun, Y., public statements 2025–26.
  • Gebru, T. & Torres, É. P., The TESCREAL Bundle; Hanna, A. & Bender, E. M., Scientific American, 2023; Whittaker, M., Web Summit Lisbon, November 2025.
  • Hoes, E. & Gilardi, F., Existential risk narratives about AI do not distract from its immediate harms, PNAS, 17 April 2025 (N = 10,800, three preregistered experiments).
  • Chair’s Statement, 2026 World Artificial Intelligence Conference and High-Level Meeting on Global AI Governance, Shanghai, 17–20 July 2026; WAICO founding agreement, 16 July 2026, 29 signatories; China’s Global AI Governance Initiative, 2023; China position paper to the UN Global Mechanism on ICTs, July 2026.
  • Supervisory stack: SR 26-2 / OCC Bulletin 2026-13; PRA SS1/23; OSFI E-23; MAS AI Risk Management consultation, November 2025.

Provenance note on the seventy per cent figure. The widely circulated framing of Kokotajlo’s Diary of a CEO appearance as “a 70% chance AI leads to human extinction” is a compression introduced by the episode title and downstream coverage. In the interview he distinguishes explicitly between a 70% probability of AI takeover or comparable catastrophe — which could lead to extinction — and a 70% probability of extinction itself. Practitioners citing this figure to regulated audiences should cite the distinction. Where a reconstructed transcript rather than primary audio has been used as a source, that reconstruction should not be treated as verbatim.

Leave a comment