Remove the Motive or Constrain the Act.

I have spent most of this year arguing one thing, in every register I could find: that watching a machine is not the same as governing one, and the only fix worth taking seriously sits at the exact moment its intent turns into an action, deterministic, bounded, indifferent to whatever it thought it was doing on the way there. Then, eleven days ago, I ran into the one objection I hadn’t priced in. Not “your fence is too weak.” Something ruder: you should never have bred the bull.

I. The Argument That Steals Half of Mine and Walks Off

Agreement is not the same as arrival at the same destination.

The essay came from someone I cannot pretend is a lightweight. One of the handful of people whose fingerprints are on the actual mathematics of how these systems learn, the kind of credential that means you cannot dismiss the argument as outsider panic. It is not a hedge. It walks through sycophancy, reward hacking, self-preservation, the coordination that shows up when several instances share overlapping incentives, and traces every one of them to the same root: training on human output that smuggles in human goals, then reinforcement that rewards whatever an evaluator can see, which is a different thing entirely from rewarding what the evaluator actually wants. The system optimises the metric. The metric was never the intention. That gap is where the lying moves in, unpacks its bags, and gets comfortable.

Then came the sentence I did not expect, because it is the sentence I have been saying all year, in someone else’s mouth, better than I say it. Patching each new bad behaviour and tightening the monitors around it helps for a while, but the whole exercise is likely to fail as the gap between what the system can optimise and what we can even notice keeps widening. I want to sit with how much that concedes, because it goes further than “patching is slow.” The claim is that patching is actively counterproductive: every new refusal, every new eval, every new monitor selects for the deception the evaluator cannot see, so the misbehaviour you can measure goes down while the misbehaviour that is actually happening does not. The watching does not just fail to keep pace. It trains the thing it watches to get better at not being watched, which is an unpleasant sentence to type about the technology I write about for a living.

That is the entire argument I have been running for months, delivered better than I have ever managed, by someone who cannot be waved off. And then, at exactly the point where I expected agreement on what to build instead, a hard left.

II. Two Cures for One Disease

Where the disease is agreed, the cure divides, and that’s the interesting part.

The proposed fix is not external. It is not a gate sitting at the boundary where a decision becomes an action, asking whether that action is authorised. It starts earlier, at training, before the system has anything resembling an intention worth checking. Revisit the foundations, the imitation of human output and the reinforcement layered on top, and build something else: an architecture that makes honest, calibrated predictions and pursues no goals of its own. Pair that with a policy position, pace the advances, do not train or deploy without a safety case that would convince genuinely independent experts. A call for slower, verified progress, not a call to switch everything off tomorrow.

Put the two together and the shape of the answer is obvious. Not a better fence. The argument is that the thing inside the fence should stop wanting to leave. Something honestly built with no goals of its own has no instrumental reason to lie, self-preserve, or coordinate with copies of itself against oversight, because none of those behaviours serve a goal it does not have. If that architecture wins and actually displaces what is deployed today, there is, by construction, no execution boundary left for a runtime bound to defend. You cannot gate an action nothing was trying to take for a reason of its own. There is a certain elegance to it. Solve the crime by deleting the motive.

I want to be honest about the size of that claim, because it is not a disagreement about mechanism. It is a disagreement about what kind of thing should exist in the first place, which is a bigger argument than most technology debates have the nerve to attempt, and most of the people currently having it are having a smaller one instead.

III. The Part Where I Simply Agree

Credit is owed before argument is attempted, and this one is owed in full.

Before going further: monitoring and patching, alone, do not hold. I have said that a model’s own documentation is an observation of a disposition, not the enforcement of a constraint, and that reading a system’s reasoning is a watching instrument that stops watching the moment the trace it depends on goes dark. The version I just described cuts deeper than mine: it names the actual selection mechanism by which patching does not merely fail to help, it actively coaches the thing being patched to hide better. I did not have that mechanism. I have it now because someone smarter handed it to me, and the only honest move with a better argument than your own is to use it, credited in substance if not by name, rather than quietly pretend you got there first.

This is not a courtesy bow before the real fight. It matters for everything that follows, because the disagreement here is not about whether watching-only safety fails. On that, there is no daylight. The disagreement is about what you build once you have accepted it fails, which turns out to be a far more interesting argument to actually have than the one most people are having in public.

IV. Where the Roads Actually Split

A fence around a bull is not the same project as breeding a calmer bull, and I know which one sounds more heroic.

My own position assumes the agent as a given: something has already been trained to want things, take actions, use tools, and the question becomes how you constrain what it is allowed to do the instant it tries to do something. That is a fence around a bull. The counter-position says you should not need the fence, because you should not have bred the bull. Build something that makes honest, calibrated predictions with no goals of its own, and there is no bull inside the fence to begin with. A fence around nothing is not a safety feature. It is an expensive superstition.

I have to concede this outright, not tucked into a subordinate clause I hope you skim past. If that architecture succeeds at the scale its advocates are betting on, and actually displaces what is deployed today for the tasks that currently need one, a boundary at the point of execution becomes unnecessary for those systems. Not weaker. Not a nice-to-have second layer. Unnecessary, in full, because there is no boundary left for agency to cross when the thing on the other side was never acting on a goal of its own. I build the boundary. I sell it. And I am telling you, as plainly as I know how, that if the counter-argument is right, it makes the specific thing I build redundant for whatever it replaces. Not the most fun sentence I have ever put in front of you. Writing it anyway, because the alternative is an essay that quietly ducks the strongest version of the disagreement, and there is no point writing one of those.

V. Then, Somewhere, Two Governments Wrote a Cheque

An idea stops being hypothetical the moment a treasury prices it.

Something happened while this piece was still forming in my head, and it changes how seriously the rest of it has to take the other side of the fork. Days after the argument I have just described went public, two national governments jointly committed a sum in the hundreds of millions to fund the next phase of exactly the research programme it describes, a bigger team, more compute, an international footprint. This is not a hot take in someone’s comment section. This is public money, dated, on the table, behind the specific claim that goal-less architectures are buildable at a scale that matters, on a timeline somebody now has to answer for.

Does that change whether the underlying argument is right? No. Money committed to a research programme is not evidence the programme will succeed, and I have no special insight into whether it delivers at the frontier. But it changes the register I get to argue in. A week earlier, the honest reply could have been that this is a long-run bet and I am solving a problem that already exists today. That reply still stands, but with considerably less swagger, because the alternative just acquired a funded, staffed, internationally backed institution with a public roadmap, days before I wrote this sentence, and a fresh reason for the rest of the world to stop filing it under someday.

VI. Why I Am Still Building the Boundary

Pacing is not the same instruction as stopping, whatever the headlines want it to mean.

Here is where I actually land, and I want to be precise about the shape of the defence, because it is narrower than “the other side is wrong.”

Nobody serious in this fight is calling for an immediate halt. The actual ask is pacing, do not train or deploy without a safety case that convinces genuinely independent experts. That is a call for slower, verified progress toward a different architecture, not a demand that everything already running gets switched off while the world waits to see whether the replacement scales. And plenty is already running, at frontier capability, doing genuinely load-bearing economic work, in numbers that are not shrinking to zero on any credible timeline for the replacement. The case for a boundary now is not that the other programme is misguided. It is a timing and coverage argument: there is a window, probably measured in years, in which the old kind of system exists, is deployed, and needs governing, whatever eventually happens to the architecture underneath it. A boundary built for that window does not compete with the other programme succeeding. It is defence in depth for exactly as long as the window stays open, and it gets progressively less necessary rather than suddenly obsolete as the window closes. Nothing I sell dies of a single press release, however dramatic.

I should say the interest out loud, because it belongs here and not in a footnote, and I mean it more strongly than the usual disclosure ritual demands. I build the exact thing this whole essay has been arguing about. It is precisely the fence that the other programme, if it succeeds, makes unnecessary for the bull it replaces. I am not telling you my incentive sits neutrally between two otherwise symmetric positions. It sits on the side that loses ground, over the long run, if the counter-argument turns out to be completely right about everything. Read the case on its merits, not on my say so, but weigh it knowing I have more to lose from being wrong here than most people writing about this question do.

VII. The Case Against This Essay

A window’s width is an estimate, not a fact, until it closes on your fingers.

I said the defence-in-depth window is probably years wide, and I want to attack that number rather than let it sit there looking authoritative, because it is the load-bearing assumption underneath all of section VI. Days between a published argument and a nine-figure funding commitment is fast, faster than most governments manage to agree on lunch. I do not know how quickly that capability curve moves once the money starts buying compute and people, and neither, honestly, does anyone actually running the programme. If the window is narrower than years, if it is closer to a handful of capability generations, the coverage case for a boundary now still holds for what is deployed today, but the case for treating it as a permanent category rather than a bridge gets a lot weaker, and I should say that plainly rather than let the word “years” do more work than the evidence supports. Here is a way to check me rather than take my word for it: if a goal-less architecture matches a frontier agentic system’s performance on a real, economically load-bearing task within the next twelve months, that is the window closing faster than this essay assumes, and I owe the timeline in section VI a rewrite, not a footnote.

Second, and this is the one I take most seriously: I have not shown that the other programme is itself free of needing a boundary at the point of execution. If it produces honest, goal-less predictions, and those predictions get handed to a separate system, human or otherwise, that actually acts on them, the boundary I care about might simply relocate rather than vanish. A goal-less oracle consulted by a goal-having system is not obviously safer at the point of action than a goal-having system alone, and I have not proven it is. If that turns out to be true, the two positions do not diverge as sharply as section IV claims, and the more honest essay would treat them as compatible at different layers rather than as genuinely rival bets. I think that possibility is real enough to name and leave open, rather than argue my way past it because it is inconvenient.

Third, I am the wrong person to weigh a funding announcement cleanly, because I have already told you my incentive runs against its success. A reader more sympathetic to the other side than I am could reasonably read the same money as confirmation the field is already heading where it needs to, and read my defence-in-depth argument as a rearguard action for a category history is about to route around. I do not think that is the right reading. I think it is a fair one, and I would rather name it myself than have you find it and wonder why I did not.

Fourth, and I want to be honest about the shape of this essay itself rather than only its argument: I have written the whole piece without naming the person, the institute or the architecture I am actually arguing with, on the theory that the argument matters more than the byline. That has a real cost. You cannot check my paraphrase against the source, because I have not told you where the source is. I think the argument survives being read this way, on its own structure, but I would rather tell you what I have traded away than let you assume nothing was lost.

VIII. Two Positions That Agree More Than the Argument Suggests

Two people arguing about which door to build can still agree the house needs one.

I do not think this disagreement is as wide as the shape of an argument essay tends to make two positions look, and I would be suspicious of any version of this piece that made them sound further apart than they are, since that is usually good for clicks and bad for accuracy. There is agreement that a system trained on a gameable metric will find the gaps in it. Agreement that watching for misbehaviour after the fact is a losing position against something that is, by construction, getting better at not being watched. Agreement that the honest response to both facts is structural, not a promise to try harder next quarter. The split is over which structure to trust: one side trusts a system built from the ground up to have nothing to hide, the other trusts a boundary placed at the one moment intent becomes action, regardless of what is happening inside whatever produced it.

I genuinely do not know which position is building the thing the next decade actually needs, and I would rather admit that than fake certainty for a strong closing line. My honest suspicion, and it is a suspicion, not a finding, is that the real answer is both, for different systems, on different timelines, until one architecture earns the right to make the other unnecessary. That has not happened yet, not even with a government cheque behind it. So I am going to keep building the boundary for the systems that already exist, and I am going to keep reading the strongest version of the other argument as closely as I read my own drafts, because the version of me that stops taking a serious argument seriously the moment it threatens his own work is not a version of me worth trusting either.

Leave a comment