Start with a finding that is neither alarmist nor consoling, which is why I trust it. Earlier this year Max Tegmark and colleagues published a study of what they call scalable oversight, the practice of using weaker systems to supervise stronger ones (Engels et al., Scaling Laws for Scalable Oversight, arXiv:2504.18530, NeurIPS 2025). Their result is not that oversight fails. It is more precise, and more awkward. Oversight degrades as the gap in capability between the watched and the watcher widens, and stacking watchers on watchers to recover the loss buys less at each layer, until there is an optimal depth past which adding another tier makes things worse rather than better. Supervision has a ceiling, and the ceiling drops as the thing being supervised gets cleverer. This essay is about a structural reason for that, one I think most of the debate has walked past, and about the one kind of control that escapes it.
I. The Watchers’ Watchers
A measured result, and the quiet problem hiding inside it.
Two findings from the last stretch make the problem concrete rather than rhetorical.
The first is quantitative. Engels and Tegmark model oversight as a capability-mismatched game: a weaker overseer tries to catch a stronger adversary, and the arrangement is stacked, oversight of oversight of oversight. The numbers are sobering in a specific way. At a capability gap of four hundred Elo points, the overseer’s chance of winning ran from about nine per cent in one of their four games to about fifty-two per cent in another, and in three of the four it fell below fifteen per cent. Their headline structural result matters more than any single figure: there is a finite optimal number of oversight levels, and past it each added tier lowers overall success rather than raising it. Adding a watcher is not a free reduction in risk. It has a cost, and past a point the cost dominates.
The second finding is a confession. The UK AI Security Institute, writing up its control red-teaming, concedes that when it comes to trusting an automated monitor, “we so far have no way to establish trust” (Control Red Team blog, 23 July 2026). I want to hold that sentence at exactly its weight. The load-bearing part is that hedge, so far. This is a candid admission of a present methodological gap, not a verdict that monitors are untrustworthy in principle. A gap that might close. But note what kind of gap it is. It is not the gap you close by building a better monitor. It is the gap of not yet knowing how you would ever certify that a monitor deserves to be believed, and that is a different sort of problem, because it does not obviously shrink when the monitor improves.
The usual objection to behavioural monitoring is that catching a deceptive system is an intelligence race you can lose. True, and by now well worn; I will not re-run it. The objection I want to build is not about the race. It is about recursion, and it is why those two findings belong together.
II. Why Monitoring Is the Control That Recurses
The property is not difficulty. It is self-reference.
Most controls get harder as systems get more capable. Monitoring is worse than hard. It is self-referential, and the self-reference is the whole problem.
Let me make the structure explicit, because the precision is the point. Let g be the operation that maps a control to the control you must place above it to warrant your trust in the first. For a behavioural monitor M, applying g asks a simple question: why do you trust M’s judgement? If the only available answer is another evaluator, then g(M) is some M prime, itself a monitor, and the regress is the orbit of g: M, then g(M), then g(g(M)), and onward. The regress terminates exactly when g has a fixed point, a control c for which g(c) equals c, meaning the warrant for c is c itself and nothing further is owed.
Now the two findings from Section I attach to this map, and they are independent forces.
The first force is that every point in the orbit is untrusted for the same reason the first one was. This is the lesson of the AI Control programme (Greenblatt et al., AI Control, arXiv:2312.06942). Its protocols are built for exactly the world where the monitored model may be adversarial, and its premise is that a monitor’s own judgement is a surface that can be gamed rather than a foundation you can stand on. A monitor is not a neutral instrument. It is an agent, drawn from the same population as the thing it watches, built by the same methods, prone to the same failures. If M can be fooled or can collude, then M prime is introduced to catch a fooled-or-colluding M, and M prime inherits the same status. Nothing about adding a layer changes the kind of thing each layer is.
The second force is that the layers are not even equally weak. Engels and Tegmark’s result says the map degrades as it iterates across a capability gap. A weaker overseer watching a stronger agent is the normal case, not the pathological one, because the reason we deploy the strong agent is that it exceeds us. So each application of g is a step across a gap that widens as capability climbs, and the composition loses reliability rather than accumulating it. The finite optimal depth is the point where the loss overtakes the gain.
Here is the join that matters. The first force says no judgement-based monitor is ever a fixed point of g, because you can always ask why you trust its judgement and always get another monitor back. The second says you cannot outrun that by piling on layers, because the pile has an optimum and then declines. So within the space of judgement-based monitors, g has no fixed point, and the search for one does not merely take a long time. It does not converge.
I should be exact about the word, because the dramatic version of this claim is false and I will not lean on it. Calling this an infinite regress overstates it. Engels and Tegmark do not prove infinity; they prove degradation and a finite optimum. The correct statement is sharper and less theatrical: the regress does not run forever, but it never lands. It is truncated by cost, not resolved by success. A stack that stops because the next layer would make things worse has not found its ground. It has run out of budget while still standing on an agent whose judgement it cannot warrant. Compare a control that does sit still. A rate limit does not need a second rate limit to vouch for it. A signature check does not grow more doubtful when the thing it checks gets smarter; the maths is the maths. Those controls have fixed points. Monitoring does not, because monitoring is an act of judgement, and the moment you demand assurance about a judgement you have summoned another judge.
III. It Is Already Happening
Not a thought experiment; the field is living in it.
The reason I am confident this is not a chalkboard puzzle is that the practitioners have started saying it out loud, from more than one direction, and I want to be careful about how much weight each source can bear.
The clearest external framing I have comes through Concordia AI’s recap of this year’s World Artificial Intelligence Conference in Shanghai, and I am relaying it as secondary, eyewitness testimony, resting no argument on the exact wording. As Concordia AI reported it, Yang Xiaofang of Alibaba made the point that watching what an agent does stops being sufficient once agents start building other agents, because the agents they build would themselves need watching. At the same forum, again via Concordia AI, Zhou Bowen of the Shanghai AI Laboratory is reported to have argued that safety cannot be established once and inherited, but has to be re-proven with each new generation of system. I cannot vouch for either man’s precise phrasing, no public transcript or recording surfaced, and I will not pretend otherwise. But two practitioners, reported converging on the claim that per-behaviour and per-generation assurance runs out exactly when agents produce agents, is at least a signal that the regress is being felt in the room and not only derived on paper.
Behind this sits a fuller taxonomy I have located but not read, a framework from Concordia AI and collaborators grading agent autonomy from assisted through to fully independent, an L1 to L5 scheme. The document is in Chinese and I have not read it, so I place it only as a marker, not as evidence: the people closest to the deployment frontier are already reasoning in terms of ascending autonomy, and ascending autonomy is exactly the axis along which the gap between watcher and watched widens. I flag the paper so a reader who wants the primary can go and get it, and so I am not quietly implying that I have.
The Western labs are circling the same drain in more guarded language. The UK AISI’s hedge is one instance. The AI Control literature’s insistence on protocols that survive an untrusted monitor is another. Nobody serious is announcing that monitoring has failed. Everybody serious is building in the assumption that the monitor’s word cannot be the last word. That shared assumption is the regress, observed from inside the institutions that would most like it not to be there.
IV. Where the Regress Bottoms Out, and Where It Doesn’t
Some orbits do land. The question is which space you are searching.
I do not want to smuggle in a result by choosing the space conveniently, so let me give the strongest version of the opposing case before I answer it.
The AI Control programme has a genuine anchor, and it is worth stating with its real numbers. In their setting, a stronger untrusted model writing code is policed by a weaker trusted model and a human audit budget of about two per cent of cases. Their best protocol reached roughly ninety-two per cent safety, measured as the chance no backdoor slips through, while keeping about ninety-four per cent of the useful work. That is real safety extracted from an explicitly untrusted system, and it is a fixed point of a sort: a place the recursion is allowed to stop because you have decided, by fiat backed by evidence, to trust something, a weaker model and a slice of human attention. I concede without reservation that in a bounded domain this can suffice. If the trusted anchor is genuinely reliable for the tasks in scope, the stack above it can be short and effective, and the regress bottoms out on the anchor.
The concession has a boundary, and the boundary is the whole of my disagreement. The anchor works as a fixed point only where its reliability is not itself in question, which means where the capability gap is small and the domain is narrow. That is the regime Engels and Tegmark’s optimum describes: a few layers, a modest gap, an anchor you can still vouch for. Push the capability of the monitored system up, widen the gap, and the trusted anchor becomes the weakest thing in the room. A trusted weak model overseeing a much stronger agent is trusted precisely where it is least able to see. The anchor does not vanish, but its warrant thins exactly as the stakes rise, and at the limit you are trusting it because you must, not because you can.
There is a second concession I owe before the payoff, and it is the one that most constrains what I am allowed to conclude. Neither primary I have actually read terminates the regress the way I am about to. Both terminate it with another evaluator. AI Control anchors on a trusted weaker model plus human audit; the oversight-scaling work grounds out in a finite optimal depth of more oversight. When the surveyed literature stops the staircase, it stops it with a judge it has decided to trust. The terminator I am going to argue for, a capability bound rather than a better judge, is my synthesis. It is not a result I can cite from these papers. I am defending it, not quoting it, and I want that line drawn before I cross it.
V. The Fixed Point Outside the Game
Declare the interest, then state the result: a bound that never asks what the agent intends.
Let me declare the interest first, because you should read what follows knowing it. I build enforcement kernels. The terminator I am about to defend is, in essence, the deterministic capability gate my company sells. I am not a disinterested narrator here. Discount accordingly, and then judge the logic on its own terms, which is all I would ask of any argued case.
If no judgement-based monitor is a fixed point of g, the way out is not a better monitor. It is a control that g cannot iterate on, because there is no judgement in it to question. Consider a control that does not evaluate behaviour at all. It does not ask what the agent intends, whether the plan is benign, or whether the output looks safe. It checks one thing: is this action a member of a permitted set. When an agent tries to act on the world, it does not act directly; it requests a specific call to a specific tool, with specific arguments, in a specific context. At that boundary the control reads the request against a fixed specification of what is allowed, and admits it or refuses it. There is no assessment of purpose anywhere in the mechanism.
Apply g to this control and watch it stall. The question g asks is: why do you trust its judgement? The answer is that it makes no judgement. There is nothing to fool, because the gate is not forming a belief about the agent that could be false. There is nothing to re-evaluate, because the decision is a deterministic function of the action and the permitted set, not an opinion a higher layer would need to audit. To inspect the control you read the permitted set, and the set is a static object written by a person, not an agent whose reliability degrades across a capability gap. So g(c) returns c: the warrant for the gate is the gate and its set, and nothing above it is required. That is what a fixed point looks like, and it sits outside the space of monitors entirely, because it is not playing the observation game. It never observes intent. It bounds capability.
That invariance is what makes it hold under recursion. Put an agent behind it, or an agent that builds agents behind it, or a tower of agents building agents; the bound at the tool-call boundary sees the same thing each time, a request to be checked against a specification, and answers it the same way. Agents building agents change everything on the intent side of the boundary and nothing on the action side, and the bound only ever looked at the action side. There is a real instance of the shape, which is why I take it seriously rather than as a wish. The CaMeL design pattern enforces what an agent may do at the point where actions are dispatched, on the basis of provenance and a capability policy rather than a reading of the agent’s intent, and its guarantee holds regardless of what the model does, because the enforcement does not depend on the model’s cooperation. Whatever an injected instruction persuades the model to want, the dispatch either sits inside the permitted capability or it does not.
That is the whole of the positive claim, and I am keeping it to one section on purpose, because it is not the novel part of this essay and pretending otherwise would be cheating. The recursion is the contribution. The fixed point is the resolution. A durable control is one you can stand behind without needing to trust the thing in front of it, and the only controls with that shape are the ones that decline to observe.
VI. The Case Against This Essay
The strongest objections, including the one that keeps me honest.
I owe the argument its best rebuttals, so here they are without softening.
The first is the terminology, and I accept the constraint it imposes. If a reader hears the phrase infinite regress and I deliver a finite optimal depth, I have oversold. The correct claim is the weaker and truer one: the stack is truncated by rising cost, not completed by success, and it stands on an unwarranted judgement when it stops. If that weaker claim does not move you, the essay does not get to fall back on the dramatic version.
The second is the anchor objection, which I conceded in Section IV and now restate as an opponent would. AI Control shows that judgement-based oversight with a trusted component delivers real safety in a bounded threat model. If your domain is narrow and your capability gap is modest, my whole argument may be answering a question you do not have. The regress lands on your anchor and you are done. I claim only that the anchor thins at the frontier; I do not claim it is worthless, and a critic is entitled to say most deployments are not at the frontier. There is a stronger form of the same objection: you may not need a fixed point at all. If stacked oversight plus a human audit drives the probability of harm low enough for the stakes, then adequacy, not a resting point, is the real engineering target, and the finite optimal depth is precisely the best reduction on offer. My reply is that at the frontier, low enough stops being available, but that is a claim about stakes, not a refutation of the method.
The third is provenance. My cleanest illustration of the deployed nesting pattern, the Yang Xiaofang and Zhou Bowen remarks, reaches me through Concordia AI’s WAIC recap, a secondary source, and the underlying L1 to L5 document is one I have located but not read. If that ladder is not what I take it to be, Section III loses its named example, though not its logic, which rests on the primaries.
The fourth is the most serious, and it is the boundary of my own fixed point, which I cross-reference rather than dodge. The capability bound is a fixed point under recursion: it does not spawn a monitor above it, because it makes no judgement to question. It is not a fixed point under coverage. If a harmful outcome can be reached through a sequence of individually permitted calls, or through a tool nobody thought to place behind the boundary, or through orchestration that drifts outside the mediated path, the bound is silent, not because its judgement failed but because it was never asked. Everything I have claimed is that the boundary is invariant to what the agent intends. None of it claims the boundary is complete. Recursion-hardness and coverage-completeness are different virtues, and I have argued for exactly one of them. A serious deployment needs both, and the second is a genuine, unsolved problem that no amount of the first will discharge.
The fifth objection is the sharpest, and my own recent work hands it to me. I have said the permitted set is a static object written by a person, and that this is what lets the gate escape the regress. But who writes the set, and does that not recurse? If the set must be authored and re-authored to keep pace with agents that build agents, and if the only thing fast enough to write it is another model, then I have pushed the judgement I evicted from the gate back into the thing that configures the gate. The regress of watchers is terminated; a regress of set-authors may open in its place. I do not think it is quite the same regress, because a policy can be reviewed at rest, offline, by many eyes and against a fixed target, in a way a runtime judgement cannot. But the escape is clean only if the set can be written and checked more slowly than the agents move, and that is not guaranteed. Where it fails, what I have is a fixed point for the enforcement and not for the specification, which is a narrower prize than the framing implies.
The sixth is the interest, once more and in the open. I sell the terminator. That does not make the fixed-point argument wrong, but it gives me a motive to find it compelling, and a careful reader should discount accordingly and check the structure without me. I have tried to make the case against my own product at full strength precisely because the alternative is to be believed for the wrong reasons.
VII. Coda
Where the recursion actually stops.
The shape of the whole thing is simple once the object is in view. Behavioural monitoring is the one control whose cost rises with recursion, because every watcher is itself a watchable agent drawn from the same population it must police. You cannot escape that by adding watchers, because the map from a monitor to its warrant has no fixed point among monitors, and because the stack degrades across the capability gap and has a finite best depth after which it worsens. The people building these stacks say, in their own documents, that they cannot yet establish the trust the stacks assume.
A regress ends one of two ways: by landing on a point it can rest on, or by running out of the resource pushing it forward. Judgement-based oversight ends the second way at the frontier. It stops because the next layer costs more than it returns, still standing on an agent it cannot vouch for. The only way to end it the first way is to step off the observation game altogether, onto a control that asks nothing about intent and so offers nothing to game. Check membership in a permitted set. Admit or refuse. Read the set, not the mind.
That is the durable ground, and it is deliberately unglamorous. It does not understand the agent. It bounds what the agent can reach. Under agents building agents, the watchers multiply and each new one needs watching; the bound does not, because there is no belief inside it to be wrong. I have a commercial reason to want that to be true. I have tried to give you a structural reason that does not depend on my wanting it. The monitor needs monitoring, and that is not a flaw you fix by adding monitors. It is a reason to build at least one control that does not need to be believed.
Go deeper. The paid companion, The Monitor: The Dossier, lays out the oversight-scaling numbers in full, formalises the regress as an operator, prices the rival trusted-anchor terminator protocol by protocol, and adds a recursion-property scorecard, each to a named source with its confidence marked.


Leave a comment