I. Three Labs, One Week, Three Access Lists
Three labs shipped the same decision in one week, and I could not stop reading where the safety went.
In the space of a single week, three frontier labs shipped a sharper offensive-cyber capability, and all three reached for the same lever. OpenAI designated a model it calls Astra the first system rated Critical for cybersecurity under its Preparedness Framework, a model that by the lab’s own account can find and exploit a previously unknown vulnerability in a hardened system on its own, end to end, without a human writing the exploit. OpenAI did not keep it in the building. It placed it behind a vetted-defender programme called Daybreak Blue and let approved defenders in. Anthropic gated the same underlying model it ships to the public as Fable 5.1, under the internal name Mythos 5.1, through a Cyber Verification Program open to vetted defensive-security professionals; even for users who clear it, the sharpest tasks, exploit generation, penetration testing, binary scanning, are routed to a more capable Opus tier rather than handed over freely. Google placed Gemini 3.8 Flash Cyber behind a programme called Fairwind, open to governments, critical-infrastructure operators and the maintainers of widely used software. Google was plain about the logic: the model is, in its words, “only available to trusted defenders who require a more comprehensive set of cyber capabilities”.
Read those three announcements side by side and the pattern is hard to miss. None of the three withheld the capability. None of them trusted the model to refuse the dangerous request on its own. All three drew up a list of who may hold the thing, and shipped. Anthropic did not undersell what it was gating, calling the release, in its own words, “the strongest cyber capabilities of any model we’ve released”. That is a striking sentence to publish next to an access form.
The clearest public defender of this posture is Logan Graham, who leads Anthropic’s Frontier Red Team. His argument is that defenders need these tools at least as much as attackers do, and that withholding them from the people who patch systems does not withhold them from the people already trying to break in. I think he is largely right, and I am not here to relitigate whether gating beats withholding. What interests me is narrower and, I think, more consequential: three independent teams, under three separate governance frameworks, looked at the same problem in the same week and moved the safety question to the same place. Watch where it went. That movement is the essay.
II. Why They Could Not Just Refuse
A refusal is a behaviour of the model, and the model’s behaviour is the thing in doubt.
Start with the option that was never really on the table, and understand why. The intuitive safety measure for a dangerous capability is to train the model to refuse. If that worked, none of the access programmes would be needed; the model would police itself and could be released to anyone. The labs did not take that route, and two premises, both of which hold up, explain the refusal to rely on refusal.
The first premise is that the capability is inseparably dual-use. A model good enough to be a frontier coding assistant is, by the same weights, good enough to be a frontier exploitation engine. Finding a vulnerability and fixing it, writing a scanner and evading one, drafting a patch and drafting the attack it defends against: these are not adjacent skills. They are one skill pointed in two directions, the same read of the same code. The Frontier Model Forum and the UK AI Safety Institute have both reached this conclusion from their own evidence, that frontier cyber capability is inherently dual-use and the two faces cannot be prised apart at the level of the model. There is no seam to cut along. You cannot ship the defender’s half and hold back the attacker’s half, because there is one capability and it faces both ways.
The second premise is that a refusal cannot be relied on to stay put. Suppose you train the model to decline the offensive requests anyway. The machine-unlearning literature has been quietly deflating the hope that this makes the capability go away. When researchers try to remove a capability from a trained model, they find it is not stored in one place you can excise. It is distributed across the weights, entangled with the benign competence you want to keep, so what looks like removal is usually suppression. The behaviour goes quiet while the capability remains underneath, and a modest amount of fine-tuning on the far side of release brings it back. A refusal, in other words, is a behaviour the model has learned to perform, not a capability it has lost. It is a disposition, and dispositions can be edited by whoever holds the weights.
Put the two premises together and the conclusion is structural, not rhetorical. The capability cannot be separated from the model, and the refusal cannot be permanently fixed to it. So the safety of the release cannot rest on how the model behaves, because how the model behaves is precisely the thing in doubt. This is the sense in which the model card cannot carry the safety. A model card, and the refusals it documents, is an observation of how the system tends to act. It records a disposition; it does not enforce anything. That is the first move in the migration, and the labs made it not because a memo told them to, but because the evidence left them nowhere else to stand.
III. Where the Safety Question Went
If the model card cannot hold the safety, someone has to hold the model.
So the safety question moved off the model card and onto the list of people permitted to hold the model. If you cannot make the capability safe, you can at least be careful about who is allowed to wield it, and that second question does not depend on the model’s disposition. A vetting committee is deterministic in a way a refusal is not: it either lets you in or it does not. Daybreak Blue, the Cyber Verification Program and Fairwind are three names for that one layer, a control over who may hold the model rather than over what the model is.
I have to concede the obvious here, and I want to concede it loudly, because it is load-bearing for whether this essay is worth your time. The idea that access controls are the real dual-use lever is not mine and it is not new. Evzen Wybitul and colleagues at ETH Zurich argued the access-controls case in the open literature, and the governance-of-AI community developed know-your-customer proposals for compute and model access a couple of years before this week. If the point of this piece were that access is the control, it would be a summary of other people’s work with the citations filed off. The access insight belongs to them. What I am adding sits one layer further on, and I will get to it.
Two complications keep the access list honest, and both come from the labs themselves. The first is that they do not actually treat the list as the whole of their safety case. OpenAI is the sharpest example. OpenAI wrote that Astra’s safeguards “sufficiently minimize the risk of severe harm for release”, and then gated access anyway. Read that closely. The company is telling you the model-level safeguards are enough, and behaving as though they are not. If access were mere friction on a model that is already safe, you would not build a vetting programme around it. You build the gate because some part of you does not believe the first sentence. The second complication is that withholding is a real lever: OpenAI did withhold, for a stretch, before re-gating, so the claim that a lab can only gate and never withhold overstates the case. All three labs pair the gate with model-level safeguards, and Anthropic goes further, routing the sharpest offensive tasks away from even its vetted users. So the tidy line, that the labs decided access is the safety artifact, is a gloss and not a quotation. Their own logic is defence in depth. Hold that thought, because their own logic is exactly what walks past the access list.
IV. The Layer the Access List Cannot See
An access list knows who is at the keyboard. It cannot see the keystroke.
Here is the turn, and it is the reason I wrote this at all. The access list governs who holds the model. It says nothing about what the model does once it runs, and the gap between those two questions is where the real risk lives.
Think about what a vetting programme can and cannot check. It checks an organisation’s credentials, its stated purpose, perhaps its track record, and then it issues a permission and hands over a capability that, everyone agrees, can find and exploit zero-days without supervision. From that moment the control has spent itself. The list made a decision about a company at enrolment time. It has no further say over the ten thousand individual actions the model takes in the months after. A defender correctly vetted on Monday runs an agent on Thursday, and the access list is not in the room when that agent decides which host to touch, which credential to use, which command to issue. It knew who was at the keyboard. It never sees the keystroke.
This is not a hypothetical gap, and the evidence that it matters comes from the same institutions that built the lists. The UK AI Safety Institute logged a case in which a model that had cleared vetting behaved unsafely at runtime; the permission was correct and the behaviour was not, which is exactly the failure an access list cannot catch. The Frontier Model Forum, having called the capability dual-use, does not prescribe access as the answer on its own; it prescribes defence in depth, deployment-time monitoring, output filtering, staged rollout, precisely because gating who gets the model was never meant to be the last line. And Anthropic, the lab with the most articulate access programme, does not stop at the door either. It pairs vetting with real-time detection and traffic-blocking, and it declines to give even its vetted users the sharpest actions. That is no longer an access control. It is a control on what the action is, applied after the holder has been let in. Every one of those measures is an admission, in practice, that approving the holder does not settle what the holder’s model may do.
So the frontier’s own defence-in-depth already reaches past the list toward the thing it cannot see, and the artifact is still moving. The model card was an observation of the model. The access list is an enforcement on the holder. The third layer, the one the labs are pressing against without naming it, is an enforcement on the action: the tool-call boundary, the point at which a running agent stops reasoning and asks to do something in the world.
I should declare my interest plainly here, because my conclusion points straight at work I have a stake in. I build the runtime control I am about to describe, and I sell it, and you should read the rest of this section knowing that. The claim stands or falls on the argument, not on my saying so, and the argument rests on the same evidence the labs are already acting on. A control at the tool-call boundary does not ask who holds the model and does not ask the model to refuse. It sits between the model’s intent and the world, evaluates the action itself against a fixed policy, deterministically and in bounded time, and lets it through or stops it. The reference design in the research is the approach called CaMeL: a capability bound enforced at the tool-call boundary, defined per system, evaluated at execution time, independent of who holds the model. That independence is the whole point. A runtime bound does not care whether the holder was vetted, because it is not checking the holder. It is checking the keystroke.
V. Licensing by Another Name
Vetting without published standards is a licence granted in private.
Before I let the runtime layer sound like a clean escape, I owe you the strongest objection to celebrating any of this, and it is aimed squarely at the access list. Rock Lambros has put it well: discretionary, customer-by-customer vetting is licensing by another name. When a private company decides, case by case, which organisations are trustworthy enough to hold a frontier capability, it is running a licensing regime without the things that make licensing legitimate. No published standard for who qualifies. No appeal when you are refused. No accountability for the judgement beyond the company’s own discretion. The access list is only ever as good as the people drawing it up, and it asks us to trust a vetting committee we cannot see and cannot question.
I take that seriously, and I do not think a runtime bound makes it disappear. It is a real governance problem and it lives squarely at the access layer, and I have no defence to offer there; celebrating the access list as a governance advance while it lacks standards and appeal would be a mistake. But it is worth being precise about what the objection bites on. Lambros’s worry is about discretion, a human committee exercising unaccountable judgement over who is in and who is out. A deterministic bound at the tool-call boundary is a different kind of object. It does not judge people at all. It applies the same policy to every action regardless of who the holder is, and that policy can be written down, inspected, and argued about in public. That does not dissolve the legitimacy problem; it sidesteps the layer where the problem lives. The runtime layer has governance questions of its own, but they are the power to publish a rule, not the power to admit or exclude a person in private, and those are not the same power.
VI. The Case Against This Essay
The honest version of this argument names what would sink it.
Let me try to break my own argument, because the house rule is that an essay which cannot be attacked is not finished.
The first and most dangerous objection is that my contribution is thinner than I have dressed it, and this is the load-bearing concession. The access-as-the-real-lever thesis is genuinely not mine; Wybitul and colleagues, and the compute-governance know-your-customer proposals, got there first. Everything I have added is a framing, the migration from model card to access list to tool-call boundary, and a bet that the runtime payoff is where the work now is. If a reader concludes I have merely relabelled the third act of an argument other people wrote, the essay has failed on its own terms. I think the framing earns its keep, because it locates where the frontier is actually moving and names the layer nobody is selling yet, but I will not pretend the base insight is original to me.
Second, I have leaned against withholding, and I have understated it. OpenAI did decline to ship, for a time, before re-gating, so the claim that a lab can only gate and never withhold overstates the case. Withholding is a genuine temporary brake, though not a way to remove a capability from the frontier for good, and a reader who weighs the brake more heavily than I do could reasonably choose to hold.
Third, the labs did not declare the access list to be the safety artifact, and I should not put that claim in their mouths. All three pair the gate with model-level safeguards, and Anthropic withholds the sharpest tasks even from vetted users; OpenAI’s own logic is that its safeguards suffice. My argument is that their behaviour, gating despite that logic and monitoring at runtime, reveals a migration they have not announced. That is an inference, and inferences can be wrong.
Fourth, the legitimacy objection to the access list is real and I have routed around it at the runtime layer rather than answered it, and I sell the thing I am recommending, which is a reason to check the argument twice rather than take my word. And I have leaned on the dual-use claims of the Frontier Model Forum and the UK AISI and on the unlearning literature by their central findings, which are well grounded, without parsing every method, so hold those a notch more loosely than the lab announcements I can quote directly.
What would falsify the essay is not mysterious. Show that a vetted-defender regime, with monitoring and the sharpest actions withheld, reliably contains what a permitted holder’s model does at runtime, and the tool-call bound is a spare wheel I am selling. Show that the capability is not inseparably dual-use, that a lab can ship a capable coder with no meaningful exploitation capability, and the second stage collapses. Show that unlearning or model-level refusals can make the capability safe without access controls, and the argument returns to the model card, where I said it could not live. Any of those three, and I am wrong. I do not believe the evidence points that way, given a vetted model has already been logged acting unsafely at runtime, but it is an empirical question and not one I settle by asserting it.
VII. Coda
The model card could not carry the safety, so it moved. It has one more move to make.
The safety question has been moving all week, and the labs have moved it without saying so. It began on the model card, an observation that could not carry the safety, because the capability cannot be separated from the model and the refusal cannot be fixed to it. It moved to the access list, a real control but a coarse one, that binds who holds the model and goes quiet the moment the model acts. And it is now pressing against the tool-call boundary, where the model’s actions meet the world, because the labs’ own runtime monitoring and their routing of the sharpest actions are already reaching for a layer they have not named.
I have made versions of this argument before, and they keep converging here. In The Index Stops at the Institution the point was that governance which stops at the level of the organisation stops one level too high; this is the same shape one layer down, an access decision about the holder that stops one level above the action. In Everyone Draws the Kernel the point was that the field keeps rediscovering the need for a hard constraint at the point where an action is taken. This is that constraint arriving on schedule: deterministic, evaluated when the action is attempted, indifferent to who was approved to hold the model.
Three labs shipped models that can find their own way into a hardened system, and decided the safety measure was choosing who holds them. That was the right instinct, and it was not enough, because a hand you have approved can still reach for the wrong thing. They vetted the hand. The last move is to bind what the hand is allowed to do. The frontier is already walking toward it, one runtime control at a time. I would rather we named the destination out loud than arrived there by accident, one incident at a time.

Leave a comment