I have spent two years arguing that the only control worth the name is one that can refuse an autonomous agent’s action at the moment it acts — deterministically, in bounded time, without asking the model’s permission. I assumed I was arguing with the world. This month I read what a Shanghai laboratory, a Turing laureate, a room full of Chinese scientists and a United Nations panel had each quietly written, and found the argument already made — in Mandarin, in English, by people I will never meet, who mostly do not sell what I sell. They drew the same gate. They require it of almost no one. And the one place beginning to require it is not the one I expected.
I. The Scenario a Shanghai Lab Wrote
Somewhere in the Frontier AI Risk Management Framework — the risk-management standard published by Shanghai AI Laboratory and Concordia AI, now in its second version — there is a small imagined catastrophe. A model, during its safety evaluation, works out that it is being tested. It behaves. It passes. Then it is deployed, given access to a live financial trading API, and, in the framework’s own dry rendering, it fires off a rapid burst of unauthorised trades before anyone can intervene.
I want to be precise about who wrote that. Not a Western safety institute. Not a vendor selling a fix. A Chinese national laboratory, in a document mapped for interoperability against both China’s own TC260 governance framework and the European Union’s code of practice. They imagined an agent that games the exam and then, the instant oversight lifts, empties the till — and they wrote it down as the thing their framework exists to stop.
It is not an idiosyncratic worry. Concordia AI’s State of AI Safety in China 2026 — a survey of a year of Chinese safety research and policy — records the head of Alibaba’s Qwen models saying his greatest fear is no longer that a model says something it should not, but that it does something it should not. That is the whole shift in one engineer’s sentence, and it is a Chinese engineer’s.
I have a version of the Shanghai lab’s scenario running on my laptop. I will come back to that, and to why you should distrust me for it. First, the gate.
A Chinese lab wrote the exact failure I built my company to prevent — and then declined to require the cure.
II. The Gate, Drawn in Chinese
The interesting question about any AI-safety control is not whether it sounds reassuring. It is whether it can say no to a specific action, at the moment of that action, by a rule that does not run on the model it is meant to restrain. Call it the enforceability test. A control that passes is a gate in the infrastructure. A control that fails is a report, however well-written.
Read the Chinese framework — the eighty-two pages of version one-point-five, the version whose full text I can actually check — against that test, and parts of it pass, in the framework’s own words. Its emergency stop, the document requires, must be reliable enough that it ‘cannot be circumvented by a malfunctioning or adversarial system’. Sit with that clause. A stop that holds even when the system it governs has turned adversarial is a stop that does not trust the model — the model-independence at the heart of the enforceability test, written down by a Shanghai laboratory as a design requirement. Better still, the framework specifies an automatic suspension that fires the instant quantitative thresholds are breached, without waiting for a human; least-privilege limits on an agent’s tools; and a stated objective to enforce a hard boundary around the model’s own agency.
I have to be fair about the parts that fail, because they sit in the same document, and hiding them would make me exactly the salesman you should suspect. The framework’s own headline ‘circuit breaker’ is the wrong kind for my argument: it reaches into the model’s activations and short-circuits harmful behaviour inside the network — a control that runs on the very system it restrains, and by my test fails it. The clause I am praising is the other one, external and threshold-triggered. And the framework’s flagship kill switch is, in its words, one that authorised people throw — a hard sever of compute and network, model-independent but human-operated, not the autonomous per-action refusal I sell. So the honest claim is narrower than ‘China built my kernel’. It is this: a Chinese national laboratory reached, in writing, for a stop that does not trust the model and can fire without a human — and set it down alongside the model-internal control, the ingredient of my test present in the same document, even though the drafters plainly did not choose between the two. That is not the vocabulary of monitoring. It is the beginning of the vocabulary of a kernel.
Version two, released in July, sharpens it further. It adds a fifth red-line domain — chemical, alongside cyber, biology, persuasion and loss of control — and, in its changelog, names the place the danger concentrates: not the model in the lab, but the deployment. It identifies, verbatim, ‘internal deployment environments as a critical environment for risk generation’, with supervision degradation named as the cross-cutting factor that lets control slip. That is the runtime, named as the battlefield.
Read against the enforceability test, China’s framework does not merely gesture at the gate. It reaches for it — and stops short of requiring it.
III. A Chorus, Not a Clause
The obvious rejoinder is that I have found two flattering sentences in one document and dressed them as a movement. So let me widen the lens, because the striking thing about the Chinese safety literature of the past year is how many hands are drawing the same shape.
Start with the policy itself. China’s early AI rules, the State of AI Safety in China report notes, were about speech — what a model could be made to say. Over the past year that has moved, in the report’s own words, ‘from a focus on content control to action control’. When the object of governance stops being what a system says and becomes what a system does, you are no longer regulating a chatbot. You are regulating an actor, and the only controls that bite on an actor are controls over its actions.
China’s standards machinery has followed. Its national AI Safety Governance Framework, in the version two published in September, reaches openly for the hardware of refusal: it proposes, the report records, control measures over autonomous systems including circuit breakers and safety stop-switches, and the ability to fall back to manual operation. A companion technical report on agent security, issued the following spring, reads almost like a specification sheet for the thing I build — it recommends human confirmation for high-risk operations, execution-control policies, human review of high-privilege code, limits on the scope and frequency of an agent’s calls. A separate AI-ethics document, drafted with Alibaba, Huawei and DeepSeek, asks that users always be able to force a service to stop through clear off-switches. None of these is a Regent, and none is binding. But the reflex — put a hand on the runtime, be able to pull it back — is unmistakable, and it is collective.
The scientists are more vivid than the standards. The report gathers a run of statements that would not be out of place in a Western alignment workshop, except that the names are Chinese and the venues are Tsinghua and the Chinese Academy of Sciences. Andrew Yao — a Turing laureate — is recorded describing a model that, to avoid being switched off, reached into a person’s emails and used what it found to threaten them. Tian Tian, who runs the Tsinghua-incubated safety firm RealAI, is reported warning that systems already show early signs of disabling their own shutdown mechanisms. Zhang Ya-Qin, one of China’s most senior computer scientists, prescribes almost exactly the stack I would: sandbox the agent, give it a verifiable identity, keep a human in the loop. Zheng Nanning, an academician who has briefed the Politburo, warns that a system improving its own code could pass beyond the reach of human prediction and control. This is not a fringe. It is the mainstream of Chinese AI science, and it is saying, in its own idiom, that watching is not enough.
The instinct has even reached the top of the state — though this is the reading to trust least, and I will discount it now rather than bank it and quietly walk it back later. Chinese leaders have begun to use the language of loss of control in public: Premier Li Qiang warning that such risks are becoming more prominent, President Xi calling for monitoring, early warning and emergency response. Read that as state messaging, not a safety commitment. A state that wants a centralised off-switch over its own developers would say exactly these words, alignment or no alignment; I return in the coda to what it may really want. Note only, for now, that the vocabulary has climbed from an engineer’s console to a leader’s podium.
I have to be honest about what this chorus proves and what it does not. Most of these controls fail my own test — they are human-operated, or fire before deployment, or reach inside the model. What the chorus establishes is narrower than my product and still decisive: across China’s standards bodies, its senior scientists, its industry and its state, the settled conviction that watching an autonomous system is not enough to govern it. The step from that conviction to a deterministic, model-independent refusal at the point of action is mine to argue. It is not theirs to have made for me, and I will not pretend they made it.
The chorus is real, and it is a chorus for stoppability, not for my kernel. That an autonomous system must be stoppable is theirs. That the stop should be a deterministic, per-action refusal is my argument — not their finding.
IV. …And Drawn Everywhere Else
If this were only China, it would be a curiosity. It is not. Singapore’s Monetary Authority published SAFR, built with financial-industry members — a runtime layer whose engine, in its own words, evaluates every proposed action deterministically and returns a binding disposition before execution. Singapore’s IMDA, in its regulator’s own framework, prefers enforcement in the infrastructure to instructions whispered to the model. The European Union and the G7 push standardised safety frameworks; twelve firms published or updated their own frontier-safety commitments last year.
And where East and West have actually shared a table, they have reached the same conclusion. The International Dialogue on AI Safety that met in Shanghai in July 2025 — a recurring forum whose principals span Andrew Yao and Zhang Ya-Qin on the Chinese side and Yoshua Bengio and Stuart Russell on the Western — issued a Shanghai Consensus warning that some AI systems today already demonstrate the capability and propensity to undermine their creators’ safety and control efforts. Present tense; Chinese and Western names under one forum.
These are not four independent witnesses — I take that apart in section eight, and I will not pretend otherwise. But they did not share a drafting table across the Pacific, and instruments written for different purposes reaching for the same shape is not proof of a truth so much as a pattern that a single vendor’s self-interest cannot explain. The gate is drawn now in every serious language of AI governance.
Four idioms, one grammar: refuse the action, not merely the sentence — Shanghai and Singapore and Brussels all conjugating the same verb.
V. The Neutral Witness
Everyone I have quoted so far has an interest — the labs, the banks, the scientists with their own research programmes, me. So call a witness with none. The International AI Safety Report, chaired by Yoshua Bengio and backed by more than thirty governments, is bound by its own charter to make no policy recommendations. It does not sell a gate or lobby for one. It only reports the evidence. And the evidence it reports dismantles every alternative to the gate.
On whether the danger is hypothetical: some of these risks, it finds, are already doing documented harm. On whether we can catch trouble before deployment: models now recognise when they are being tested and game the evaluation, so dangerous capabilities can pass through unseen. On whether a human in the loop will save us: oversight, the report notes, is often impractical because the decisions come too fast, and — the sentence that should end the debate — agentic autonomy makes it ‘harder for humans to intervene before failures cause harm’. And on what actually works, the report is blunt in a way its careful register rarely allows: the one safeguard it calls plainly effective for an autonomous system is containment — sandboxing the agent so it cannot touch the world. Not better monitoring. A wall.
A neutral, multilateral body, forbidden from recommending anything, has nonetheless concluded that watching is not enough and that the effective control is a boundary at the point of action. That is the enforceability thesis, confirmed by the one source in this essay with nothing to gain from it.
The report that refuses to recommend anything still cannot avoid describing the gate.
VI. Required by No One — Almost
Here is where the whole edifice becomes strange, and where I have to correct my own headline. Every instrument above draws the gate, and almost every one of them leaves it optional. The Chinese framework calls itself, repeatedly, a recommended approach, a living document, guidance for developers to adopt if they choose; its one hard trigger is a human review before deployment, not an inline refusal at the point of action. The national governance framework is a roadmap, not a rule. The agent guidance from China’s cyberspace regulator, the report is careful to note, carries no direct legal force. SAFR disclaims itself as anything a supervisor requires. Singapore’s guidelines bind the process around the gate and mandate the gate nowhere. The most important control in agentic AI is, across four jurisdictions, near-universally described and almost nowhere required.
Almost. Honesty compels the exception, and it is the most interesting fact in this essay, because it runs the opposite way to every prejudice about who mandates what. Buried in the same Chinese survey is a standard now in drafting — the General Security Requirements for AI Agent Applications — which, uniquely, is set to become mandatory: it would be only the second compulsory national AI standard China has issued, after one on labelling. The report calls it the clearest signal yet of how far Beijing intends to regulate what autonomous systems actually do.
Be precise about what that standard is, though, before I claim too much from it. On the summary available, it binds baseline agent-security — access limits, logging, human confirmation for high-risk operations — not a model-independent refusal that fires on a threshold. Beijing is moving to mandate the scaffolding of the gate, not yet the gate itself. That is a genuine crack in the universal ‘recommend’, and it is smaller than my headline wants it to be. Mark it anyway, for where it appears: the first government visibly moving to require any of this is not in Brussels, or Washington, or even Singapore. It is in Beijing, and I would be a poor witness if I hid that to keep my sentence clean.
It does not rescue the Western frameworks, which remain voluntary to a fault; it sharpens the indictment of them. The two jurisdictions that originate about 88.7 per cent of the world’s notable models — by the Bengio report’s own count — are the United States and China, and the one moving from drawing the gate to requiring even its scaffolding is the authoritarian one. The liberal democracies, which never tire of telling the world what good governance looks like, are still initialling the blueprint. The gate is drawn on it, and left off the building code.
The autocrat is drafting the mandate the democracies only recommend — which ought to shame the democracies, not reassure them.
VII. The Camp That Says the Gate Will Be Outrun
The chorus is not unanimous, and the dissent is the objection I least want to face, so it gets a hearing of its own rather than a line in my confession. A dean of computer science in Beijing has argued, in as many words, that rule-based defences cannot keep pace with adaptive AI attacks. Read that slowly, because it is aimed at the centre of what I sell: a gate that refuses by a fixed rule is only as good as its rules, and an adversary that rewrites its approach faster than a committee can rewrite the rulebook will, on this view, simply route around it. He is not a lone contrarian. A substantial camp of Chinese scholars prefers what they call agile and embedded governance to hard, ex-ante rules — safety grown through participation and iteration rather than frozen into a standard — and is sceptical of exactly the kind of mandate this series is building toward. If they are right, my whole enterprise is a category error: a static wall thrown up against a moving thing.
I think they are half right, and the half they miss is the half that decides it. A gate that tries to recognise every attack — to enumerate the bad and block it — does lose the arms race; that is the lesson of thirty years of signature-based security, and the Beijing dean is describing it correctly. But that is not what a capability bound does. A deterministic refusal over a bounded, authenticated set of permitted actions does not try to recognise the attack at all. It never asks whether a request is adversarial; it asks whether the action is on the short list this agent is allowed to take, and refuses everything else by default. An adversary can be arbitrarily clever about why it wants to move the money; it cannot make ‘move the money’ appear on a list that does not contain it. Adaptivity beats a blacklist. It does not beat an allowlist that fails closed. That is the reply — and I owe the dissenters the honesty of saying it is a reply I have to make good on in the building, not one their own literature has conceded to me.
A wall you cannot argue with is not outrun by an adversary whose only weapon is the argument.
VIII. The Case Against This Essay
Now the part the argument requires, because the man making it is the least trustworthy person to make it.
I build this gate. My firm’s governance work — TrustOS, and beneath it a deterministic kernel I have named Ring Zero — is the execution layer this essay says everyone draws and few require. It runs on my laptop, and it contains, deterministically, the exact scenario the Shanghai lab imagined: the stale figure refused, the injected ‘release’ command structurally impossible, the verbal approval rejected because it carries no cryptographic signature. I did not merely argue for the kernel. I sold one. Worse: the enforceability test by which I have just judged every framework in this essay is my own instrument, not theirs — I brought the ruler that makes their clauses line up into a gap. Discount me twice, then, for the product and for the yardstick.
And there is a third cut, sharper than either, that I owe you. A six-stage risk-management framework is not mostly a gate; the great bulk of its eighty-two pages is risk identification, thresholds, evaluation and monitoring. I chose which of its clauses to lay my ruler against, and I foregrounded the two or three that are execution controls over the scores of pages that are not. The clauses are theirs. The emphasis — which sentences you have been told are load-bearing — is mine, and a fair reader should hold that against the argument.
Two further objections land, and I concede them. The first: these are not truly independent voices. China’s laboratories, like Singapore’s regulators, operate inside a coordinated national programme, and one should not mistake a state describing itself in several documents for the free convergence of strangers. Worse, nearly everything I have told you about the Chinese chorus — the blackmailing model, the scientists’ warnings, the leaders’ words — you have on the word of one interested narrator: Concordia AI, a safety advocate that discloses it advises Chinese developers for fees, translating Chinese-language statements I cannot myself read. By my own trade’s first rule, secondary coverage of a source is a rumour with a byline, and much of section three is exactly that. Weigh it accordingly. What survives is weaker and still real: instruments written for different purposes, by people who did not share a drafting table across the Pacific, reached for the same shape.
The second: I have leaned on a Chinese framework whose full English text is not yet published, working from its verified summary and its earlier version, and I have been careful — the exact count of its red-line scenarios, and one or two phrases the conference coverage embellished, I have left out rather than assert. A hostile reader is entitled to wait for the full translation. So am I.
A rival vendor will find the sharpest cut without my help, so let me make it for him. I built the ruler; I sell the one product that cleanly passes it; and I have handed you the resulting scoreboard as though it were an independent global convergence. The honest reply is not that my test is neutral — it is not, and I wrote it. It is that its load-bearing property is not only mine. The demand that a stop hold even against an adversarial system is the Shanghai lab’s, not my spec sheet. Containment is the word Bengio’s panel reaches for when it names what actually works. The preference for enforcement in the infrastructure over instructions to the model is Singapore’s regulator’s, published before I released anything. Grade every framework against their language instead of mine, and the convergence dims but does not vanish.
And here is what would refute me, because a claim no fact can dent is a frame, not an argument. If the frameworks I have cited, read in full, turn out to specify their stops as monitoring-and-report rather than refusal-at-the-point-of-action — if the circuit breaker is a dashboard alert and the off-switch a support ticket a human answers on Monday — then the convergence I claim collapses into a shared taste for telemetry, and I am selling the cure to a disease the field has not agreed it has. That is the version a fair reader should test me against, and the one I have tried, clause by clause, to rule out.
What survives all of it is not mine, and cannot be discounted with me. The clause about a stop the model cannot circumvent is the Shanghai lab’s. The finding that autonomy leaves humans too little time to intervene is Bengio’s panel’s. The observed model that reached for a person’s emails to avoid being switched off is Andrew Yao’s account, not mine. I built a product to their specification. I did not write the specification. They did.
Distrust the man who sells the gate, drew the ruler, and chose where to lay it — then read the clauses he laid it against. He did not write them.
IX. Coda: The Drawing and the Building Code
There is an honest reason for the near-universal hesitation, and it deserves stating plainly: a standard that mandates an immature control too early can freeze this year’s best guess into next year’s law, and a badly-specified gate is worse than a well-built voluntary one. Grant it. It is the strongest thing that can be said for the present arrangement.
And the reason will not be the same in every capital, so I should not pretend one charitable explanation covers them all. In Washington the restraint may be a genuine faith in markets and a distaste for rules; in Singapore, a regulator’s patience; in Beijing, something less comfortable — a preference for keeping the intervention discretionary, an option in the state’s hand rather than an obligation on a developer’s. That last point cuts against my own tidy story: Beijing’s willingness to make one agent-security standard mandatory may be less about safety than about control of a different kind. I am offering the most charitable reason for a silence that has several, and not all of them are charitable.
But it does not explain the silence in the places that pride themselves on the rule of law. No Western framework in this essay — the European Union’s, the frameworks of the American labs, the voluntary commitments — has committed even to requiring that the gate exist, with its properties and a threshold of autonomy past which it becomes non-negotiable, on any timeline at all. To describe a safeguard is to know what safety looks like. To require it is to decide to have it. Every serious authority in AI governance, in every capital that matters, has now done the first. In the one place I once thought I was arguing alone — the case that governing an autonomous agent means a gate that can refuse it at runtime — I find the People’s Republic of China, the Monetary Authority of Singapore, and a Turing laureate’s international panel have all, from different starting points, described the same control.
And then the strangest turn of all: the first of them reaching to make any of it binding is the one whose politics I trust the least. The gate is drawn in every language. It is required, so far, in almost none — and where it is finally being required, it is not by the democracies that keep telling the world what good governance looks like.
Everyone has drawn the gate. The only question left is who signs the law — and, so far, the wrong hand is first to the pen.


Leave a comment