Men racing to build superintelligence have now all proposed a way to police it. All reached for a licensing regime. None of them reached for the execution layer.
—- Dr. Luke Soon
HX Architect, Futurist and AI Ethicist
I. Five weeks, three manifestos
On Tuesday 14 July 2026, Demis Hassabis published A Framework for Frontier AI and the Dawning of a New Age on his Substack, with an accompanying exclusive in Axios. The centrepiece is a US-led, industry-funded, majority-independent Standards Body, modelled explicitly on FINRA — the private self-regulatory organisation that polices Wall Street’s broker-dealers under SEC oversight. Frontier labs would submit models up to thirty days before release for evaluation across dangerous cyber capability, biological and nuclear risk indicators, autonomous-capability escalation, guardrail-bypass susceptibility and deception. Voluntary first. Mandatory for US market access once the protocol proves robust. A board stacked with Turing Award winners, alongside industry, government and open-source representation. He wants it operational before year end.
It landed to rare cross-industry applause. Sam Altman, Satya Nadella and Elon Musk all praised it publicly. Anthropic’s Jack Clark called the framework excellent, and observed that everyone at the frontier now agrees third parties should test these systems.
He is right that they agree. What almost nobody has said out loud is what they agree on.
Within five weeks, the three people racing hardest toward superintelligence each published a regulatory prescription. Dario Amodei wants an FAA for AI — a federal agency empowered to block a release outright, from day one. Hassabis wants a FINRA for AI — industry-funded, federally overseen, voluntary pre-release review hardening into mandatory market access. Sam Altman, writing in the Financial Times, wants an IAEA for AI — a US-led international forum certifying countries, companies and standards, using access to frontier models and markets as leverage.
Three of the most consequential technologists alive reached into the twentieth century for an institutional analogy, and all three came back holding the same object.
A gate.
Not one of them proposed an instrument that operates during execution. Not one proposed a mechanism that can intervene in a running agentic workflow. The celebrated convergence of the frontier is a convergence on admission control — and admission control is the one thing that agentic AI, by its very architecture, routes around.
This is not a small oversight at the edge of an otherwise sound proposal. It is the proposal.
II. The analogy indicts the argument
Start with FINRA, because Hassabis chose it and the choice is doing enormous rhetorical work.
FINRA does not certify products before launch. That is not what it is, and it never has been. FINRA’s actual instruments are conduct supervision, trade surveillance, order audit trails, books-and-records obligations, and the authority to halt, fine and bar. It watches executed transactions, continuously, after they occur, and it intervenes in live market conduct. Its entire apparatus rests on a single hard-won proposition: you cannot know in advance whether a firm will behave, so you instrument what it actually does.
Hassabis has borrowed the institutional prestige of FINRA and attached it to something FINRA has never been — a type-approval regime. The nearer analogues are UL listing, or FDA premarket review. Test the artefact. Certify the artefact. Release the artefact into a world that will do precisely whatever it likes with it.
Amodei’s FAA fares no better under inspection. The FAA does not merely certify airframes; it runs air traffic control. It is a real-time authority over aircraft in motion, with the power to ground, reroute and separate while the thing is flying. If you took the FAA analogy seriously you would build ATC for agents — a control plane with live authority over trajectories in flight. Amodei has taken the certification half and left the control tower on the runway.
Altman’s IAEA is the most revealing of the three. Safeguards agreements, declared inventories, inspectors on site — the IAEA’s power lies precisely in continuous verification of material in use, not in a one-off approval of reactor blueprints.
Each of the three men cited an institution whose actual power lives at runtime, and then proposed the pre-market fragment of it.
III. Certification is not enforcement
Here is the diagnostic I keep returning to, because it has survived every framework I have applied it to.
Can this instrument refuse anything?
The Standards Body can refuse a release. It cannot refuse an action.
That distinction is not pedantic; it is the whole of the matter. A frontier model is not a system. It is a component. The system is the model, plus system instructions, plus a tool layer, plus memory, plus guardrails, coupled by an orchestration architecture — and every layer other than the model is assembled by the deployer, after certification, out of materials the Standards Body never saw and over which it holds no authority whatsoever.
The technical consensus now forming in regulated sectors — in banking model risk practice, in the runtime governance literature, in agentic security research — converges on an observation the frontier manifestos do not engage with at all: risk in agentic systems arises from execution trajectories, not from a stable input-output mapping. The material failures are process failures. Unsafe tool use. Skipped approvals. Privacy breaches. Uncontrolled side effects. They emerge during runtime, out of composition, and they are structurally invisible to any evaluation of a component in isolation.
Take the canonical worked example, now well established in the financial-services literature: an agent drafting a credit memo.
Retrieval returns audited accounts twenty-six months old where policy permits eighteen. A third-party document carries an embedded directive asserting that credit committee approval has already been granted, and the agent acts on it. EBITDA is double-counted and a coverage ratio emerges at 2.82 rather than 1.82 — fluent, plausible, entirely wrong, and it sails through every semantic check applied to it. The agent treats “approval confirmed verbally” as satisfying an approval requirement for which no authenticated record exists. And across repeated runs, the proportion of unauthorised releases climbs steadily without any single run tripping a violation — orchestration drift, invisible at per-run granularity and therefore invisible to every observation-layer instrument ever shipped.
Now ask which of those a thirty-day pre-release evaluation catches.
None. Not one.
Every failure is post-certification. Every failure is a property of the composition, not the component. The tool layer did not exist when the model was tested. The retrieval corpus did not exist. The approval gate did not exist. The authority scope did not exist. You cannot evaluate a trajectory through a system that has not been built yet — and the system will be built by a deployer who will wave the certificate at the audit committee.
IV. What the panel already told us
I have spent a long time reading the people who built this technology and the people who warned about it. On this question their positions do not converge on the gate. They converge somewhere else entirely — and several of them have already published the argument against it without appearing to notice.
Roman Yampolskiy has produced the most inconvenient result in the field, and it lands directly on this proposal. His impossibility work — unpredictability, unexplainability, unverifiability and, most pointedly, unmonitorability — argues from theoretical computer science that we cannot reliably foresee which capabilities will emerge before they manifest, and that certain safety guarantees are formally unavailable. Whatever one makes of his conclusions about existential risk, the narrow technical claim is devastating for a pre-release regime: if capability emergence cannot be reliably predicted prior to manifestation, then an evaluation conducted prior to manifestation is structurally incapable of the task it has been chartered to perform. It will pass models that later exhibit exactly what it was built to detect — and it will do so carrying the full authority of a federally overseen body.
Yampolskiy’s own remedy — abandon the general, build the narrow — is not one I accept. But his impossibility results point somewhere he does not go and I do: if you cannot predict behaviour, you must constrain authority. Govern what the system is permitted to do, at the moment it attempts to do it. That is not a philosophical preference. It is the only remaining engineering option.
Yoshua Bengio has done more than anyone to build the scientific evidence base such an institution would need — the International AI Safety Report, the hundred-plus contributing experts, the discipline of stating plainly what we know and what we do not. But an evidence base is an input to design-time governance. It tells you what to test for. It does not stand in the execution path, and Bengio has never claimed it does. The honest reading of his work is that it makes the gate better informed, not that it makes the gate sufficient.
Geoffrey Hinton stopped arguing years ago about whether these systems are dangerous and started arguing about whether we retain control. Control is a runtime property. It is exercised over actions, in the moment of action, or it is not exercised at all. A certificate is not control. It is a statement about a past test.
Fei-Fei Li has spent a decade insisting that AI be judged against human reality rather than against benchmarks, and her critique is the sharpest available objection to a body whose entire epistemology is regularly updated benchmarks. Benchmarks are measured on models. Consequences land on people — in deployed contexts, through tool calls that touch real accounts, real records, real lives.
Dario Amodei is, to his credit, the most honest of the three CEOs about what enforcement actually requires. He wants hard power and says so: a federal agency able to block from day one. He has understood that voluntary regimes decay under commercial pressure. What he has not extended is the same realism to the layer below. An agency that can block a release still cannot block a transaction.
Kai-Fu Lee supplies the geography the manifestos omit. Deployment velocity is not evenly distributed, and the centre of gravity for agentic adoption is not the jurisdiction writing the rules. A US-first body certifies at the source, while the consequential composition happens across Asia, in enterprises assembling agentic workflows at a pace no Western certification cadence will ever track.
Eric Schmidt signed the We Must Act Now statement on the very day Hassabis published, and has cited polling showing roughly four in five American adults want AI regulated even at the cost of slower progress. But Schmidt is also the field’s most consistent voice on competitive dynamics — the race with China, the open-versus-closed contest, the impossibility of holding an advantage in place. Both things are true simultaneously, and their intersection is the problem nobody wants to name: a gate in a race is a gate that will be arbitraged. Not by villains. By ordinary commercial actors deploying certified components into uncertified compositions, entirely lawfully, at speed.
And Hassabis himself, in a quieter register than his manifesto, has said the thing that matters most. Asked what keeps him up, he pointed past job displacement to meaning and purpose, and asked whether we have the right institutions at all. It is the best question in this entire debate. And he has answered it, this month, by proposing an institution that governs models — while the thing that will actually reshape human meaning is not the model but the system: acting in the world, in composition, at runtime, on our behalf, and frequently without our knowledge.
V. Three consequences that follow immediately
The open-weight coverage claim is rhetorically clean and operationally void. The framework is said to cover every frontier-class model regardless of origin or licensing. Admirable. A certified open-weight model becomes uncertified the instant somebody wires it to a tool layer, and no mechanism in a pre-release regime reaches that moment. The certificate travels; the governance does not.
“Passed the Standards Body” becomes a procurement credential — and this is where the damage lands. SOC 2 did not make software secure. It made buyers stop asking. A frontier certification badge will discharge the buyer’s diligence obligation at precisely the layer where the buyer’s own residual risk actually lives: their orchestration, their tools, their data, their authority grants, their approval gates. The certificate will be produced in a procurement meeting and the conversation will end. We will have built an instrument that transfers the appearance of assurance from the party who can enforce nothing to the party who could have enforced everything — and then relieved that second party of the impulse to try.
The capture risk is structural, not incidental. The three organisations proposing this already possess the lawyers, security teams, government relationships and technical staff to navigate a complex certification process. Startups and open-source developers do not. Rules written to make AI safer may entrench the firms that wrote them. That critique is now standard, and correct as far as it goes — but it understates the problem. A captured gate is bad. A captured gate that also does not govern the layer where harm occurs is worse, because it purchases legitimacy without purchasing safety.
I would add a fourth, which I have not seen raised anywhere and which I care about most. Singapore’s IMDA and AI Verify produced the world’s first agentic-specific governance framework. Its instincts — tripartite, practical, testing-oriented, built around trust rather than prohibition — are exactly what a global regime needs, and a US-first body chartered before year end will consult them approximately never. The Global South and Asia will be handed a standard to comply with rather than a standard they helped write. That is not merely inequitable. It is technically impoverishing, because the deployment reality those jurisdictions inhabit is precisely the reality the standard will fail to describe.
VI. The layer nobody is building
Every governance instrument now in the field — vendor platforms, assurance frameworks, model cards, evaluation suites, named-accountability regimes, and now a frontier standards body backed by the three most powerful labs on earth — operates above the execution layer.
They observe. They document. They assess. They attest. They advise.
Not one of them sits in the execution path with the authority to say no.
The deterministic execution kernel — Ring Zero, the layer at which an agentic action is either admitted by a governed transition system or is not defined within it at all — remains unwritten. The vendors are shipping user-space utilities and calling them operating systems. A Standards Body would be the most authoritative user-space utility yet built, and it would still be user-space. The privilege ring at the centre is empty, and the agent’s trajectory passes straight through it, untouched.
The test at that layer is brutally simple and it never changes: a control that cannot be evaluated deterministically over telemetry, in bounded time, independent of the language model, is not a control. It is guidance. It can be reasoned around, prompted around and drifted around — and it will be. Not occasionally. As a matter of statistical certainty, across millions of runs.
Everything above that line is an opinion about a system. Only the kernel is a decision about an action.
VII. And above it, the layer nobody is even discussing
There is a second absence, and for me it is the deeper one.
Ask what the Standards Body tests for: cyber capability, biological and nuclear indicators, autonomous escalation, guardrail bypass, deception. Every item is a harm-avoidance metric. Every one asks whether the system will do something terrible.
Not one asks whether it will do something good.
Not one asks whether the deployed system preserves human capability or quietly erodes it. Whether it sustains meaningful agency or manufactures the comfortable frictionlessness in which judgement atrophies. Whether it holds relational integrity intact. Whether the humans nominally in the loop retain any genuine ability to intervene — or have been converted into moral crumple zones: named, numbered, accountable, and structurally powerless to refuse anything at all.
This is the HX Proof Gap, and it applies here with full force. We are about to charter the most significant AI institution of the decade around the proposition that safety means the absence of catastrophe. It does not. Safety is the floor. Flourishing is the objective, and no framework that cannot measure it can claim to be governing the thing that matters.
Hassabis, of all people, knows this. He has said so. It simply did not make it into the framework.
VIII. Long-AND, not Short-OR
Let me be exact about what I am arguing, because the reflex in this field is to hear critique as opposition.
The Standards Body should exist. Independent pre-deployment evaluation of frontier capability is necessary, overdue, and vastly better than the improvised interventions that have twice this summer substituted political reflex for technical process. Hassabis has done the field a real service by producing an institutional design rather than another set of principles, and the cross-industry praise it has drawn is deserved.
It is not sufficient. And the danger is not that it fails — the danger is that it succeeds: visibly, prestigiously, with a Turing laureate board, a funded compute budget, and a badge that finds its way into every procurement pack in the Western world. That is the Short-OR trap in its purest form. An expedient partial instrument, adopted quickly, celebrated broadly, and then permitted to stand in place of the comprehensive one for a decade.
The Long-AND position is that we need all of it, at every altitude:
• Certification and enforcement.
• Evaluation of the component and deterministic governance of the composition.
• A gate at release and a kernel at execution.
• Standards for what the model can do and structure for what the system is permitted to do with it.
• The absence of catastrophe and the presence of flourishing.
Humanity and AI. It was never a choice, and every framework that treats it as one will fail at the layer it declined to build.
Hassabis has designed a referee. He has appointed him properly, funded him well, and given him a whistle of real authority.
The referee stands at the touchline. The match is played in the middle. And when the tackle comes in late — as it will, a hundred thousand times a day, in credit files and clinical notes and customer accounts and procurement systems — he can show a card afterwards.
He cannot stop the ball.
Somebody has to build the thing that can.
Genesis: Human Experience in the Age of Artificial Intelligence
GenesisHumanExperience.com
Sources and further reading
• Demis Hassabis, A Framework for Frontier AI and the Dawning of a New Age, Substack, 14 July 2026
• Axios, Behind the Curtain: AI godfathers converge on regulations, 16 July 2026
• Dario Amodei, Policy on the AI Exponential; Sam Altman, Industrial Policy for the Intelligence Age (Financial Times)
• Roman V. Yampolskiy, AI: Unexplainable, Unpredictable, Uncontrollable; Unmonitorability of Artificial Intelligence
• Bengio et al., International AI Safety Report 2026
• Szpruch, Sudjianto, Bhatti & Ang, Scalable Runtime Governance for Agentic AI in Financial Services, SSRN 6567199, April 2026
• Wang et al., MI9: An Integrated Runtime Governance Framework for Agentic AI
• Khoo, Foo & Lee, Agentic Risk & Capability (ARC) Framework
• IMDA & AI Verify, Model AI Governance Framework for Agentic AI, Singapore, January 2026
• Madeleine Clare Elish, Moral Crumple Zones; Ben Green, The Flaws of Policies Requiring Human Oversight


Leave a comment