Trusted Agents: The Most Dangerous Euphemism of the 2020s
I. Give the Index Its Due
They refused to count frameworks, and that refusal is the reason to trust them.
Start with the thing most readers skip. Before I say a word about what the Global Index on Responsible AI cannot do, I want to be honest about what it does, because the whole of this essay rests on the index being good. Not adequate. Good. The Global Center on AI Governance built its second edition across 135 countries and 68,138 individual data points, over a period running from 1 November 2023 to 30 September 2025. That is not a poll. It is closer to a census of a field most of us only gesture at, and the discipline of it shows on every page.
Here is the part I admire most. The lazy version of this index would have counted laws. It would have tallied how many countries published an AI strategy, added the strategies up, ranked the total, and called the ranking governance. Everyone in this space knows the game. A country writes a glossy national framework, the framework earns a press release, the press release earns a citation, and somewhere a league table ticks upward while nothing on the ground has moved. The GIRAI team saw that trap and walked around it. Their own foreword rejects the idea that tallying frameworks could ever amount to measuring governance, and they built the harder instrument instead.
So they did the harder thing. They asked whether a law was binding or merely aspirational. They asked whether a government that announced a principle then built anything to carry it: an oversight body, a budget, a register, a route to redress. They looked for evidence that implementation followed the promise, and they were willing to record its absence. They even subtracted a penalty where they found credible evidence of a state itself using AI in ways the field treats as unacceptable. This is not framework-counting dressed up. It is an instrument built by people who clearly resented framework-counting. I think it is the best measure of responsible-AI governance anyone has produced. I want that on the record before I draw a single line around it, because the line I am going to draw is not a complaint. It is a boundary the index draws on itself, and you can only see it once you have taken the index seriously.
II. What It Actually Measures
Existence of an initiative is the proxy doing the work.
The architecture is worth walking slowly, because the argument I am making lives inside it, not against it. GIRAI scores each country across five equally weighted dimensions, and within each dimension it reads three pillars that carry very different weights. Government AI Policy accounts for 60 per cent. Civil-society and non-state action accounts for 10. Enabling conditions, the surrounding capacity that makes governance possible, account for the remaining 30. From that composite the index then subtracts a penalty for real-world misuse, which I will come to. The centre of gravity, by design, is what governments themselves have put in place.
The 60 per cent government pillar splits again, and this split is the load-bearing part of the whole method. Half of it is government frameworks: the laws, policies, strategies and standards a state has adopted. The other half is government-led initiatives, and this is where GIRAI lifts itself above framework-counting. An initiative, in the index’s coding, is evidence that something followed the framework into the world: an implementing body, a budget line, a monitoring function, a register, a redress route. The presence of an initiative is how the index detects implementation. That is the mechanism I want to name plainly, because the entire argument turns on it. Implementation is measured as the existence of an accompanying initiative. The proxy for a framework having teeth is that a corresponding institution can be shown to exist.
Two coded attributes matter enormously here, and both show the seriousness of the design. The first is enforceability, scored as a binary property of each instrument: is this framework legally binding, or is it guidance a regulated party may ignore without consequence. The second is operationalisation, scored by whether a framework arrives with the apparatus that would let it work at all, a responsible body, a budget, a monitoring and evaluation function. This is a team refusing to accept a beautiful document at face value, insisting instead on the machinery a state actually stood up around it.
The numbers this produces are sobering, and they are the reason the naive critique fails. On GIRAI’s own count the global average is 35 out of 100. The Global North averages 55; the Global South, 27. Most striking, and the figure I would put on the first slide, is that among countries with an active framework, the index finds evidence of implementation in only about 55 per cent of cases, and closer to 45 per cent across the Global South. Binding-ness tells the same story from another angle: roughly 78 per cent of active frameworks in the South are non-binding, against about 42 per cent in the North. Independent oversight bodies exist in just 28 countries. Public disclosure of a government’s own algorithmic systems is required in roughly 18 per cent of them. And the single largest source of enforceable-protection gains between the two editions is one law: the European Union’s AI Act drives 51 of 76 such gains, almost the entire movement of the needle. Anyone who says the index measures whether frameworks exist rather than whether they bite has not read it. It measures, quite precisely, how far each state carried a framework into binding force and standing institutions, and it performs that measurement better than anyone.
III. The Line It Will Not Cross
A register can exist without one action being refused.
Now the turn, and I want to make it in the index’s own terms rather than mine. GIRAI describes its own result with care. It presents itself, in the disclaimer it prints on page twenty-six, as “a measure of publicly verifiable governance conditions, not as a complete account”. That is not marketing modesty. It is a methodological fence, put there on purpose, and it tells you exactly what the instrument is not. It does not assess individual AI systems; its unit of analysis is the national governance condition, not the deployed model, and certainly not the single action a model takes. Search the hundred-odd pages for the words runtime or technical control and you will not find them, because they belong to a different layer than the one the index surveys.
Follow the consequence through each of the signals I praised. Enforceability is coded as an attribute of a framework: this law is binding. That is a true and useful fact about a state. It is not a fact about any particular action a deployed system took on a particular day. Implementation is coded as the existence of an initiative: a register was established, a body was funded, a monitoring function was stood up. Again true, again useful, and again a fact about the institution rather than about an event. Oversight is coded by counting the supervisory bodies that exist. The index can verify that such a body exists, is independent, and holds a mandate. It cannot, and does not claim to, verify that on a given Tuesday that body’s remit caused one specific model action to be stopped before it completed.
This is the distinction the whole essay turns on, so let me put it as cleanly as I can. Every one of GIRAI’s strongest signals is a measurement of what a state has built. A binding law is a built thing. A register is a built thing. A budget, an oversight body, a redress route, all built things, all verifiable, all correctly counted. What none of them measures, and what the index is careful never to claim to measure, is whether a single system action was bound, blocked or contained at the moment it was attempted. Existence of the institution is the proxy. It is a good proxy for its purpose. It is still a proxy, and it sits one layer above the event it stands in for. I have made a related argument before, in Everyone Draws the Kernel, about how much of the governance conversation describes the building and calls it the lock. GIRAI is the most rigorous description of the building I have found, and it is scrupulous about not describing the lock.
IV. Promise and Protection
The promise is national. The protection, if it exists at all, is per action.
The report has a phrase for the space I am pointing at, though it uses it more broadly than I will: the distance between promise and protection. I have not been able to stop thinking about it, because it names, without meaning to, the layer sitting beneath the one the index can see. The promise is institutional. It is the law, the strategy, the oversight body, the budget. The protection, in the strict sense of an action refused at the moment it is attempted, is a single event, to one person, inside one system. Those are different altitudes, and no amount of excellence at the first guarantees anything at the second.
The layer beneath the index has a name in the engineering literature now, and it is worth being concrete, because the abstraction dissolves the moment you look at a real mechanism. Consider the pattern described in the CaMeL work. It places a deterministic capability bound at the tool-call boundary: at execution time, when a model attempts an action, a separate and deterministic layer checks that action against a fixed policy and either permits it or refuses it, in bounded time, independently of how the model was trained, graded or persuaded to want. It runs per system, per action, and it fails closed. That is protection in the strict sense. It is also, by construction, invisible to an index of national institutions, because it is a property of one system’s execution path, not of any law, register or oversight body a state has built.
Take the strongest possible counter-example from inside the report, because I want to test my claim against the best case rather than a weak one. The index finds that the largest single driver of enforceable-protection gains this edition is the EU AI Act. If any instrument were going to reach the runtime layer, it would be this one. Read its human-oversight provision, Article 14, closely, though, and it requires high-risk systems to be designed so that they can be effectively overseen, so that a person can understand the system, intervene, and if necessary interrupt it. Those are institutional and design duties. They require that the capability to intervene exists and is resourced. The Act mandates that a human be positioned to stop the system. It does not itself constitute the mechanism that stops it, and it does not specify that a given action is deterministically refused at execution time independently of the model. The best enforceable protection the index can find is real, and it is one layer up from where enforcement finally bites.
I should declare my interest here, in the body, where the house rules say it belongs and not in a footnote. I build the runtime layer I have just described. I sell deterministic, per-system, execution-time binding, the exact thing I am arguing an index of national governance cannot measure. So read the next two sections as adversarially as you like. If my argument only works because it flatters what I sell, it does not work, and I would rather you found that out than believed me.
V. Why the Index Cannot Go There, and Should Not
An index can measure the building. It cannot measure the lock.
Here is where I want to be most careful, because the cheap move now would be to treat the boundary I have found as a failing. It is not. A national-governance index measures the conditions a state establishes, and its unit of analysis is the country. To reach the runtime layer it would have to inspect the execution path of individual deployed systems, thousands of them, most proprietary, most changing week to week, none reachable through the desk research and expert survey that give the index its comparability and its defensibility. The moment it tried, it would lose the very properties that make it authoritative across 135 countries, and collapse into noise. The boundary GIRAI draws is not the edge of its ambition. It is the definition of what it is measuring.
So the two layers are not rivals. They answer different questions. The index answers, with more rigour than anything else available, whether a country has built the institutional conditions for responsible AI: the binding law, the funded body, the register, the redress route. The runtime layer answers whether one system’s one action was refused when it mattered. You need both, and you cannot substitute either for the other. A country can score well on the first and still deploy a system that, at execution time, does the wrong thing, because a funded oversight body is not the same object as a refused action.
I have circled this frontier before, from other directions. In Everyone Draws the Kernel the argument was that every serious actor, however they describe themselves, ends up drawing some line of deterministic constraint at execution time. In The Tired Human at the Gate it was that oversight resting on a human’s attention fails exactly when the volume rises, which is the same assumption Article 14 quietly makes. GIRAI, read generously, is the strongest possible portrait of the institutional gate, the observation layer at national scale. What it cannot show you is the point where an action is refused whatever the surrounding institutions look like. The observation layer and the enforcement layer are not rivals, they are different measurements, and the field keeps mistaking a good one for both. The index measures the building, superbly. Someone still has to measure the lock.
VI. The Case Against This Essay
The strongest version of the objection is that I am selling the gap I found.
Let me argue the other side as hard as I can. The most serious objection is the one about my incentive. I sell runtime enforcement, and I have written an essay whose conclusion is that the world’s best governance index cannot measure runtime enforcement. That is a conveniently self-serving shape, and you should weigh it as such. My defence is only that the boundary is checkable against the report itself. GIRAI says it does not assess individual systems; the words runtime and technical control genuinely do not appear; the disclaimer is on the page. If those facts were not there, the essay would deserve to fall, interest declared or not.
The second objection is sharper on the merits. A hostile reader will say I have smuggled in the naive claim through the back door, that when I talk about the institution layer I am really just saying GIRAI measures whether frameworks exist rather than whether they work. That claim is false, and the index’s own numbers sink it: the 55 per cent implementation finding, the binding-versus-non-binding axis, the misuse penalty. So I concede it before it is made and build on it instead. GIRAI measures enforcement at the institution layer, and measures it well; it does not measure enforcement at the runtime layer, and says so. If I have blurred those two anywhere, correct me there, because the difference is the entire argument.
The third objection is that measuring at the institution layer is a limitation I am dressing up as a discovery, when any competent governance index measures institutions and that is correct. I concede this fully. Measuring at the institution layer is the right thing for a national governance index to do. I am not saying GIRAI should reach the runtime layer. I am saying the runtime layer exists, sits beneath the one the index measures, and is where a different kind of enforcement lives. Locating a boundary is not the same as faulting the instrument for having one.
Fourth, the falsifier and its thinnest edge. I am wrong if GIRAI, or a comparable authoritative index, is shown to score whether individual system actions are bound at execution time, or if the existence of institutional activity turns out to reliably guarantee runtime enforcement, so the proxy and the thing collapse into one. The place that distinction is thinnest is a binding law that mandates a technical control which is itself a runtime bound. Even there the index scores the law’s existence and binding-ness, not the execution of the control; the mandate and the mechanism remain different objects. But I concede that is where my line is hardest to hold, and I would rather mark it than hide it. Every incident I have watched up close happened in a jurisdiction with frameworks on the books. That is the test, and I am content to be measured by it.
VII. Coda
A binding law is a fact about a state. A refused action is a fact about a system.
The Global Index on Responsible AI is a genuine achievement, and I have tried to say so without hedging, because the respect is the point and not the throat-clear. It counts what a state has built, across more of the world and in more detail than anything else in the field, and it refuses the flattering shortcut of counting frameworks and calling the count governance. That refusal is exactly why the line it will not cross is worth reading. The index stops, honestly and on purpose, at the institution.
What lives past that line does not disappear because an index cannot see it. A country can pass a binding law and stand up a register and appoint an oversight body, and every one of those facts can be true and scored and ranked, and still, in a running system somewhere inside that country, a single action can go through that should have been stopped. A binding law is a fact about a state. A refused action is a fact about a system. GIRAI measures the first with more rigour than any instrument I know, and it tells you, in its own words, that it does not measure the second. The promise is national. The protection, if it exists at all, is per action, at execution time, or it does not exist. The index measures the building, superbly. Enforcement is measured one layer down, on each action, or it is not measured at all.


Leave a comment