Every serious plan for governing frontier AI eventually promises the same thing: an independent check. The interesting question is never whether we want one. It is who sets the terms under which the check is performed – and, almost always, it is the party being checked.
I. The Evaluator That Marked Its Own Homework
On 27 February 2025, METR published its pre-deployment evaluation of GPT-4.5. Read the caveats before the findings. The evaluator had the model for roughly a week before release. OpenAI supplied the context and a curated subset of the benchmark results. And then the sentence that matters, which METR volunteered rather than concealed: its short evaluations “likely underestimate GPT-4.5’s capabilities”, and were not sufficient, in its own assessment, to rule out large-scale risks.
Begin here, and not on a critic’s accusation, because the admission is more valuable than any accusation could be. A hostile outsider claiming the evaluation was rushed proves nothing. The evaluator saying so, in its own report, on the record, proves a great deal — and what it proves is not about METR at all. METR behaved impeccably. It disclosed its constraints instead of laundering them. The finding is structural: the party being assured decided how long the assurer had, and what the assurer was allowed to see. The assurance was real work performed under conditions the assured party set.
Hold that distinction; it becomes the whole essay. Candour is not independence.
Candour about a constraint is not the same as the absence of the constraint.
II. Two Routes, One Signature
Now change scale, from one lab’s voluntary evaluation to a continent’s binding law, and watch the same shape appear. The European Union’s AI Act is the most demanding AI statute in force. Under Article 43, a high-risk system reaches market by one of two conformity-assessment routes. The first, Annex VII, is third-party assessment by a notified body. The second, Annex VI, is internal control – and Annex VI, in the regulation’s own construction, does not provide for the involvement of a notified body. The provider verifies its own quality-management system. The provider examines its own technical documentation. The provider signs.
Which route governs most high-risk AI? The self-certifying one. Third-party assessment is reserved, narrowly, for parts of the biometric tier; for the rest of Annex III, internal control is the default path to a CE mark. The flagship regime for high-risk AI, the one every other jurisdiction measures itself against, routes the majority of its subjects to verifying themselves.
Read that again, because the reflex is to assume the strong regime must have closed this door. It did not close it. It built it into Article 43 and called it proportionate.
The most demanding AI law in force lets most of its subjects verify themselves.
III. The Sort
This publication has one instrument it applies to every control, and it applies it here. Ask of any assurance mechanism: can it refuse an action at execution time, deterministically, in bounded time, independently of the party whose behaviour it governs? Four conjuncts, and all must hold. Fail any one and the mechanism is not enforcement; it is observation, however rigorous, however well-funded, however sincere.
Most of what the field calls audit fails the fourth conjunct, and fails it not on rigour but on arrangement. A pre-deployment evaluation performed on access the developer grants, for a window the developer sets, over information the developer curates, is not independent of the party it governs. It may be excellent. It may catch things. It cannot, structurally, catch the things the granting party preferred it not see, because the granting party chose the aperture. An internal-control conformity assessment fails the same conjunct more plainly still: the assured party and the assuring party are the same legal person.
None of this makes the work worthless. Observation has enormous value. It is the confusion of observation for enforcement that this essay is about — the moment a report that could only ever describe is treated as a gate that can refuse.
An instrument that samples what the assured party reveals cannot refuse what the assured party conceals.
IV. Audit-Washing
So why does the confusion persist, when stated this plainly it seems obvious? Because a check that cannot refuse still does something valuable to the party that commissioned it: it confers legitimacy. This is the mechanism the accountability literature named audit-washing, and it is the load-bearing move.
Hartmann and colleagues, writing in 2024 on the EU audit ecosystem, put the structural version of it: the Act mandates audits but withholds from researchers and civil society the data and model access that independent verification would require. An audit conducted on the auditee’s terms verifies compliance with what the auditee chose to show. It cannot verify conduct, because conduct is precisely what a curated aperture omits. The output is a document that reads as scrutiny and functions as endorsement.
Notice what has happened to the incentive. If an assurance the assured party scopes still yields the legitimacy of having been assured, then every rational assured party will supply exactly enough access to produce the document and not one byte more. The arrangement does not merely permit shallow assurance. It selects for it.
A check that cannot refuse, but can still legitimise, is not a safeguard. It is a laundry.
V. Independence Is a Property of the Arrangement
Here is the distinction the regimes discovered and did not name. We keep asking whether an auditor is independent, as though independence were a virtue an auditor could possess – a matter of character, credentials, professional ethics. It is not. Independence is a property of the arrangement, not of the auditor, and it has exactly three load-bearing components: who controls access, who controls funding, and who controls scope. Where the assured party controls any of the three, the auditor’s virtue is irrelevant, because the auditor never sees the thing the arrangement was built to keep from it.
This is why the first fix is not a better auditor, or a code of conduct, or a professional body. Those improve the character of a party whose character was never the binding constraint. The fix is a separation the assured party cannot re-scope: independent access, independent funding, independent selection, set by someone other than the party under assurance and not revocable by it. That separation is a necessary condition – the thing without which the METR caveat can only ever be a disclosure. It is not, on its own, sufficient, and §VI will make me concede exactly why: someone still has to specify what the separated gate may refuse, and that specification is a fallible human act no arrangement removes. The separation buys the capacity to refuse. What is refused is a further, upstream question.
I should declare the interest, in the body, where the argument becomes convenient to me and not in a footnote where it is easy to miss. I build toward exactly this separation — deterministic, independently scoped enforcement — and a reader should therefore discount this section accordingly, and weigh it against the fact that the same incentive I am describing operates on me. So discount it. Then notice that the primary evidence for the claim comes from the assured parties themselves — from METR’s own report and from the European Union’s own statute — and not from me, which is the part the discount cannot reach.
Independence is a property of the arrangement, not the auditor — and arrangements can be engineered.
VI. The Case Against This Essay
Four objections, of four kinds, answered honestly.
The category objection. The essay speaks of independent and captured as though assurance were a switch, when it is a dial. METR itself notes more access and more editorial latitude than in prior engagements; some regimes separate one of the three levers and not the others. Answer, partial. The objection lands, and I concede the binary: independence is a spectrum, and the essay should not imply that a partially separated arrangement is worth nothing. What survives is narrower and still bites — that wherever the assured party retains any one of access, funding or scope, the assurance cannot refuse along that axis, and the marketing rarely says which axis was retained.
The regress objection. The essay demolishes assurance for depending on a party who should not be trusted to set its own terms — so why is the separated, deterministic gate it recommends trustworthy, when someone still had to enumerate what that gate may refuse? Conceded. This is the strongest objection and it is correct. A separation of access, funding and scope removes one dependency; it does not remove the semantic act of specifying the permitted set, which remains a fallible human judgement made upstream. The claim must therefore shrink, and I shrink it: structural separation is a necessary condition for an assurance to be able to refuse, not a sufficient one. Anyone selling it as sufficient — including me, on a careless day – is selling the same legitimacy this essay is attacking.
The scope objection. Self-certification is ordinary product-safety practice; most goods bearing a CE mark self-declare, and the AI Act inherits a mature framework that reserves third-party scrutiny for the highest tier by design. Answer. Granted for static artefacts — but the analogy breaks on the thing being assured. A conformity-assessed kettle does not act after certification; a high-risk model does, and self-verification of a static file is a weaker instrument against conduct than against composition. The essay’s claim is not that self-certification is always scandalous. It is that self-certification of an actor is weaker than the CE precedent assumes.
The interest objection. The author sells the separation he recommends, so discount the conclusion. Answer, and already declared. Discount §V accordingly — I have said so in the body. Then weigh it against the fact that every primary here comes from the assured parties, not from me. The incentive I describe operates on me too; the evidence does not.
The strongest case against this essay is that separation is necessary, not sufficient — and it is right.
VII. Coda
Return to the evaluator that marked its own homework. The point was never that METR marked it badly; it marked it honestly, and said so. The point is that it was handed its own homework to mark, by the party whose grade depended on it, and that this arrangement is not an aberration at the frontier but the default in the statute — the same shape at two magnifications, discovered independently, named by neither. We have spent the debate asking whether our auditors are trustworthy. It is the wrong question, politely phrased. The right one is colder: by what arrangement could they refuse, and who holds the key to the room.
We asked whether the auditors were honest. We should have asked who held the key to the room.
Sources
- METR, GPT-4.5 pre-deployment evaluations, 27 February 2025.
- Regulation (EU) 2024/1689 (AI Act), Article 43 and Annex VI.
- Hartmann, Laranjeira de Pereira, Streitbörger & Berendt, Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society, 2024.


Leave a comment