The Tired Human at the Gate!

Dr Luke Soon · Genesis: Human Experience in the Age of Artificial Intelligence · August 2026

I have spent a year applying one test to governance instruments: can this control refuse an action at execution time? I have never once applied it to the component every framework in the world depends on most. So let us do that now, and accept what comes back.

I. P TO PROCEED

Between 1985 and 1987 a radiation therapy machine called the Therac-25 delivered massive overdoses to six patients. Three died. The machine had a human in the loop. It had an approval gate. The operator was required to confirm before the beam fired.

The machine threw cryptic error messages constantly, and clearing them required the operator to press P to proceed. Dozens of times a day. Every day. The confirmation became a reflex, then a muscle memory, then nothing at all. The operator was nominally in command of a system that had quietly trained her, through its own reliability, never to question it.

The failure was not that the gate was absent. The gate was present, staffed, and logged. The failure was that the gate had no capacity to refuse, because the only component capable of refusing had been conditioned out of the capability by the system it was meant to constrain.

I have written at length about the distinction between instruments that enforce and instruments that observe, and about the strange fact that almost every framework in the world has built the second and called it the first. What I have not done, and should have done first, is turn the test on the component that every one of those frameworks names as its ultimate backstop.

A control that was trained not to refuse is not a control. It is a ritual with a timestamp.

II. THE TEST, TURNED AROUND

The test is one question, and it has served me well against documents, platforms and national frameworks alike.

Can this control refuse an action at execution time, deterministically, in bounded time, independently of the model whose behaviour it governs?

Everything that can is enforcement. Everything that cannot is observation: valuable, often necessary, categorically different. I have used it to sort model cards, standards bodies, semantic classifiers, kill switches and packet recomputation. Every time, the affirmative list turned out to contain only structural things: membership tests, threshold comparisons, provenance checks, undefined transitions. Ring Zero can be trusted precisely because it is too stupid to be persuaded.

Now put the human reviewer through the same sorting.

THE HUMAN APPROVER, TESTED CLAUSE BY CLAUSE

Can refuse an action — YES
At execution time, at machine speed — NO
Deterministically, given identical inputs — NO
In bounded time — NO
Independently of the model being governed — NO
Fails closed when the component degrades — NO
Produces replay-deterministic evidence of its own decision — PARTIAL

One yes, five no, one partial. By my own test, the human at the gate is not an enforcement control. It is an observation instrument that has been handed a switch, and its reliability has never been characterised.

This is not a claim that humans are useless in the loop. It is a claim about category. We have taken a component with unbounded latency, non-deterministic output, a degrading duty cycle and no fail-closed behaviour, and installed it as the load-bearing member of every governance architecture in regulated industry. Then we wrote its name on the incident report in advance.

The reviewer is the only part of the stack we never characterised, and the only part we made responsible for everything.

III. THE CORRELATED VERIFIER

Here is the finding I did not expect when I started writing this, and it is the reason the essay exists.

Sudjianto and Wingyan make an architectural point that has become foundational to how I think about guardrails: when both the system and its verifier reason probabilistically in the same semantic space, the verifier is blind to precisely the errors it exists to catch. A fluent, well-formed, entirely wrong output passes. I have used this repeatedly to argue against LLM-as-judge, against embedding-similarity checks, against semantic classifiers dressed up as controls.

I have been using it to expel machine verifiers from the kernel. I never noticed it applies, without modification, to the human one.

A human reviewer assessing an agent’s output is a probabilistic verifier operating in the same semantic space as the system it verifies. Fluency reads as competence. Structural coherence reads as correctness. A well-formed credit memo with a double-counted EBITDA passes, because it looks exactly like a well-formed credit memo. The reviewer is not checking the arithmetic; the reviewer is checking whether it reads like something a competent analyst produced, which is the one property a language model is optimised to guarantee.

Ethan Mollick’s jagged frontier makes this structural rather than merely unfortunate. Capability is superhuman on some tasks and incompetent on adjacent ones, and the boundary bears no relationship to how difficult the task looks. A model can produce an accurate summary of a paper it has never read and a plausible summary of a paper that does not exist, and the two are indistinguishable to anyone who does not already know the answer. The failure boundary is invisible, so vigilance cannot be allocated proportionally. The reviewer must treat everything with equal suspicion, which is the workload no human sustains past mid-morning.

So the semantic verifier I spent a year expelling from the guardrail layer has been sitting inside the kernel the whole time, with a staff pass and a salary.

We removed the machine judge because it shared a semantic space with the system. We kept the human judge, who shares it too, and gets tired.

IV. THE FOUR INVERSIONS

Szpruch, Sudjianto, Bhatti and Ang dismantle human oversight on three structural grounds: machine speed against human speed, reviewers seeing compressed summaries rather than trajectories, and cognitive overload across multi-step workflows. That is correct and it is architectural. What it does not carry is the empirical literature on what happens to the human over time, which turns out to be worse than the architectural case suggests.

I organise it as four inversions, because in each case the intuitive expectation runs backwards from the evidence, and each one weakens the human at the gate precisely as the machine at the gate gets stronger.

INVERSION I — LOAD
Automation removes the easy work first. It does not reduce cognitive intensity, it concentrates it. The agent handles forty routine cases and escalates eleven hard ones, so the human’s day becomes uninterrupted exception handling. Total hours fall. Difficulty per hour rises to the ceiling and stays there. The organisation books the headcount saving and never notices it has invented a shift no brain can hold.

INVERSION II — CONFIDENCE
Lee and colleagues, across 319 knowledge workers and 936 real tasks, found higher confidence in the AI predicted less critical thinking, while higher confidence in oneself predicted more. Read it as an engineering statement. Every improvement you ship to model accuracy is a quiet reduction in reviewer scrutiny. Reliability manufactures complacency. The control degrades as the system improves.

INVERSION III — PERCEPTION
A randomised trial of sixteen experienced developers across 246 real tasks measured them roughly 19 per cent slower with AI assistance, while they estimated afterwards they had been 20 per cent faster. Thirty-nine points between the clock and the professional judgment. Not novices. Experts, in codebases they had worked in for years.

INVERSION IV — KNOWLEDGE
Around 40 per cent of desk workers report receiving polished, hollow, AI-produced work in a single month, each instance costing close to two hours to untangle. Holweg and Davenport name the compound version knowledge decay: once one person offloads their thinking, the next reasons that if AI is reading it anyway, AI may as well write it. The record against which the reviewer checks is itself degrading.

Inversion III is the one that should frighten a chief risk officer, because it invalidates the instrument. Nearly every enterprise AI programme in the world measures its own effect through self-report: user surveys, perceived time saved, satisfaction scores, adoption telemetry interpreted as value. If professionals are confidently and measurably wrong about their own performance, self-report is not weak evidence. It is anti-evidence, and it is what your business case is made of.

Take the four together and the picture is worse than any one of them. A reviewer carrying maximum intensity per hour, whose scrutiny declines as the model improves, who cannot perceive their own degradation, checking against an institutional record that is quietly rotting, in a semantic space they share with the thing they are checking.

Every one of these curves moves in the wrong direction as the programme succeeds.

V. THE PANEL

The panel convenes because each member does distinct work, and because two of them break parts of this essay that deserved breaking.

LISANNE BAINBRIDGE — Cognitive psychologist, Ironies of Automation, 1983

She arrived four decades early. Automate most of a process and the operator monitors a system they no longer practise operating. The skills that atrophy during smooth running are exactly the skills required when it stops. Her closing irony is the one I cannot put down: the most successful automated systems, needing intervention least often, demand the greatest investment in human skill.

Refines. Everything in this essay is a rediscovery. This is a documented human factors problem arriving in an industry that never read the literature. Parasuraman and Riley called it the out-of-the-loop performance problem in the 1990s. We are not pioneers here, we are late, and forty years of aviation and process control evidence sits unread while we write our eleventh maturity model.

SZPRUCH, SUDJIANTO, BHATTI AND ANG — Scalable Runtime Governance, April 2026

Their three grounds against human oversight are speed, compression and overload: the human cannot act at machine tempo, sees a summary rather than a trajectory, and is asked to hold multi-step workflows in working memory. Their requirement, which I treat as the operational definition of a control, is that governance decisions be expressible as deterministic functions over governed state, in bounded time, independent of the language model.

Anchors the thesis and bounds it. They already proved the human is not a control. What they did not do, because it was not their problem, is say what the human is for once you accept that. My answer, developed below, is that the human is the semantic residue layer, and residue layers need instrumentation and fail-closed defaults, not job titles.

NATALIYA KOSMYNA — MIT Media Lab, cognitive debt, 2025

Her EEG work found the strongest, most distributed neural connectivity in participants writing unaided, moderate engagement using search, weakest using a language model. Participants in the model group struggled to quote work they had produced minutes earlier and reported the lowest ownership of it. Her framing lands: this is cognitive debt, and unlike technical debt there is no credit facility.

Breaks part of the thesis. Fifty-four participants, eighteen in the critical session, an essay task, a preprint. Stanković and colleagues flag sample size, EEG methodology and reproducibility directly, and note the search group offloaded heavily without the same impairment, which complicates the delegation story. I will not build a control framework on a neural connectivity finding and neither should you. Anyone selling brain scans as an enterprise risk metric is selling something.

METR — Becker, Rush, Barnes, Rein, developer productivity RCT

They produced the perception gap, then did what almost nobody in my industry does. In February 2026 they published an update: the follow-up showed some evidence of speed-up, selection effects had become severe because developers who benefit most declined to participate in no-AI conditions, and they were redesigning the study. Revised position, in effect: we do not know.

Breaks part of the thesis and improves it. I cannot use the slowdown as a general claim about AI productivity, and I will not. The finding that survives is the one I actually need: professionals were confidently, measurably wrong about their own performance. That is not a claim about AI, it is a claim about self-report, and it holds regardless of which direction the productivity number eventually settles.

JEFFREY HANCOCK — Stanford Social Media Lab, on workslop

His distinction between workslop and ordinary bad work is the sharpest observation in this literature. Sloppy work still required effort, and that effort was a signal. The signal is gone. Anyone can now generate unlimited plausible, polished, substanceless material at zero cost, and the burden of discovering it is substanceless transfers entirely to the recipient.

Extends past the gate. The reviewer is not only fatigued, they are being fed. Volume requiring review rises at machine speed while capacity to review stays biological. That is a throughput mismatch, and throughput mismatches are solved by architecture. Never by training, never by culture, never by a wellbeing programme.

GLORIA MARK — UC Irvine, attention research

Her field studies find people averaging forty-seven seconds on a screen before switching, and her earlier work put the cost of a genuine interruption at over twenty minutes to full refocus. Set that against a governance model assuming considered human judgment can be summoned on demand, dozens of times a day, in the gaps between everything else.

Quantifies the absurdity. We ask for deliberative reasoning from a working environment engineered to make deliberative reasoning impossible, then treat the resulting click as evidence of oversight in a regulatory filing.

VI. THE CASE AGAINST THIS ESSAY

A position that has not survived its best objection has not been tested, it has only been asserted at length. Four objections. Two I can answer, one partially, one I concede.

Objection 1: this is an expert’s complaint

The deskilling literature is overwhelmingly about experienced practitioners losing edge. The counter-evidence is strong that the largest gains from these systems accrue to less experienced people, who are lifted toward the frontier rather than dragged from it. If AI raises the floor faster than it lowers the ceiling, the aggregate cognitive effect on an organisation could be positive.

Answer. Correct on the aggregate, irrelevant to the argument. I am not making a claim about mean organisational capability. I am making a claim about a specific person occupying a specific structural position: the approver at a gate on a consequential action. That role is staffed by experts precisely because it requires expertise, and it is exactly the population where the deskilling evidence bites hardest. Raising the floor does not help when the failure mode lives at the ceiling.

Objection 2: if the human is not a control, why instrument them at all

Sharpest version, and it goes to the heart of Section II. If I have just proved by my own test that the reviewer is an observation instrument rather than an enforcement one, then instrumenting them is polishing a component that should be removed from the critical path entirely. Build the kernel and stop worrying about who is tired.

Answer. Because the semantic residue does not go away. The kernel governs everything reducible to structure: membership, threshold, provenance, ordering. It provably cannot reach the questions that are irreducibly semantic, and those questions remain, permanently, in human hands. So the correct architecture is not to remove the human but to stop pretending they are Ring Zero: instrument them as the fallible component they are, and surround them with deterministic guards that fail closed when they degrade. That is Section IX, and it is the entire prescription.

Objection 3: the evidence base is thinner than the rhetoric

One preprint with fifty-four participants and a published methodological critique. One randomised trial its own authors have walked back. One survey of self-reported time lost to a phenomenon defined by the people who named it. A forty-year-old paper from process control. A great deal of this essay’s force comes from the accumulation of studies that individually would not survive a careful read.

Answer, partial. Substantially fair, and it is why the two strongest planks are deliberately not the neuroscience. They are the confidence correlation, which is peer-reviewed and published at CHI on 319 workers and 936 tasks, and automation bias, which carries four decades of evidence across medicine, aviation, military and public administration and is named in statute in at least two jurisdictions. Strip out every contested study and those two remain, and they are sufficient. The rest is corroboration, and I have marked it as such rather than laundering it into certainty.

Objection 4: the thesis is commercially convenient

I work on execution governance. The essay concludes that execution governance is the missing layer. Again. Discount accordingly.

Answer. Discount away, and discount symmetrically, because there is no disinterested position in this discourse. What I can offer instead of disinterest is falsifiability. The enforceability test returns an unambiguous answer and anyone can run it on my work. And the specific claim here is testable this week, at zero cost, with data you already hold: pull your approval dwell times and your override rates. If the distributions look healthy, I am wrong about your organisation and you should say so publicly.

A proposal you cannot attack is not a strong proposal. It is an unfalsifiable one, and the two are opposites.

VII. WHAT THE FRAMEWORKS ALREADY CONCEDE

The uncomfortable part is that the regulators got there first and then stopped.

Singapore names it outright

The proposed Guidelines on AI Risk Management address human oversight at paragraph 4.10, and they name automation bias and decision fatigue explicitly. Many frameworks do not. That is an accurate diagnosis in a supervisory document, written before most vendors had noticed the problem existed. What follows the diagnosis is guidance, not mechanism.

The agentic model governance framework goes further still. Version 1.5 proposes measurable indicators for whether human oversight is real: human override rates and human response times during review. That is exactly the right instinct, and it is the first time I have seen a national framework treat oversight as something with a measurable failure mode rather than a box with a name in it. An oversight function that never overrides is not oversight, it is a signature.

And then the same framework concedes the whole game, in its own text: the speed at which agents take decisions makes it difficult for oversight mechanisms to detect and prevent unauthorised actions in real time before they cause harm. That is a national agency writing down that the oversight model it recommends cannot arrive in time.

Europe names it in statute and then defers everything except it

Article 14 does not merely require human oversight. It requires that the assigned person be enabled to remain aware of automation bias, named in the text. Article 26 requires deployers to assign oversight to natural persons with the necessary competence, training and authority. For the most sensitive category, two such persons.

Now the timing, which is the part almost everyone has read backwards. Following the Digital Omnibus agreement, high-risk obligations for standalone systems moved from 2 August 2026 to 2 December 2027, and embedded systems to August 2028. Transparency obligations held their original date. And the AI literacy duty under Article 4 has applied since February 2025 and was never deferred at all.

Every obligation about paperwork moved. The one obligation about whether your people can actually think did not.

Sixteen months of relief landed on documentation, conformity assessment and registration. Zero months landed on competence. If you read the deferral as breathing room, you read it precisely backwards.

So we have three regulators who between them have named automation bias, named decision fatigue, proposed override rate and response time as indicators, conceded that oversight cannot arrive at machine speed, and mandated competence since early 2025. Not one has specified a mechanism, a threshold, a conformance test or a fail-closed default. The diagnosis is complete and the prescription is absent, which is the same shape as every other gap I have written about.

Naming the failure mode in statute and specifying nothing is how a regulator writes an incident report eighteen months early.

VIII. REVIEWER DRIFT

Szpruch and colleagues name a failure mode I regard as the most important conceptual contribution in their paper. Even when no individual trajectory violates policy, the trajectory distribution may drift: shorter paths favoured, one retrieval source leaned on, fewer abstentions. They call it orchestration drift and insist it be monitored separately from data drift and model drift, because it concerns execution behaviour over the transition system rather than inputs or model performance. No single run offends. The distribution moves.

The same thing happens to the reviewer, and nobody monitors it at all.

No single approval violates policy. Every one is defensible in isolation, and each would survive a case-by-case audit. But the distribution moves: dwell time compresses, override rate decays, escalation frequency falls, disagreement with the agent trends toward zero. Call it reviewer drift, and note that it has a direction and a driver. Inversion II tells us it decays as the model improves, because confidence in the system predicts less scrutiny of the system. The better your programme performs, the faster your control erodes.

This is why per-instance compliance checking cannot see it. Every instance complies. Reviewer drift is visible only in the distribution, over time, across a population, which means it requires exactly the kind of telemetry nobody has built.

It is also why the annual control testing cycle is the wrong instrument. Sampling twenty approvals once a year and confirming each was properly authorised will return a clean result in an organisation whose oversight function has completely dissolved, because dissolution is not a violation.

A regime that can only detect violations will not notice itself dissolving, because dissolution is not a violation.

IX. MAKE THE TIRED HUMAN FAIL CLOSED

Here is the prescription, and it is not what a wellbeing consultancy would tell you.

You do not fix the reviewer. You cannot. Fatigue, automation bias and semantic correlation are properties of the component, not defects in the individual, and no amount of training removes them. What you do instead is change what happens when the component degrades.

Today, a fatigued reviewer produces an approval. The system fails open. Exhaustion, distraction, time pressure and complacency all converge on the same output, which is the output the agent wanted, which is why nothing ever looks wrong. Degradation is silent because degradation and success produce identical log lines.

The runtime safeguards work published with eight tier-one institutions contains the fix, and I do not think its significance has been widely noticed. Escalations carry a defined window, and if no decision is made within it, the action defaults to block or escalates to a senior reviewer. Most human-in-the-loop designs specify who reviews and never specify what happens when nobody does. This one specifies the timeout, and the timeout is deterministic.

That single design decision converts the reviewer from a fail-open control into a fail-closed one. Fatigue no longer produces approval. Fatigue produces denial, or produces escalation to someone fresher. The human’s degradation is absorbed by the kernel rather than transmitted through it.

And this is where my own thesis does real work rather than decorating a slide. Long-AND, not Short-OR. Short-OR stops evaluating the moment one branch resolves, which is exactly what a rubber-stamp approval is: the reviewer finds one satisfying condition, the output looks right, and the check terminates. Every remaining condition goes unexamined and the log records a pass. Long-AND evaluates every condition, every time, and fails closed if any fails. A tired human is a Short-OR machine. The kernel’s job is to hold the conjunction that the human can no longer hold.

The question is not whether your reviewer will degrade. It is what your architecture does on the day they do.

X. THE REVIEWER INTEGRITY LAYER

My framing has been Prompts to Loops to Loop Governance: the unit of value moves from the individual prompt, to the automated loop, to the governed loop. The correction I am making is that a governed loop containing an uninstrumented human is not governed. It is supervised in name.

Six instruments. Two of them are not mine and I have marked them so. None requires new science.

  1. Approval dwell distribution
    Time between presentation and decision, per reviewer, per decision class, tracked as a distribution rather than a mean. Set a floor derived from actual reading time for that case type. Alert when the median crosses it. This single metric will tell you more about your control environment than your entire policy library.
    [Extends the response-time indicator proposed in the agentic MGF]
  2. Override rate and direction
    Rejections and modifications tracked as a positive signal, not a friction cost, and watched for decay rather than level. A gate trending toward zero overrides is decaying into a signature. Expect the decay to accelerate as your model improves, and treat that acceleration as the primary reviewer drift alarm.
    [Extends the override-rate indicator proposed in the agentic MGF]
  3. Trajectory presentation, not summary
    Szpruch’s second ground is that reviewers see compressed summaries rather than execution trajectories. Present the path: what was retrieved, what was inferred, which approvals were consumed, what preceded the proposed action. A reviewer shown a conclusion is being asked for a plausibility judgment. A reviewer shown a trajectory is being asked for a structural one, and structural questions are the only ones a human can answer reliably in a shared semantic space.
  4. Calibration probes
    Route known-defective outputs through the live queue periodically and measure catch rate per reviewer, per shift hour. This is the only honest test of whether oversight is functioning, and it is standard practice in every other assurance discipline from airport security to laboratory quality control. If your reviewers cannot catch a planted error at 16:00, they are not catching the real ones at 16:00 either.
  5. Exception density ceilings
    Cap consecutive high-intensity exception decisions per individual, and design deliberate variation back into the shift. Not as wellbeing. As a documented control tolerance with a limit, a monitored parameter and a breach procedure, exactly like any other operating constraint. Then roster protected unassisted work for anyone holding approval authority, because Bainbridge’s remedy is forty years old and still the only one that works.
  6. Fail-closed escalation with replay determinism
    Every escalation carries a timeout defaulting to block or onward escalation. Every disposition is logged such that re-running the recorded envelope against the recorded controls reproduces the recorded outcome. Without that guarantee your audit trail is a narrative of what the system reported having decided, not a reconstruction of the decision, and for a supervisory review or a court those are not the same artefact.

Every one produces evidence, which is the point. Three of them convert a psychological state into a monitored parameter, which is the harder point, and the one the compliance industry has avoided because psychological states do not fit in a document.

Instrument the reviewer with the same seriousness you instrument the model, or admit the reviewer is decoration.

XI. CODA: THE ACCOUNTABILITY SINK HAS A PULSE

I have described the accountability sink in its most refined form: not an absence of names, but an abundance of them, arranged around a hole. Designated control functions, accountable persons, named senior management, board responsibilities, all specified in careful detail around a layer that cannot refuse an action.

This essay adds one thing to that picture, and it is the thing that makes it personal rather than architectural. The names belong to people. When the instrument that names the accountable individual cannot reach the layer at which harm is committed, and the individual it names is a biological component with a duty cycle nobody characterised, working in a semantic space they share with the system they are checking, then we have not merely built an accountability sink. We have staffed it.

That person will be named in the incident report. They will be asked why they approved it. The honest answer, which they will not be permitted to give, is that the architecture was designed so that approval was the only thing exhaustion could produce.

You will not be asked whether your model was right. You will be asked who approved it, and whether they were in any state to do otherwise.

I have written before about The Fork, and about how the dystopian branch is quieter than the scenarios that get attention. It is not machines seizing control. It is machines being handed control, one exhausted click at a time, by people who technically held the authority to refuse and materially did not have the capacity, in organisations that recorded each click as evidence a human was in charge. Nothing dramatic happens. The audit trail stays immaculate right up to the moment it becomes an exhibit.

Long-AND is not only a claim about technical conjunctions. It is a claim that the machine layer and the human layer must both hold, and that neither redeems the other. A perfect kernel with a hollowed-out judgment layer is a well-instrumented system with nobody home. A thoughtful, capable, rested reviewer with no kernel underneath is a person waiting to be blamed for something they were never given the means to stop.

Two December 2027 is not your deadline. Your deadline is the next incident, and it does not appear on any regulatory calendar.

So run the test this week, on the component you never tested. Pull the dwell times. Pull the override rates. Pull the escalation timeouts, if you have any, and find out what happens when nobody responds. If the numbers embarrass you, that is the useful outcome, because right now the thing standing between your agentic estate and a serious failure is a person who has been pressing P to proceed since nine o’clock this morning.

We are extremely good at knowing who approved it. We remain unable to say whether anyone decided.

DECLARED INTEREST

I work on AI Execution Governance, the deterministic enforcement layer for agentic systems in regulated environments. This essay concludes that the enforcement layer must absorb the reviewer’s failure modes, which is convenient for me. Section VI is my attempt to argue against it. The claim is falsifiable with data you already hold, and I have said where to look.

Dr Luke Soon is an HX Architect, Futurist and AI Ethicist. He writes Genesis: Human Experience in the Age of Artificial Intelligence at genesishumanexperience.com and is the author of Synthesis: The Superintelligence Protocol.

Long-AND, not Short-OR.

SOURCES

Bainbridge, L., Ironies of Automation, Automatica, 1983. Parasuraman, R. and Riley, V., on the out-of-the-loop performance problem. Parasuraman, R. and Manzey, D., on complacency and automation bias, Human Factors, 2010.

Leveson, N. and Turner, C., investigation of the Therac-25 accidents, 1985 to 1987.

Szpruch, L., Sudjianto, A., Bhatti, T. and Ang, G., Scalable Runtime Governance for Agentic AI in Financial Services, SSRN 6567199, April 2026. Sudjianto, A. and Wingyan, L., FinstructBench, SSRN 6506403, 2026.

Lee, H. P. et al., The Impact of Generative AI on Critical Thinking, CHI 2025. 319 knowledge workers, 936 first-hand task examples.

Kosmyna, N. et al., Your Brain on ChatGPT: Accumulation of Cognitive Debt, arXiv:2506.08872, 2025. Commentary by Stanković et al., arXiv:2601.00856, 2026, on sample size, EEG methodology and reproducibility.

Becker, J., Rush, N., Barnes, E. and Rein, D., randomised controlled trial of AI tooling on experienced open-source developers, July 2025, and the authors’ February 2026 update on selection effects and study redesign.

Niederhoffer, K., Hancock, J. T. et al., research on workslop, Stanford Social Media Lab with BetterUp Labs, survey of 1,150 US desk workers, September 2025. Holweg, M. and Davenport, T. H., on slopification and knowledge decay, Harvard Business Review, June 2026.

Mollick, E., on the jagged frontier of AI capability. Mark, G., Attention Span, 2023, and Mark, Gudith and Klocke, The Cost of Interrupted Work, CHI 2008.

Monetary Authority of Singapore, Consultation Paper on Guidelines on AI Risk Management, P017-2025, November 2025, paragraph 4.10. Infocomm Media Development Authority, Model AI Governance Framework for Agentic AI, Version 1.5, May 2026. Monetary Authority of Singapore and BuildFin.ai, Safeguards for Agentic Finance at Runtime, July 2026.

Regulation (EU) 2024/1689, Articles 4, 14 and 26. Digital Omnibus on AI, provisional political agreement 7 May 2026, Council confirmation June 2026: Annex III high-risk obligations deferred to 2 December 2027, Annex I to 2 August 2028, Article 50 transparency retained at 2 August 2026, Article 4 AI literacy unchanged and in force since 2 February 2025.

Provenance note. The approval trace is illustrative and uses synthetic data. Adoption and value-realisation figures are drawn from the 2026 global AI adoption survey literature. The randomised trial result should be cited alongside its authors’ own subsequent qualification, not in isolation.


Leave a comment