Singapore has published the most complete agentic AI governance stack in the world. Only one part of it can stop an agent from acting -,-and that is the part nothing requires.
Genesis: Human Experience in the Age of Artificial Intelligence
I. The Set, Not the Documents
Between May 2025 and July 2026, four documents were published in Singapore that, taken together, describe the governance of artificial intelligence more completely than any other national body of work in existence.
The Association of Banks in Singapore, through its Standing Committee on Data Management, published the Handbook on Generative AI Guardrails in Banking in May 2025 ;;; nine guardrails distilled from thirty-plus live deployments across ten member institutions. The Monetary Authority of Singapore issued Consultation Paper P017-2025, the proposed Guidelines on Artificial Intelligence Risk Management, on 13 November 2025; consultation closed on 31 January 2026. The Infocomm Media Development Authority launched the Model AI Governance Framework for Agentic AI at Davos on 22 January 2026 – the world’s first framework written specifically for agents – and updated it to Version 1.5 on 20 May 2026 after feedback from more than sixty organisations. And on 3 July 2026, MAS and eight financial institutions published Safeguards for Agentic Finance at Runtime under the BuildFin.ai initiative.
Each has been read individually. Each has generated its own law-firm client alert, its own compliance checklist, its own consulting offer. Almost nobody has read them as a set.
Read as a set, they do something none of them does alone. They map the entire-control surface of an agentic system – from the shaping of a prompt to the authorisation of a payment – and in mapping it completely, they make visible precisely where the surface has no floor.
II. The Test
There is one question that separates governance from documentation, and it is not is this control robust? or is this control proportionate?
Can this control refuse an action at execution time?
Everything that can is enforcement. Everything that cannot is observation – valuable, often necessary, but categorically different. Observation instruments record, detect, report, and inform. They do not stop anything. A control that can only tell you afterwards that an unauthorised payment was made is not a payment control; it is a payment record.
The distinction is not pedantry. Szpruch, Sudjianto, Bhatti and Ang put it with unusual bluntness in April 2026: the industry’s reliance on prompt-based guardrails constitutes a fundamental failure in risk mitigation, presenting an illusion of control while offering little binding enforcement. Their formal requirement is that governance decisions be expressible as deterministic functions over governed state, executing in bounded time and independent of the language model. Anything failing that test is non-binding, and its residual risk must be explicitly carried.
Apply that test to four instruments, and they sort themselves not by quality but by position — by where they sit relative to the moment an agent acts.
III. Before the Model Speaks — the ABS Guardrails
The ABS Handbook is the earliest and the most operationally grounded of the four. Its nine guardrails — Enterprise Governance and Training; Filtering and Control; Customised Model Design; Red Teaming; Prompt Design; Monitoring and Validation; Human-in-the-Loop Moderation; User Feedback and Iterative Improvement; User Transparency and Consent — are mapped against ten risks and seven use-case categories, with an Excel companion that resolves each guardrail into named controls. It is a genuinely useful document, and its candour about cost is refreshing: Gen AI use cases that are not economically viable to implement responsibly should not be implemented at all.
But its position is unambiguous. Eight of the nine guardrails act before the model produces output or after a human has read it. Enterprise Governance is structural. Customised Model Design alters weights. Red Teaming is pre-deployment. Prompt Design shapes instructions. Monitoring and Validation measures. User Feedback and User Transparency operate downstream of the output entirely.
Only Filtering and Control touches the runtime, and its controls — content filters, relevancy checkers, output verification tools, input quality assurance — screen text. They ask whether the model said something undesirable. They do not ask whether the system is permitted to do what it is about to do. The Handbook is explicit that its scope is the usage and outputs of Gen AI, not the actions of agents.
This matters more than it appears, because several of these controls are themselves probabilistic. Control 2.3 deploys an output verification tool whose role can be filled by the principal LLM or a secondary model. Control 2.5 screens prompts with another LLM. Sudjianto and Wingyan’s finding applies directly: when both the system and its verifier reason in the same semantic space, the verifier is blind to precisely the errors it exists to catch. A fluent, well-formed, entirely wrong output passes.
The ABS guardrails are the layer SAFR would later describe, with some tact, as not a substitute for runtime governance of financial action.
IV. Around the Lifecycle — MAS AIRG
The AIRG is the only one of the four that will carry supervisory force. It is also the one that comes closest to diagnosing the problem correctly and then walks away from it.
Paragraph 1.10 is accurate. It names an agent with access to tools autonomously executing actions misaligned with the institution’s objectives or a customer’s interests, and compromised agents exfiltrating data or executing malicious commands at scale. That is the runtime failure, correctly identified, in a supervisory document, in November 2025.
What follows is a lifecycle control set: AI identification, AI inventory, risk materiality assessment, and then data management, transparency, fairness, human oversight, third-party management, selection, evaluation and testing, technology and cybersecurity, reproducibility, pre-deployment review, post-deployment monitoring, change management.
Run the test across all of it. The residue is thin.
Kill switches appear at 4.4 and 4.23(b), defined in footnote 15 as mechanisms to deactivate AI expeditiously if it exceeds risk tolerances. That is system-level termination, not action-level refusal — a circuit breaker for the building, not a lock on the door. And at 4.23(b) the verb is consider implementing.
Access control at 4.16(b) and least-privilege review at 4.22(a) do bind. But they bind at the identity and infrastructure layer. They answer may this credential reach this system. They cannot answer may this agent dispatch this payment given the approvals recorded on this trajectory.
Human oversight at 4.10 is thoughtful — it names automation bias and decision fatigue explicitly, which many frameworks do not. It is also the control that Szpruch et al. dismantle on three structural grounds: machine speed against human speed, reviewers seeing compressed summaries rather than trajectories, and cognitive overload across multi-step workflows.
Paragraph 4.23(a) is the only clause in the document that reaches an execution trajectory: where relevant, information flow and decision-making paths across workflows that use AI, such as reasoning processes, actions taken and tools used, should also be monitored. One sentence. The verb is monitored. The qualifier is where relevant.
There is a second gap, structural rather than lexical. The risk materiality methodology at 3.10 requires impact, complexity and reliance. Complexity is a property of the model, inherited from classical model risk management, and it tells you nothing about blast radius. Reliance is a partial proxy for autonomy. Nowhere does the methodology ask what the system is permitted to do — read, compute, write, dispatch, commit. Szpruch’s tiering uses five dimensions and puts authority alongside impact as a primary driver. Under AIRG as drafted, a read-only summariser and an agent with payment-dispatch authority can tier identically if they share an architecture.
Question 7 of the consultation asked, in terms, whether other risk dimensions should be included. The invitation was open. Whether it was taken is not yet public.
The AIRG describes the fire accurately and then specifies the building inspection.
V. On the Design — IMDA MGF v1.5
The IMDA framework is voluntary, cross-sector, and intellectually the most advanced of the four. It is also, for anyone who has been arguing that governance controls differ in kind rather than merely in strength, an unexpected ally.
Three moves in Version 1.5 deserve particular attention.
First, it separates action-space from autonomy. Action-space is the range of actions an agent is permitted to take, determined by its tools and the transactions it can execute. Autonomy is the degree to which it decides when and how to act. These are orthogonal, and treating them as one — which most frameworks do, under the single word “autonomy” — destroys the ability to reason about authority at all. IMDA gets this right.
Second, and more consequentially, v1.5 introduces a taxonomy of technical controls by type: structural controls, rule-based controls, and prompt-layer controls, with guidance on selecting among them. That is enforceability class, functioning as a property of the control rather than a footnote to it. A flat enumeration of governance controls masks the fact that a prompt-layer instruction and a structural constraint are not substitutes for one another and never were. IMDA has stopped enumerating and started classifying.
Third, v1.5 folds the enforcement layer into the definition of the agent itself — adding access controls, guardrails, human approvals, logging and monitoring as core components alongside model, instructions, memory, planning, tools and protocols. The kernel is no longer something bolted on afterwards. It is constitutive.
The framework also names cascading effects and unpredictable outcomes in multi-agent systems, adds system complexity and third-party agent use as likelihood factors, and — unusually — proposes measurable indicators for whether human oversight is real: human override rates and human response times during review. That is the right instinct. An oversight function that never overrides is not oversight; it is a signature.
And then IMDA does something almost no framework does. It concedes the gap. The updated framework states that the speed at which agents take decisions makes it difficult for oversight mechanisms to detect and prevent unauthorised actions in real time before they cause harm.
That is a regulator’s agency writing down, in its own document, that the oversight model it is recommending cannot arrive in time. Roman Yampolskiy has argued for years that exhaustive pre-deployment evaluation of a sufficiently capable system is not merely expensive but impossible in principle — the state space cannot be enumerated. IMDA has now conceded the operational corollary: if you cannot enumerate it beforehand and you cannot catch it during, the only remaining option is to refuse it at the boundary.
Which is exactly what the fourth document does.
VI. At the Moment of Action — SAFR
Safeguards for Agentic Finance at Runtime, published 3 July 2026 with Ant International, Circle, HSBC, J.P. Morgan Chase, Manulife, Mastercard, OCBC and Visa, is the only one of the four that passes the enforceability test.
Its architecture is four components joined by a Governance Envelope — a structured record carrying the proposed action, the action trace showing the steps by which the agent arrived at it, and context metadata including identity, mandate and policy constraints.
• Agent Identity resolves the agent against a registry before any other evaluation proceeds. An envelope failing this check is rejected.
• The Controls Repository holds the institution’s rulebook: regulatory requirements, product rules, and user-delegated mandates. Mandates are explicitly capability-based in the Dennis and Van Horn sense — bounded, machine-readable tokens of authority. An agent cannot extend the scope of a mandate through its own reasoning.
• The Disposition Engine evaluates each in-scope action deterministically and returns one of four outcomes: Deny, Escalate, Auto-Execute, Observe.
• The Audit Log records every decision, append-only and tamper-evident, so that events can be reconstructed without relying on the agent’s own account of what occurred.
Deny rejects the action before execution, with a specific reason recorded. That is refusal. Not detection, not escalation-after-the-fact, not a post-hoc flag. This is the deterministic execution kernel, specified — and specified with the eight largest names in agentic payments and banking attached to it.
Two design decisions in SAFR are better than they first appear. The escalation timeout is one: escalations carry a defined window, and if no decision is made, the action defaults to block or escalates to a senior reviewer. Most human-in-the-loop designs specify who reviews and never specify what happens when nobody does. The second is per-action independence: an Auto-Execute at one step confers no authority at the next, because the agent adapts to intermediate results and prior authorisation should not carry forward.
SAFR is the floor the other three instruments are built on top of. It is also, as we will see, the floor that none of them requires anyone to lay.
VII. The Definitional Exclusion
Here is the synthesis. Arrange the four by position and the stack is complete:

The stack is complete in coverage and empty at the centre.
Three instruments govern everything around the moment of action. One governs the moment itself. And that one is disowned twice over.
First disownment. SAFR states plainly that it does not constitute regulatory guidance or supervisory expectations, nor does it prescribe or anticipate future directions for such. It is a reference approach. It is not a managed service. Implementation is left to the deployment context.
Second disownment, and this is the finding that matters. MAS AIRG paragraph 1.2(d), repeated at 3.3(d) of the consultation paper, provides that calculators or tools whose outputs are solely based on predefined programming logic or rules would not be regarded as AI for the purpose of these Guidelines.
SAFR’s Disposition Engine is deterministic rule evaluation against fixed thresholds and categorical constraints. It is, by construction, predefined programming logic. It is therefore not AI under the instrument that will carry supervisory force.
Singapore has specified the enforcement kernel for agentic finance, published it under the regulator’s own copyright, attached eight tier-one institutions to it — and placed it outside the supervisory perimeter by two independent mechanisms: an explicit disclaimer in the specification, and a definitional exclusion in the binding instrument.
The only component in the entire stack capable of saying no is governed by nobody’s mandate.
This is not a Singaporean failing. It is a structural artefact of how AI governance has been built everywhere. Every framework in the world defines its scope as the probabilistic thing — the model that learns and infers — and in doing so writes the deterministic thing out of scope. But the deterministic thing is the only part that can enforce. We have built a body of regulation that reaches everything except the layer where refusal happens.
VIII. The Probabilistic Layer Returns
The kernel has a second problem, and it is internal.
SAFR’s Controls Repository distinguishes generic controls — authorisation checks, exposure limits — which are typically deterministic, from AI-specific controls such as evidence quality and envelope integrity checks, which may involve probabilistic or semantic assessment.
The Disposition Engine is deterministic in its evaluation. But determinism in the evaluation of a probabilistic control does not yield a deterministic outcome; it yields a reproducible reading of an unreliable instrument. And the paper’s own case studies confirm the concern is not theoretical: Manulife’s implementation validates outputs through content filtering, retrieval controls and LLM-as-a-Judge evaluation.
The semantic verifier, expelled from the guardrail layer by Szpruch et al. and by SAFR’s own framing, has re-entered through the Controls Repository. Sudjianto and Wingyan’s architectural point stands wherever it is placed: a verifier operating in the same semantic space as the system it verifies is blind to precisely the class of error it exists to catch. Moving it inside the kernel does not fix it. It makes it load-bearing.
If Ring Zero is to be deterministic, it must be deterministic all the way down, and every probabilistic control admitted to the Controls Repository must be explicitly classified as advisory with its residual risk carried separately — exactly the discipline Szpruch et al. prescribe and exactly the discipline IMDA’s structural/rule-based/prompt-layer taxonomy makes expressible.
IX. What None of the Four Does
Four instruments, and three requirements that none of them closes.
Envelope origin authentication. SAFR identifies the problem with real intellectual honesty: the action details and the action trace are both agent-declared contents of the same envelope, and both can be fabricated together by a sophisticated adversarial injection that maintains internal consistency while departing from the original task. Its conclusion is that the envelope must therefore be treated as a document to be authenticated against its origin, not merely as a record of what the agent reported.
No mechanism is specified. That single unstated mechanism carries the integrity of the entire kernel: if the envelope can lie coherently, everything downstream of it is theatre performed on false evidence.
Trajectory and path conformance. SAFR’s per-action independence is correct authority hygiene and a blind spot for compositional risk. Every action in a sequence may be individually permitted while the sequence itself violates policy — approval obtained for one purpose and consumed for another, a limit respected per-transaction and breached in aggregate, a path through a state machine that no single transition forbids. Szpruch et al. specify temporal and path conformance checking over trajectories. SAFR does not carry it. IMDA names cascading effects and emergent multi-agent behaviour at the risk level and monitors them nowhere at the enforcement level. Orchestration drift — the migration of the trajectory distribution over time while no individual trajectory violates policy — is diagnosed across the stack and enforced against by nothing in it.
Replay determinism. The Audit Log records the envelope, the mandate, the rules applied, the outcome and the basis. Nothing requires that re-running the recorded envelope against the recorded controls reproduces the recorded disposition. Without that guarantee, the audit trail is a narrative of what the system reported having decided, not a reconstruction of the decision. For a supervisory review, a dispute, or a court, those are not the same artefact.
X. Named, and Still Powerless
There is one more document in the Singapore set, and it changes the register of everything above.
In May 2026, IMDA published a thirty-six-page discussion paper on Legal Responsibility for AI Agents, drawn from a working group of more than twenty members of Singapore’s legal community. Its finding, in outline: AI agents are not legal persons and cannot be agents in the legal sense; existing negligence and contract frameworks may be adaptable but face serious practical difficulty given the number of actors and the opacity of the value chain; Singapore’s product liability regime does not reach AI. It echoes the Law Commission of England and Wales, which observed in July 2025 that scenarios exist in which no natural or legal person may be liable for harm caused by an autonomous system.
Now set that alongside what the AIRG requires. Designated control functions for AI identification, for the inventory, for risk materiality assessment, acting as final arbiter. An appropriate accountable person designated for ongoing monitoring and incident management. A senior management member named as responsible for AI oversight. Board members held to explicit responsibilities.
Naming is necessary. It is not sufficient. When the instrument that names the accountable individual cannot reach the layer at which the harm is committed — because that layer is definitionally not AI, and the specification that describes it is explicitly not a supervisory expectation — then the named individual absorbs consequence for a system they were never given the means to stop.
This is the accountability sink in its most refined form: not an absence of names, but an abundance of them, arranged around a hole. Eric Schmidt’s observation about regulatory arbitrage applies with a twist here. The arbitrage is not between jurisdictions. It is between the binding instrument and the voluntary one, within a single jurisdiction, on a single question.
A control function that cannot refuse an action is not a control function. It is a name on an incident report that has not yet been written.
XI. The Missing Rung
None of this is a criticism of Singapore. The opposite. No other jurisdiction has produced four instruments of this quality inside fourteen months, and no other jurisdiction has produced the fourth one at all. IMDA has classified controls by enforceability class before any regulator asked it to. MAS has convened eight of the largest institutions in agentic payments and shipped a runtime specification. The Future of Finance Institute is now standing up pilots and a sandbox, with expressions of interest open. That is a functioning innovation-and-governance loop, and it is genuinely ahead of the field.
The point is what the loop has not yet closed.
Read the four in order and they form a dependency-ordered stack, not a flat list. ABS shapes what the model produces. AIRG governs the lifecycle around it. IMDA governs the design of authority and autonomy. SAFR governs the moment of action. Each depends on the one below it holding. And the bottom rung — the one everything rests on — is specified, voluntary, definitionally excluded from supervision, and left to each institution to implement in its own infrastructure, with no conformance regime, no reference implementation, and no test of whether what was built matches what was specified.
The gap has moved. Twelve months ago the argument was that nobody had named the execution layer. That argument is finished; Singapore has named it, and named it well. The argument now is narrower and harder: specification is not implementation, and conformance to a specification nobody is required to meet is not governance.
The enforceability test remains the only diagnostic that matters. Run it on your own stack. Not on the policy, not on the framework, not on the committee — on the control itself, at the moment of action, with a payment about to leave.
Can it say no?
If the answer is anything other than yes, you do not have a control. You have a record of what happened.
Dr Luke Soon writes on AI governance, agentic AI safety, and the future of work at GenesisHumanExperience.com. His thesis is Long-AND, not Short-OR: humanity and AI must flourish together, not one at the expense of the other.
Sources
• Association of Banks in Singapore, Standing Committee on Data Management, Handbook on Generative AI Guardrails in Banking, May 2025 (a further ABS SCDM release was reported on 24 March 2026; readers should verify which edition applies).
• Monetary Authority of Singapore, Consultation Paper on Guidelines on Artificial Intelligence Risk Management, P017-2025, 13 November 2025. Consultation closed 31 January 2026; final Guidelines pending as at August 2026, with a proposed twelve-month transition after issuance.
• Infocomm Media Development Authority, Model AI Governance Framework for Agentic AI, Version 1.0, 22 January 2026; Version 1.5, 20 May 2026 (updated 5 June 2026).
• Monetary Authority of Singapore and BuildFin.ai, Safeguards for Agentic Finance at Runtime (SAFR), White Paper Version 1.0, 3 July 2026.
• Infocomm Media Development Authority, Legal Responsibility for AI Agents, discussion paper, May 2026.
• MindForge Consortium, AI Risk Management: Executive Handbook and Operationalisation Handbook; MAS AI Risk Management Toolkit, 20 March 2026.
• L. Szpruch, A. Sudjianto, T. Bhatti and G. Ang, Scalable Runtime Governance for Agentic AI in Financial Services, SSRN 6567199, April 2026.
• A. Sudjianto and L. Wingyan, FinstructBench, SSRN 6506403, 2026.
• J. B. Dennis and E. C. Van Horn, Programming Semantics for Multiprogrammed Computations, Communications of the ACM 9(3), 1966.


Leave a comment