The Return on Not Yet Knowing

I. The Gate That Eats the Returns

A project can work, teach you something real, and still be killed at the quarterly review.

I want to start with the moment the whole argument turns on, because I have watched it happen and I suspect you have too. A team ships an AI pilot. It does something genuinely new. It reads the messy inputs a rule engine choked on, drafts the memo, catches the exception a junior would have missed. Then the quarterly review arrives, someone asks what it returned this quarter, the honest answer is not much yet, in cash, and the project is closed. The headcount is reassigned. In the minutes it is filed under prudence. That decision is one of the most expensive mistakes in enterprise technology right now, and it is expensive precisely because it looks so responsible.

Here is the provocation, stated plainly. The chief financial officer who demands that an AI project pay back inside two quarters is not being careful with the firm’s money. He is forbidding the only activity that would ever make the money come back, and then writing the cancellation into the minutes as evidence of rigour. The gate that is supposed to protect returns is eating them.

I want to be precise about the claim, because a looser version of it is everywhere and I am not making it. I am not saying returns are merely slow and everyone should be patient. That take is a decade old, correct, and boring. My claim is narrower: the instrument most firms use to impose discipline, a near-term returns gate, forecloses the one activity that produces the return, and there is a different instrument that imposes real discipline without foreclosing it. The gate is not too strict. It is aimed at the wrong thing. First, the numbers, because the setup has to be honest before the turn can be earned.

II. What the Numbers Actually Say

Read the surveys in order and one line separates them: value in quarters, audited profit in years.

Start with the horizon, because everything downstream turns on it. Deloitte spent 2025 asking the AI return question at scale, across a sample of 1,854 leaders, and put the modal payback at two to four years. Only about one respondent in sixteen reports a return inside twelve months; even among the cohort these firms called their most successful, only around one in eight clears the bar inside a year. Set that against the seven to twelve months a firm will happily grant an ordinary piece of enterprise software before it expects the needle to move, and the mismatch is not subtle. We are holding a normal-technology stopwatch to something that does not behave like normal technology. Deloitte’s own framing names the mechanism rather than the mood, conceding that “not all returns are immediate or financial”, and roughly two-thirds of the sample have quietly reclassified AI as strategy rather than a line item that must clear a hurdle rate now.

Then the realisation gap, where the gate does its damage. IBM’s Institute for Business Value ran a parallel study across 2,000 chief executives and found the same shape from the other side. Its headline is blunt: “only 25% of AI initiatives have delivered expected ROI over the last few years”. Yet roughly five in six of those same executives expect positive returns by 2027, and only about a sixth have scaled AI enterprise-wide. The expectation is not that the return fails to come. It is that it comes later than the review cycle that judges it, and a gate calibrated to two quarters reads the 2027 line as a broken promise and cuts before the promise can be tested.

Now the casualties, the part most people miss. S&P Global Market Intelligence reported that in 2025 the share of firms abandoning most of their AI initiatives rose sharply, to around two in five, up from about one in six the year before, with close to half of all proofs of concept scrapped before they reached production. The tidy reading is that the projects deserved it, that the market is clearing out what did not work, and some of it is exactly that. But a returns gate set at two quarters will also abandon projects that were learning on schedule and had simply not yet been given the years their payback curve requires. The abandonment number cannot tell you which is which, and neither can the firm doing the abandoning, because the thing being cancelled, accumulated learning, is precisely what the gate does not measure.

I will not overreach on the sharpest figure, because it does not survive the overreach. MIT Project NANDA’s 2025 working paper is widely quoted for the finding that around ninety-five per cent of organisations report no measurable profit-and-loss impact from generative AI, with only about five per cent of integrated pilots extracting real value. That number has been screenshotted into a thousand posts as proof that AI projects fail. It is not that. No measurable impact is not the same as failure; the paper attributes the gap to weak learning and integration, not to the technology refusing to work. It is a preliminary v0.1 paper out of the MIT Media Lab, with small non-probability samples, and it belongs to Project NANDA, not to MIT at large. It means little on its own, and I carry it only beside S&P’s abandonment rise and the counter-evidence below, for the modest thing it can bear: the value is not yet landing on the accounts, and the cause looks like a learning gap.

McKinsey’s 2025 state-of-AI survey, across roughly 2,000 respondents, resolves the picture. Adoption is almost universal, near seven in eight firms using AI somewhere, while only about two in five report any earnings impact, and for most of those it sits under a twentieth of EBIT. The line that matters is where the value sits: it tracks how far a firm has redesigned the workflow around the tool, not how much it spent on the tool.

I promised honesty, so here is the counter-evidence at full strength, in the same breath, because the essay is not worth reading if it hides it. The Wharton 2025 adoption report reaches a more optimistic headline than anything above: “Three out of four leaders see positive returns on Gen AI investments.” That is a serious result and I will not explain it away. But read the horizon attached to it, and those same leaders overwhelmingly expect the fuller payoff over two to three years, not two quarters. And the studies with genuinely fast, sub-year payback tend to be vendor-adjacent, from Google Cloud, IDC, Forrester and the like, and they define return loosely, self-reported productivity and time saved, drawn from sponsor-friendly samples. Hold the definitions constant and the optimistic and pessimistic camps stop contradicting each other and start agreeing on a shape. Value, as productivity and capability and learning, shows up inside quarters. Audited profit at enterprise scale shows up across years. Deloitte, IBM, Wharton and McKinsey are not four disagreements. They are four instruments reading the same object at four resolutions, and the object is a return that arrives as learning first and cash later.

III. Learning Is the Return

The thing the gate cancels is not the profit; it is the apprenticeship that becomes the profit.

If value tracks workflow redesign rather than tool purchase, the early return on an AI project is not a number in the ledger. It is a stock of knowledge: where the model helps and where it quietly harms, which parts of the process to rebuild around it and which to wall off, what the failure modes look like before they reach a customer. The pilot that read the messy inputs taught you where your data is dirty. The agent that drafted the memo taught you which judgements you were willing to delegate. That knowledge is the asset. It compounds, and it is invisible to an instrument that only reads cash. The firms that get nothing from AI are not the ones that spent too little; they are the ones that never closed the learning and integration gap. The profit is what the asset produces once the workflow is rebuilt around it, and the rebuilding is slow, human and full of dead ends, because you cannot design the new process until you have watched the tool succeed and fail against your own operation.

This is why the gate is self-defeating, not merely impatient. Impatience waits badly for a return that is coming. A two-quarter gate does worse: it cancels the apprenticeship partway through, books the cancellation as a saving, and destroys the accumulated learning that was the only asset the project had produced, recording the destruction as discipline. The usual defence, that capital should flow to the things that pay, is correct in general and wrong here, because it measures the wrong quantity at the wrong time and then acts on the reading. This is the same shape as the Frozen Workforce problem I have written about before: a capability that exists but cannot move, because the surrounding system will not let it. The gate freezes it. If I stopped here I would have written the tired essay everyone else has written, and the distinctive claim is the next one, and it is not about returns.

IV. Why You Cannot Afford to Learn Without a Bound

Reckless and affordable experimentation differ by exactly one bound, and it lives below the business case.

Here is the turn. Telling a firm to be patient with AI is useless advice unless you can answer the question the CFO is really asking underneath the returns question. He is not only asking when this pays back. He is asking how much it can hurt him while he waits, and that is a fair question, because agentic AI does not fail the way a dashboard fails. When an agent goes wrong it does not merely underperform. It acts. It calls a tool, moves a record, releases a document. The downside of a spreadsheet is that it is wrong. The downside of an agent is that it is wrong and then does something about it, at machine speed, before the quarterly review it was going to be judged at ever arrives.

So separate two things that usually get run together. Whether an experiment pays is a business-case question, answered in years. Whether it can hurt you is an execution-layer question, answered the moment the agent tries to act, and it does not wait on the first. This is the thing the patient-capital crowd never addresses. You cannot ask a firm to keep experimenting through a two-to-four-year learning curve if every experiment carries an uncapped tail risk. Nobody sane signs up for years of open-ended downside on the promise of eventual learning. The real precondition for patience is not faith. It is a cap on the fall.

And here is the claim I will stake the essay on. The ability to fail cheaply is not a property of the business case. It is a property of the execution layer. Reckless experimentation and affordable experimentation are the same activity separated by exactly one thing: whether the downside is bounded at the moment the agent acts. Not a policy in a slide deck, and not a model told to behave, but a constraint enforced at the tool-call boundary, at execution time, per system, that refuses the harmful action regardless of what the model was trained to want or what an injected instruction talked it into. The research community has a clean published example in CaMeL, a capability bound applied where the agent actually acts, deterministic, independent of the model’s own goals. The point is not to make agents behave. It is to make their misbehaviour cheap, and cheap failure is the condition under which patient experimentation becomes affordable rather than reckless.

Let me declare my interest, because the argument now runs straight into what I do for a living. I build and I sell exactly this kind of execution-layer bound, a deterministic enforcement layer that contains what an agent can do at runtime. So read what follows as an argument to be attacked, not a neutral finding, and discount the conclusion for motive. But the argument does not depend on my product existing. It depends on the claim that safe experimentation is a property of the execution layer and not the business case, and that claim is testable without me: if an agent’s downside can be contained by the business case alone, with no execution-time bound, then I am wrong and a returns gate is a reasonable instrument after all. I do not think the evidence points that way, but that is the seam to attack, and I would rather hand you the knife than hide it. Governance, read this way, is not the tax you pay on learning. It is the permission slip that lets it happen at all. Take the bound away and the CFO is right to be terrified. Put it in, and the same experiment becomes something a prudent firm can run a hundred times.

V. Not an Excuse for Burning Money

Put the bubble case at full strength: the capex has run miles ahead of the revenue, and hoping is not a plan.

There is a reading of all this I want to refuse out loud, because it discredits the argument if it stands. The reading is: returns take years, so stop asking about returns, keep spending, the profit will arrive. That is not my argument. It is the bezzle, the comfortable interval where the money is already gone but nobody has admitted it, and I want no part of it. So let me steel-man the sceptic pole at full, uncomfortable strength before I say where I part from it.

The spending is real and enormous. Stanford’s HAI index puts global corporate AI investment above 581 billion US dollars in 2025, more than double the year before. Against that, the revenue actually booked against AI products is a fraction of what the spending implies, which is the substance of Sequoia’s widely cited framing, the several-hundred-billion-dollar gap between the capital being deployed and the revenue that would justify it. Voiced at full strength by the bubble camp, the case is not easily dismissed: infrastructure built miles ahead of demonstrated demand, a large share of pilots showing no profit impact, abandonment rising fast, and a horizon that keeps receding to next year, which is exactly what a bubble sounds like from the inside. If the returns never arrive at scale, the patient-learning story is just the sound a bad investment makes while it waits to be written down.

I take that seriously, and it does not move my thesis, because my thesis is not that the returns are guaranteed. It is a claim about which discipline is the right one, and it cuts against the spenders as hard as against the gate. The capex-versus-revenue gap is a real reason for discipline. The question is what kind. A returns gate forbids trying: it kills the experiment before the learning that might have closed the gap can happen, and mistakes the killing for control. A containment bound makes trying cheap: it lets the experiment run and caps what it can cost when it fails, so you can find out fast, and cheaply, which projects actually pay and which should be killed on the evidence rather than on the calendar. If you believe, as I do, that a great deal of this spending is being wasted, you should want the instrument that surfaces the waste cheaply, not the one that cancels good and bad projects together at the same quarterly gate. The bound is not a licence to burn money. It is what lets you stop burning it on the projects that deserve to die, without extinguishing the ones that were three quarters from paying. The sceptic and I agree on more than either of us agrees with the keep-spending camp. We both want discipline, both distrust the receding horizon, both want waste killed. We disagree only on where the discipline sits: at the business case, gating whether you may try, or at the execution layer, bounding what trying can cost.

VI. The Case Against This Essay

If this argument is going to fail, here is the seam where it fails, marked for you.

Five objections, put as well as I can, because an argument that cannot be attacked is a gesture.

First, the load-bearing one: the setup is a crowded take. That AI returns arrive slowly, that learning takes time, that firms are too impatient, has been written a hundred times, and I concede I have added nothing to it. If the essay is judged on the slowness claim it is worthless. It should be judged only on the enforcement claim, that cheap failure is a property of the execution layer and not the business case. Knock out the slowness and that claim is untouched; knock out that claim and the essay is just another patience sermon.

Second, the counter-evidence is real and I have leaned on horizon to reconcile it. Wharton finds three in four leaders already seeing positive returns, and the vendor studies report sub-year paybacks. If those fast returns are real and general rather than loosely defined and selectively sampled, the gate forecloses little and my premise weakens. I have argued the definitions do not hold up, but arguing that the inconvenient studies have bad definitions is exactly what someone protecting a thesis would do.

Third, the sceptics may simply be right that the returns are not coming at scale. If the capex-versus-revenue gap never closes, no amount of cheap, well-bounded learning conjures a return the underlying economics cannot support. A containment bound makes a failed experiment cheaper. It does not make a bad bet good.

Fourth, my sharpest supporting figure is my weakest evidence. The NANDA ninety-five-per-cent number is a preliminary working paper with small, non-probability samples, and it measures the absence of a booked impact rather than a failure. A hostile reader could say I want it both ways, citing it when it dramatises the gap and disowning it when it overreaches. Fair. I have used it only in company and only for its mechanism; strike it entirely and the S&P and McKinsey evidence still carries the argument.

Fifth, and most directly, I sell the instrument this essay concludes toward. Every step that makes the execution-layer bound look like the right instrument also describes what I build and would like you to buy, and you should discount accordingly. The honest reply is that the mechanism does not depend on me: the logic of a deterministic bound at the tool-call boundary holds whether you buy it, build it, or borrow it from the open literature. But the motive is real, and I have put it on the page rather than in a footnote.

VII. Coda

You cannot gate your way to a return that only learning produces; you can bound your way to affording the learning.

The argument reduces to two layers firms keep collapsing into one. Whether an AI project will pay is a business-case question, resolved over years as workflow redesign and accumulated learning turn into audited profit. Whether it can hurt you while you find out is an execution-layer question, resolved in the instant the agent acts. A returns gate answers the first at the wrong time and uses the answer to foreclose the second. A containment bound answers the second at the right time and leaves the first to resolve on its own schedule.

I have written before about the way the durable thing is the constraint and not the forecast drawn on top of it, in The Map Expires Not the Kernel. This is that shape once more. The map of where the value sits keeps expiring; the kernel that contains the downside does not. The CFO who demands payback in two quarters is not being disciplined. He is banning the only activity that produces the payback, and recording the ban as prudence. The discipline he wants is real, but it lives one layer down, in a bound that makes failure cheap, and that bound is the precondition for being allowed to learn at all. That is the return on not yet knowing. You do not collect it by gating. You earn it by bounding the downside, and then letting yourself try.

Leave a comment