August 10, 2026
By: Deepak Dastrala
Every institution now buys the same intelligence. The advantage has moved to the one thing only you can build.
Some months ago I sat in a review at a large bank where a system we had built was performing well above the threshold the business itself had set. Nobody disputed that it worked. It did not ship, because the second line could not establish which version of which policy had driven a given conclusion, on a given date, for a given customer.
The model had never been the problem. The institution could not produce evidence about its own rules at the speed its own technology now demanded. That meeting is the industry in miniature.
The argument
Frontier intelligence has become abundant and nearly uniform. Any large financial institution can buy roughly the same reasoning capability as its competitor, in the same quarter, at a price that keeps falling. Whatever advantage exists there is temporary.
Knowledge in a form that intelligence can use has not become abundant. Most institutions now carry a large capability surplus alongside an equally large knowledge deficit, and the deficit binds.
The determining metric of the coming decade is therefore not how many models you have deployed or what you have spent on them. It is what proportion of your decision surface is backed by knowledge that is engineered, current and traceable. I call that the knowledge coverage ratio, and I think it is the most predictive figure an institution can put in front of its board.
The Evidence
MIT’s NANDA initiative reported in 2025 that around ninety five per cent of enterprise generative AI pilots produced no measurable effect on the profit and loss statement. Roughly eighty per cent of organisations explored tools, twenty per cent ran pilots, and five per cent reached production with measurable impact. The authors were explicit that the divide came down to how far the tools were integrated into the way organisations actually work, rather than to model quality or regulation.
The headline figure has since been challenged on methodology, and some of that criticism lands. Anyone quoting ninety five per cent as a measurement rather than an indication is overreaching. The direction survives, and so does the diagnosis.
My colleagues and I have taken this question into more than two hundred conversations with people in your seat, across banking, capital markets and insurance, and catalogued sixty four friction points that stop enterprise AI reaching production. Only a small minority are model problems. Almost nobody says the system was not clever enough. Nearly everybody says the same four things: the output could not be traced, the reasoning could not be audited, the underlying knowledge was stale or incomplete, and nobody could sign the risk memo.
Two samples, one academic and one commercial, point at the same constraint.
Knowledge is not data
Ask a bank what it owns, and data appears early in the answer. Data is what institutions say when they mean knowledge.
Institutional knowledge sits in three states, none usable by a reasoning system as found. It sits in documents written for humans who already carry twenty years of context, full of implicit reference and inherited exception. It sits in systems, in four decades of encoded logic whose original authors have retired, executable but no longer legible, so that the institution runs on rules it cannot read. And it sits in people, in the credit officer who knows which exception the committee will actually accept, which is the most valuable state and the least durable, because it leaves the building every evening and one day does not return.
All three are raw material. None of that material appears on your balance sheet, and none of it sits in your AI budget.
The arithmetic
Output accuracy in an enterprise AI system is not a property of the model. It is the product of three terms. Knowledge is the facts: what the institution holds as its own governed truth, connected and citable. Reasoning is the engine: how the model derives an answer from what it has been given. Context is the situation: whether retrieval and live signals put the right material in front of the model for this client, this decision, this moment.
Knowledge × Reasoning × Context. The formulation came out of a review run in late 2025 by Arun Jain, chairman and managing director of Intellect Design Arena. He had been working back through the root causes of knowledge accuracy failures across our enterprise AI products, and those products had little in common: different domains, different models, different teams. The failures did not vary with any of that. What they varied with was the state of the knowledge underneath them.
The equation compresses that finding, and its entire argument sits in the multiplication signs. Read additively, it says a strong enough model can compensate for weak knowledge. Read as written, it says nothing compensates for anything. We were slower than we should have been to see what followed from that. Almost every strategic error I describe in this article begins with the first reading.
Take three figures that would pass any steering committee without comment. Knowledge at 85 per cent, reasoning at 95 per cent, context at 90 per cent. Nothing in that row looks like a problem. The system runs at 73 per cent.
In production there is no mostly right. There is trusted, and there is compromised. A regulated institution needs something close to 94 per cent before a decision can be automated without a human reading every output, and 73 per cent is not a near miss. It is a system that manufactures review work rather than removing it. Eighty per cent accuracy has no value at all.
Now watch where the money goes. Most of it flows into reasoning, already the strongest of the three. Move reasoning from 95 to 97 per cent and the system reaches 74.2. Move knowledge from 85 to 95, which is an engineering programme rather than a purchase, and it reaches 81.2.
Neither clears the bar, and here is the part that should change how you fund this work. Hold reasoning at 95 and context at 90, then make knowledge perfect. Not better. Perfect. The system reaches 85.5 per cent, and it stops there. The enterprise threshold cannot be reached by fixing one term however completely, because the other two impose a ceiling on it. Clearing 94 requires something near 98 on all three at once.

You do not experience 73 per cent as a number. It arrives as a system that demonstrated well in March, entered review in June and is still in review in December, with the budget consumed, the team dispersed and the business case quietly withdrawn. A knowledge deficit rarely announces itself as a knowledge deficit. It presents as delay, and delay is charged to the technology rather than to the ground it was built on.
The weakest term governs the outcome, and in nearly every institution I have seen the weakest term is not the model. The point gets missed because the model is the term with a vendor attached.
The objection
The serious counterargument runs like this. Context windows are vast and getting cheaper. Agents read policy corpora directly at inference. Retrieval improves with each model generation. So is knowledge engineering not scaffolding that the next generation quietly absorbs, leaving the firms that invested in it holding an obsolete asset?
Part of this is right, and conceding it precisely is what makes the rest load-bearing. Much of the mechanical work of knowledge engineering today is transitional. It is already being absorbed into the models and the tooling around them, and it will keep getting cheaper. Any institution building a large permanent team around those activities is building the wrong team.
What does not get absorbed is this. A larger context window lets a model read all your policies. It does not tell the model which of two contradictory policies governs, because your institution never decided. It does not supply an effective date nobody recorded. It does not resolve that customer means one thing in origination, another in the risk model and a third in the regulatory return, because that divergence is not a retrieval failure but an absence.
A model cannot retrieve a decision the institution never made.
There is a second limb, and in regulated industries it matters more. Suppose a capable enough model could infer the correct resolution of a policy conflict. An inferred resolution is not an evidenced one. Supervision does not ask for a system that reaches the right answer. It asks for an institution that has decided, on the record, with a named owner and a date, and can show that its systems applied that decision consistently. Capability does not touch this, and no amount of it will.
The volume of this work will fall, probably sharply. Its criticality will rise, because what remains once the mechanical layers are absorbed is the irreducible part: deciding what your institution means, and proving it.
The ratio
KRC tells you whether a system’s output can be trusted today. It does not tell you whether the knowledge term can be raised at all, and that is the question a board has to fund. For that you need a second measure.
Banks know how to express a complex condition as a single governed ratio, agree how it is calculated, and hold themselves to it. Liquidity was once an argument. Then it became a number and the argument ended.
The knowledge coverage ratio has two components. Coverage: the proportion of decisions you intend to automate that are supported by knowledge explicitly engineered, with defined concepts, resolved relationships and stated exceptions. Currency: the proportion of that knowledge verifiably current, with a known effective date and a known successor. Multiply them. Sixty per cent coverage at seventy per cent currency gives forty two per cent, and that figure forecasts next year’s deployment rate better than any pilot count.

Three rules keep it honest, and this is where most measurement attempts fail.
The unit is the decision, not the use case. A decision is any point where the institution takes a position it would have to defend to a customer, a court or a supervisor. Counting use cases lets you claim credit for scope.
Partial coverage counts as zero. If a decision draws on eleven rules and ten are engineered, that decision is uncovered. Averaging across the eleven produces a number that flatters you exactly where the risk sits.
The second line scores it. Every institution reading this already knows from capital adequacy what a self-assessed ratio is worth.
Ask for the number at your next technology review. The argument you get back about why it cannot yet be calculated will tell you more than any demonstration on the agenda that day.
Governance is throughput
Twenty years of technology programmes have trained financial institutions to treat governance as the function that arrives late and slows things down. Applied here, the instinct is expensive.
What blocks deployment is rarely accuracy. It is the inability to demonstrate accuracy to someone with the authority to accept the risk. The bottleneck is evidentiary.
Evidence production capacity is therefore deployment capacity. Institutions that can evidence their systems ship them. Institutions that cannot are left producing pilots that pass every technical test and fail the only one that counts, and they conclude, wrongly, that the technology is not ready.
Designed in from the start, governance produces what clears the queue as a by-product of normal operation: lineage for every assertion, a version for every rule, evaluation history, review points, an explanation a supervisor can follow without a data science degree. Added afterwards, it becomes a nine-month remediation programme. The institutions moving fastest are not those with the loosest controls but those whose controls emerge automatically from how their systems are built.
Whose problem is this?
A central bank does not make loans. It holds reserves, sets the standard, clears between participants and guarantees that what circulates can be trusted, which is what allows everything above it to compound safely. Institutional knowledge needs that layer, and almost no firm has built one.
Someone must hold the authoritative version of what the firm’s terms mean, arbitrate when risk and commercial disagree, guarantee currency, and clear that knowledge to every system and agent that needs it. Without it, each function engineers its own private version, the versions diverge, and the divergence surfaces later as a control failure nobody can trace to its origin.
This will not sit with the chief information officer, and giving it to them is the most common way the problem gets postponed. The work requires someone to rule that the risk function’s definition of exposure governs, or the commercial function’s definition of customer does, and technology leaders do not have the standing to make such a ruling hold. It is a chief executive’s appointment, made once, in public, with the authority stated rather than implied.
This Monday
Three things, none of them a purchase.
Ask for your knowledge coverage ratio, and keep asking until the number exists. It will be worse than you expect, and knowing how much worse is the whole point.
Fund knowledge as infrastructure rather than as a project, with the permanence and budget treatment you give the core platform. This work has no natural owner on most organisational charts, which is why it goes unfunded until it is rediscovered as a crisis.
Make evidence a design requirement rather than a reporting obligation, so that systems produce their own audit trail while they run.
By 2028
I would rather be specific and wrong than general and safe, so here is what I expect.
Supervisors will start asking about knowledge currency directly. Not model documentation, which they ask about already, but questions of the form: which version of which rule governed this decision, and who owned it. Once one regulator puts that question in an examination, the others follow, and every institution finds out in the same quarter whether it can answer.
The org charts will move before the technology does. The title will vary. The function will not.
And the institutions that scale furthest will not be the ones that spent most. I expect several of them to have spent conspicuously less on models than peers still stuck in pilot.
That last one is where I can lose, and it is the one to hold me to. If AI outcomes across this industry turn out to line up with AI budgets, then the mechanism I have described is not the mechanism, and this article is wrong at its foundation.
Why this compounds
Financial institutions became powerful by working out how to make capital compound, and then building the governance that made compounding safe enough to scale. Knowledge behaves the same way with one additional condition. Capital compounds when it is invested; knowledge only when it flows. Held in silos, in documents nobody reads, in systems nobody can decode, in the heads of people about to retire, it depreciates quietly and is written off without ever appearing as a loss.
The models will keep improving without your help. The knowledge will not.
Author:
Deepak Dastrala,
Chief Executive Officer, Purple Fabric


