Enterprise AI Has a Compounding Problem – And Most Organisations Don’t See It Yet

Why more AI tools are producing less enterprise value – and what makes the gap structural rather than temporary

Not long ago, I sat with a technology leader at a financial services firm who opened a slide listing fourteen AI tools his organisation was running. He had built a real programme – real budget, real teams, real deployments. Then he said something I have not forgotten: ‘I cannot tell you, honestly, what the combined return is.’ He was not being evasive. He genuinely did not know. The tools had never been designed to answer that question together.

My last piece was about the governance gap – the space between management’s confidence in deploying AI and a board’s ability to oversee it. What I want to write about now is a second gap, one that sits closer to the ground. It is the gap between the AI investments organisations are making and the business outcomes they are actually measuring. And unlike the governance gap, this one is not about intent or awareness. It is about architecture.

“The typical enterprise AI programme is not building capability. It is building inventory. The two look similar from the outside. They produce very different results.”

The Tools Are Working. The System Isn’t

I have seen some version of this landscape in almost every large organisation I work with: a productivity copilot for knowledge workers, a document processing tool for operations, an analytics layer for the data team, a fraud detection model for risk, and two or three additional point solutions acquired by business units that moved faster than central IT could track. Each tool was chosen on its merits. Each delivers something measurable in its lane. None of them were chosen together.

Ask what they share – a common data layer, a unified audit trail, a single control plane that governs how decisions flow between them – and the answer is almost always nothing. They were never designed to work together. They were designed to work.

This is the difference between an AI inventory and an AI ecosystem. An inventory grows by addition – each new tool expands the list. An ecosystem grows by compounding – each new capability makes the existing ones more valuable. Harvard Business Review found that AI-driven productivity actually drops when individuals use four or more AI tools concurrently. At the enterprise level, I have seen the same dynamic play out at a much larger scale. More tools, more coordination overhead, less compound value.

The Productivity Plateau

The numbers at the task level look encouraging. Forty percent faster on document review. Thirty percent reduction in query handling time. Meaningful gains in specific workflows, reported with confidence in every management presentation. These improvements are real, and the people who achieved them deserve credit.

But Deloitte’s 2025 AI survey found that while 88% of organisations now use AI in at least one business function, fewer than 40% report measurable impact on business outcomes. The tools are present. The results are not. That gap has a specific cause: enterprise value does not live inside individual tasks. It lives in the workflows that connect them. When the tools serving those workflows were built independently, the gains stay isolated inside the tools that generated them. They do not travel.

In practice, this shows up as a plateau – impressive early deployments, genuine enthusiasm, and then a quiet flattening that arrives in year two or three and is hard to explain in a board presentation. It is not a failure of the individual tools. It is a failure of integration that was never designed in from the start.

Data Is Not Knowledge. And the Difference Is Costing You

I have sat in more than one meeting where someone gestures at a data lake and says: ‘The AI has access to all of that.’ The implication is that access equals capability – that pointing a model at a repository of documents produces the same result as building structured knowledge from it. It does not. And this is one of the most consequential misconceptions I encounter in enterprise AI today.

Data and knowledge are not the same thing. Data is what an organisation collects – transactions, records, documents, logs, interactions. Knowledge is what an organisation understands – the rules, relationships, context, and meaning that turn raw data into something that can answer a question reliably. Data has volume. Knowledge has structure. One is stored. The other has to be built, curated, and maintained. They are not the same investment, and one does not automatically produce the other.

The practical consequence shows up in how AI is implemented. The common pattern is to point a language model at a document repository – policy manuals, regulatory guidelines, product rules, historical case notes – and expect it to function as an intelligent assistant. What it actually does is probabilistic search over unstructured text. It retrieves passages that are statistically similar to the query. Sometimes those passages contain the right answer. Sometimes they contain something that looks right but is outdated, contextually incomplete, or applies to a different product or jurisdiction. The model cannot tell the difference, because it does not know the rules. It knows the text.

In a regulated environment, this is not a minor distinction. A customer service model that retrieves the wrong policy condition and presents it confidently is not unhelpful – it is a compliance event. A claims assistant drawing on a superseded document version has created an audit exposure. Data retrieval dressed as knowledge access does not carry the certainty that regulated decisions require. And the investment distortion compounds the problem: organisations keep spending on more data – larger lakes, richer pipelines, more ingestion – when the actual constraint is knowledge structure. More data does not solve the absence of curated knowledge. It often makes it harder to find, because the signal has to compete with more noise.

The sharpest version of this problem involves what I think of as deterministic knowledge – the body of information every organisation already holds but rarely structures properly. Policy rules, eligibility criteria, regulatory thresholds, underwriting guidelines. These are not inference problems. They are lookup problems: a given input has a single correct answer, and that answer does not require a language model to reason through. In a well-designed system, these queries route to a deterministic layer – answered with certainty, at low cost, with a fully auditable trail. In a fragmented landscape, they route to an LLM instead, because no shared knowledge architecture exists. The cost consequence is direct: when a significant proportion of query volume consists of questions with deterministic answers, AI operating costs are inflated by architecture, not workload. I have seen organisations experiencing rising AI spend quarter on quarter, attributing it to usage growth, when the actual driver is unnecessary LLM routing – paying inference costs for certainty that was already available, in knowledge the organisation already owned.

Accuracy at the System Level

There is a quieter version of the same problem, and it lives in arithmetic that most AI programmes never do.

Each AI component operates with its own accuracy profile. A document classification model might be accurate ninety percent of the time. A data extraction layer might run at eighty-eight percent. A downstream recommendation engine at eighty-five. Individually, these are defensible numbers – each team can point to them with confidence.

But enterprise decisions do not happen at the component level. They happen at the end of a chain – after data has moved through classification, extraction, analysis, and recommendation before it reaches a human or an automated output. Accuracy across a chain does not add. It multiplies.

0.90 × 0.88 × 0.85 = 67%.

In a consumer context, sixty-seven percent accuracy is a product feature. In financial services or insurance – where an AI-influenced decision may trigger a regulatory obligation, affect a customer’s policy, or surface in an audit – it is a liability. The enterprise bar is not seventy percent. It is well above ninety, because a wrong decision carries real consequence for a real person. Point solutions, even excellent ones, cannot solve this by themselves. They were built to perform well in their lane. They were not built to perform well in combination. And in most enterprise AI landscapes, combination is all they do.

Governance Becomes Structurally Difficult

In my previous piece, I argued that the named accountability owner – a single person responsible for AI outcomes, answerable to the board – is the highest-leverage governance change an organisation can make. Several readers picked that out as the critical one, and I agree with them.

Here is the problem: even when the will exists to make that appointment, fragmentation makes it almost impossible to implement with any substance. If your AI landscape spans five vendors, three data sources, and two business units with no shared audit trail, who does the named owner actually govern? They can own the intent. They cannot own the evidence – because the evidence is scattered across systems that were never designed to produce it together.

The boardroom accountability gap and the architectural fragmentation gap are not separate problems. They reinforce each other. A fragmented AI architecture does not just make governance harder to practise. It makes meaningful governance structurally difficult to even design.

“You cannot govern what you cannot trace. And you cannot trace decisions that cross five systems with no shared audit trail.”

Pilots Don’t Fail. Scale Does

One of the most consistent patterns I see in enterprise AI programmes is the stalled pilot. A proof of concept delivers real results – sometimes impressive ones. The business case gets approved. And then scaling takes three times longer than expected, costs twice as much, and delivers a fraction of what the pilot suggested was possible.

The explanation I hear most often is change management. Or data quality. Or organisational resistance. These are real factors. But they are rarely the root cause.

The root cause is simpler and harder to fix: a pilot built in isolation – on a curated dataset, with a focused use case, inside a controlled environment – has to be re-integrated from scratch into the broader enterprise landscape every time it needs to touch another system. There is no shared foundation to build on. Without one, every scale-up is a custom engineering project. The pilot did not fail. Scale was never designed to work.

The Structural Question

None of this is an argument for slowing down. The case for enterprise AI adoption is clear, and 2026 is increasingly being described – by analysts, by boards, by technology leaders – as the year where scale separates from stall. The competitive cost of delay is real.

But after three decades in financial services technology – through ERP, through digital transformation, through cloud – I have seen what happens when organisations move fast without asking the structural question early enough. The fragmentation accumulates. The integration debt grows. And eventually the cost of the next step is higher than the value it will return.

The fragmentation most enterprise AI programmes have built is not a temporary state they will grow out of. It compounds. Every new tool adds integration debt. Every new vendor adds governance complexity. Every new pilot adds to the scale-up backlog. The organisations that will sustainably scale AI are not necessarily the ones that moved fastest. They are the ones that at some point asked a different question – not ‘which tool solves this problem?’ but ‘how do we build an environment where our AI investments compound rather than accumulate?’

That question has an answer. But it starts somewhere different from where most AI programmes begin. I will write about that in the next piece.

I would be glad to hear from others who are living this – whether you have hit the plateau, found a way through it, or are still in the early stages where the inventory looks like an ecosystem. The conversation is useful from any of those positions.

References: Deloitte State of AI in the Enterprise (2025): 88% of organisations use AI in at least one function; fewer than 40% report measurable business impact. Harvard Business Review (2026): AI-driven productivity drops when individuals use four or more AI tools concurrently. MIT Sloan Management Review: Only 5% of companies have AI integrated into core workflows at scale.

Author:

Hitesh Kumar Arora, 
EVP & Enterprise AI Leader, Intellect Design Arena · Independent Committee Member (IT & Risk), Financial Services Sector · 30 Years in BFSI & Insurance Technology

Enterprise AI Has a Compounding Problem – And Most Organisations Don’t See It Yet