2026-08-05
The Chat Window Was Never the Product
Nearly every enterprise has piloted an AI agent and very few have put one to work. The gap is not model capability but integration and governance, and closing it takes fewer, simpler moves than the industry would have you believe.
Nearly every enterprise has now piloted an AI agent. Very few have put one to work. An IDC survey of more than 900 organisations, commissioned by AWS and published in November 2025, found 62 per cent actively experimenting with agentic systems and roughly 3 per cent scaling them across more than one department. The pilots are near universal. The production deployments are rare. That gap, not the technology, is where the interesting questions now live.
The shift underway is easy to state and easy to underestimate. For three years the default interface to AI has been a chat window: you ask, it answers, you copy the answer somewhere useful and carry on. Helpful, certainly, but the human remains the workflow engine and the tool only ever supplies parts. The systems now reaching production invert that arrangement. Given a goal, a set of tools and permission to act, they plan the steps, call the systems, check the results and escalate the exceptions. They chase the invoice, reconcile the ledger, book the engineer, file the report. The work between the messages, which is most of the work, becomes the machine's job rather than yours.
The sceptical case deserves to be stated properly, because it is strong. Gartner expects more than 40 per cent of agentic AI projects to be cancelled by the end of 2027, on grounds of escalating cost, unclear business value or inadequate risk controls. Anyone who has watched an agent confidently work its way into a dead end will recognise the concern: a system that completes nine steps correctly and invents the tenth has not saved you work, it has manufactured a liability with an audit trail. The sceptics are right about naive deployments.
But look at where the failures actually occur. MIT's Project NANDA reported last August that roughly 95 per cent of enterprise generative AI pilots were delivering no measurable return, and located the fault not in model capability but in integration: systems that could not reach the data they needed, processes nobody owned, outcomes nobody could verify. The models cleared the bar some time ago. The organisations did not. That is the heart of it.
The industry's response to this gap has been to sell complexity: orchestration platforms, multi-agent frameworks, maturity models with more levels than the organisations they assess. Complexity sells, because it flatters the difficulty of the problem. Yet the organisations that have crossed from pilot to production tend to have done a small number of unglamorous things well. They picked one workflow whose outcome can be checked mechanically, not a vague ambition to transform a department. They gave the system the smallest possible surface into their estate: one door, one purpose-built endpoint, one credential that can be revoked in a single click. They put an approval gate wherever a mistake is expensive, and removed it only where the evidence said they could. And they gave the whole thing a named owner, because a workflow nobody owns is a workflow nobody fixes.
Here is the part that should make every executive sit up: governance, done this way, is not the tax on autonomy. It is the enabling layer. A system whose permissions are narrow, whose actions are logged and whose failures are visible is a system you can trust with more, sooner. The organisations still treating governance as paperwork to be completed after the pilot are the ones stuck in the pilot. The ones that built it in from the first workflow are quietly compounding, adding a second workflow, then a fifth, on foundations that do not need re-laying.
The chat window was a demonstration. The workflow is the product. The question to put to your own organisation is not whether a system can act on your behalf, because it now can. It is whether you can say, precisely, what you have authorised it to do, and prove it afterwards. Very few can. Start there.
