By the middle of 2025, enterprises had put an estimated $30 to $40 billion into generative AI. That July, a preliminary study from MIT’s Project NANDA looked at what the money was buying, and reported that 95% of the organizations it assessed had seen no measurable impact on the profit and loss statement. Staff felt more productive. The bottom line had not moved.
That number does not surprise me. For eleven years I worked inside a large bank that adopted new technology faster than almost anyone in its market. It chased every new tool the moment it appeared, and it put its managers through course after course to keep pace.
I saw what the pressure to adopt does to people. The reward went to whoever could stand up at the strategy session and report that the latest thing was now in use. Whether it had changed any outcome was a slower question to answer, and a far less popular one to ask.
Handing every employee a ChatGPT, Claude or Copilot licence is this year’s version of that reflex. It is a purchase, and a purchase is easy to file under strategy, because from the finance system the two look identical. A strategy is a decision about which work the company should do and where a wrong answer would quietly cost real money. A licence answers neither question.
What a licence actually buys
Start with what the licence is supposed to buy. The same MIT study found that only 40% of companies had paid for an official subscription, while at more than 90% of them staff were already using AI at work on their own accounts.
Read those two numbers next to each other and the case for the seat as a strategy quietly falls apart. Access was never the scarce thing. Your staff solved access months ago, on their phones, without a budget line or a rollout plan.
To be precise about what the seat does buy: the enterprise tiers of ChatGPT, Claude and Copilot promise by contract that your prompts do not train the vendor’s models, they put usage behind the company’s own sign-on with an audit trail, and the Microsoft 365 version can read the company’s own mail and files, which a personal account cannot. For a firm that handles client data, moving staff off personal accounts, where chats can feed training runs by default, is worth paying for on its own. Every item on that list is governance around the work. None of it changes the work, and the work is where the return was supposed to come from.
The returns are missing because the data is
So why do the returns not show up? Gartner, forecasting that a third of generative-AI projects would be abandoned after the proof of concept, named poor data quality among the reasons. That matches what I saw for years inside large institutions and what we see now building for clients.
Corporate data lives in formats that disagree with one another, in systems built never to speak. Add-backs booked inconsistently. Three versions of the same figure in three files. A trial balance that maps to a different chart of accounts than last quarter. Thin where it matters and doubled where it does not.
Drop a language model into that and you have not bought a strategy. You have hired a brilliant analyst on their first morning, fluent in everything and sure of nothing that matters, with no idea which version of the numbers is the real one or who is allowed to see it. The model was never the missing piece. What was missing is the prepared data underneath it, and no licence has ever shipped with that.
I have watched this fail in ways that look identical on the surface. An agent set to find personal data across a data room could not open several of the files, so it reported, with full confidence, that the set held none. Another read through a CRM and went straight past a customer whose name was a cartoon character. A third built a clean dashboard that set one company’s 2025 figures against another’s 2023 and never noticed it was comparing different years. Every answer looked finished. Not one was safe to act on.
The dashboard will glow green
The uncomfortable part is that while all of this happens, usage will look healthy. When the research group METR ran a controlled trial in early 2025 with experienced developers on code they knew well, the developers guessed the tools had made them about 20% faster. Measured, they came out 19% slower. METR’s own 2026 follow-up could not settle the number, because the later data was too tangled by who chose to take part. That is the point. The effect is hard to measure and easy to feel.
A UK government department handed out a thousand Copilot licences and its own evaluators, sitting in front of satisfied users, found no robust evidence that the time saved was turning into measurable productivity. They noted that measuring exactly that had not been the trial’s main aim.
Klarna is the same story with a share price attached. In 2024 it said its assistant was doing the work of 700 support agents and announced the win. By 2025 its chief executive was hiring people back, saying the chatbots were cheaper but produced lower-quality service. In each case the usage looked healthy, and the promised gain either did not arrive or did not hold.
Two fair objections
There is a fair objection here, the one I would raise myself if someone handed me this argument. People learn tools by using them, and a company where everyone has a licence will be more at ease with AI than one where nobody does. True, and worth something.
But ease with a tool is literacy, and literacy is not a strategy either. A finance team fluent in Copilot still cannot tell you which of its month-end tasks should be rebuilt around a model and which ones quietly leak money every quarter. Those answers do not come from more seats. They come from someone sitting down with the actual work, task by task, and deciding where a machine earns its place and where it only adds risk.
The second objection arrives with numbers, and it deserves them back. When the UK government put Copilot in front of 20,000 civil servants across twelve departments, users reported saving about 26 minutes a day, and 82% did not want to give the tool back. A 2026 evaluation at the Department for Work and Pensions put its own figure at 19 minutes a day. These are real numbers from real deployments, and anyone arguing the licences bought nothing has to answer them.
The answer sits inside the trials themselves. The savings are self-reported, and the largest study noted it could not establish where the saved time went. The thousand-licence trial from earlier in this piece belongs to the same family and read the same events from the measurement side: satisfied users, felt minutes, no robust evidence of the minutes reaching measured productivity.
Saved minutes are real. They become a business number only when a rebuilt workflow gives them a measured place to land, and unmanaged use can run the other way. Researchers at BetterUp Labs and Stanford found 41% of workers had received AI-generated “workslop” from colleagues, output polished enough to pass for finished work and hollow enough to cost its receiver about two hours of rework.
The order of the build
That is the actual job, and it has an order to it. Frontier models are remarkable instruments. You do not drive nails with a microscope. So the way we build runs in order of preference, the simplest dependable tool first and the frontier model only where nothing else will do:
- The workflow, taken apart with people who have run it inside due diligence and financial advisory teams.
- The data, collected and cleaned and put in order before any model sees it.
- Purpose-written software wherever the task allows it, because code does not hallucinate and does not burn a token budget.
- Our own models where the problem is narrow enough to train one.
- Smaller or open-source models where they are enough.
- A frontier model only where the work genuinely needs deep reasoning.
The expensive and unpredictable part is a last resort. Reaching for it first, which is roughly what buying everyone a chatbot does, is in my experience the most common way the money gets spent with nothing to show for it.
The question that decides whether any of the money meant anything is short: which work is actually done differently now. When no one asks it, the reflex from the top of this piece does its damage. Someone needs a line for the weekly meeting or the strategy session, and a manager under pressure to show something modern reaches for the one that always plays: we have rolled out AI, and usage is climbing. What the slide cannot say is what changed, because the number that would have answered that was never measured before the rollout. The report lands well. The business is where it was a quarter ago.
Four questions before the renewal
None of this needs an outside firm to see where a company stands. McKinsey’s 2025 survey of AI found 88% of companies now using AI and only 39% seeing any effect on profit. The factor it associated most strongly with the companies that did see one was redesigning the actual workflow around the model.
Put that 39% next to the 95% this piece opened on and the two look like they cannot both be true. They can. McKinsey records what executives say about AI of every kind. MIT went looking for a measurable effect in generative-AI deployments and found it at roughly one organization in twenty. Most of the distance between the two numbers is the distance between what gets reported and what gets measured, which is the gap this whole piece is about.
So before the renewal, ask four questions of your own rollout. An owner asks them of their company, an advisory partner of a client, a diligence partner of a target that credits AI with better margin. Each question has a pass and a fail, and a vague answer counts as a fail.
- The workflow. Name one workflow that was taken apart and rebuilt around the model, and let the person who runs it describe what moved to the machine and what stayed with a person. It passes if that person can walk you through the new shape of the work. It fails if the model was bolted onto the old steps and the slide calls it transformation.
- The data and its owner. Ask what was done to the data feeding that workflow before any model saw it, and who signed for it being correct, current and traceable to source. It passes if a named person owns that sign-off. It fails if the preparation was uploading the files. Deloitte’s 2026 work found only 20% of organizations had named anyone responsible for turning AI into value.
- The baseline. Ask which business number the rebuild was meant to move, what it read before the rollout, and what it reads now. It passes if all three answers exist. It fails if the answer is a usage dashboard, because a dashboard with no baseline is decoration.
- The expensive decision. Find the point in the workflow where a confident wrong answer costs the most, and ask what the model may decide there on its own and who reviews it. It passes if a named person or a deterministic check stands between the model and that decision. It fails if the model runs it unwatched, and it also fails if the model was kept away from every decision that matters, because both mean no one has worked out where it earns its place.
Four passes mean someone built a system. Four fails mean someone bought seats, and the renewal can wait until the work is mapped.
Prepare the data and build in the context, and a harder question arrives, the one that really separates a company that bought licences from an AI-native one: how much of itself should a business let an AI system see? That is the subject of the next piece, Should AI agents get access to everything your company knows?
See what we build for firms like yours.
The short version
- MIT’s 2025 research found that 95% of organizations got no measurable return on generative AI, against an estimated $30 to $40 billion spent.
- In the same MIT research only 40% of companies held official subscriptions while at more than 90% of them staff already used AI on their own. Access is therefore rarely the constraint, and buying seats is procurement rather than strategy.
- Prepared data matters as much as the model. Gartner counts poor data quality among the reasons AI projects are abandoned after a pilot.
- Usage dashboards mislead. In METR’s early-2025 trial developers felt about 20% faster and measured 19% slower, a gap its 2026 follow-up left unresolved, and Klarna, after saying its assistant did the work of 700 agents, moved back toward human support in 2025 citing lower quality.
- Self-reported time savings are real and are not yet a business result. UK government trials report 19 to 26 minutes saved a day, and the largest could not establish where the time went. Unmanaged use can also run negative: BetterUp Labs and Stanford found 41% of workers receiving AI-generated “workslop” that costs about two hours of rework per instance.
- Before renewing an AI subscription, a business can separate a strategy from a purchase by naming the task that now runs differently, the place a wrong answer would cost the most, what was done to the data beforehand, and how the result was measured apart from usage.
Sources
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025”, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- Gartner, “Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025”, press release, July 29, 2024. https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
- Entrepreneur, “Klarna CEO Reverses Course by Hiring More Humans, Not AI”, May 9, 2025. https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, July 10, 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, “AI uplift study: February 2026 update”, February 24, 2026. https://metr.org/blog/2026-02-24-uplift-update/
- UK Department for Business and Trade, “Microsoft 365 Copilot Evaluation”, 2025. https://assets.publishing.service.gov.uk/media/68adbe409e1cebdd2c96a19d/dbt-microsoft-365-copilot-evaluation.pdf
- McKinsey (QuantumBlack), “The State of AI in 2025”, November 2025. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Deloitte, “Closing the AI value gap”, 2026. https://www.deloitte.com/dk/en/services/consulting/perspectives/closing-the-ai-value-gap.html
- The Register, “UK govt study: Copilot AI saved workers 26 minutes a day” (GDS cross-government trial of 20,000 civil servants), June 3, 2025. https://www.theregister.com/2025/06/03/uk_government_study_ai_time_savings/
- The Register, “DWP finds Copilot saves civil servants a whopping 19 minutes a day”, February 4, 2026. https://www.theregister.com/2026/02/04/dwp_finds_copilot_saves_civil/
- Harvard Business Review (BetterUp Labs and Stanford Social Media Lab), “AI-Generated ‘Workslop’ Is Destroying Productivity”, September 2025. https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity
- OpenAI, “Enterprise privacy at OpenAI” (no training on business data from ChatGPT Enterprise, Team or the API). Accessed July 26, 2026. https://openai.com/enterprise-privacy/
- Anthropic, Commercial Terms of Service (“Anthropic may not train models on Customer Content from Services”). Accessed July 26, 2026. https://www.anthropic.com/legal/commercial-terms
- Microsoft Learn, “Data, Privacy, and Security for Microsoft 365 Copilot” (prompts, responses and Microsoft Graph data are not used to train foundation LLMs). Updated July 9, 2026. https://learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-privacy