AI Integration Projects Need a Check on the Output, Not a Better Model

AI integration projects mostly stall on workflow, not on the model, according to MIT’s 2025 study. Before scaling one, build a check that catches a wrong answer before it reaches a customer, a regulator or print.

Subject
Why AI integration projects fail
Published
19 AUG 2026
Reading time
5 min

In short

The finding
In one widely cited 2025 study, 95% of organisations running generative AI got no measurable return six months after their pilots.
Where it breaks
In a 2025 survey, 66% of people using AI at work said they had at least once relied on its output without checking it.
What it means for you
A check has to cover every answer that can reach a customer, including ones from tools you buy or content you license.
In this post · 3 sections
  1. —In short
  2. 01MIT’s study blames workflows and tools that don’t learn
  3. 02You inherit a failure the moment you publish it
  4. 03Most people using AI at work have acted on it unchecked
The same argument in 3:57. Our graphics, AI narration.

AI integration projects are the work of wiring an AI system into how a business actually runs. A widely cited 2025 study found that 95% of organisations pursuing generative AI got zero measurable return six months after their pilots. The report blames brittle workflows and tools that don’t learn, not weak models. Newspaper, bank and workplace evidence points to a second gap: nobody checked the output.

MIT’s study blames workflows and tools that don’t learn

The finding comes from MIT Project NANDA’s 2025 study of AI use across business, which measured return six months after each pilot. The report found that popular tools such as ChatGPT and Copilot mostly raised individual productivity, not profit. Custom tools mostly failed on brittle workflows, tools that don’t learn and a poor fit with daily work.

The same report found that outside-vendor tools went live about 67% of the time. In-house builds went live about 33% of the time — roughly half as often. The report warns that these are self-reported figures from interviews with 52 organisations, and that the gap may reflect which organisations chose to buy, not buying itself.

MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025,” July 2025. mlq.ai (hosted copy; MIT’s original link now redirects)

You inherit a failure the moment you publish it

In May 2025, a bought-in, syndicated “summer guide” ran in the Chicago Sun-Times with a reading list attached. 10 of the 15 recommended books didn’t exist. The section came from King Features, a Hearst syndication unit. A freelancer used an AI tool to write it and sent the result in without checking it.

The Sun-Times’s own newsroom never reviewed the section, because the circulation department didn’t submit the pages for editorial review. The Sun-Times hadn’t built an AI system, and still ran fabricated content under its own name, because nobody applied a check to work it didn’t write. A check that covers only what your own team writes has a hole the size of everything you buy in.

Chicago Sun-Times, “Syndicated content in Sunday print Sun-Times included AI-generated misinformation.” chicago.suntimes.com · Sun-Times CEO Melissa Bell, “Lessons (and an apology) from the Sun-Times CEO.” chicago.suntimes.com

Two lanes from the same source: an AI-generated, undisclosed draft that reaches print with no verification gate, fabricated content and all, versus the same draft with one verification gate inserted before publish, catching the fabricationPUBLISHONE INPUT, TWO OUTCOMESTHE GATE IS THE FIXDWG Nº 01VERIFY BEFORE PUBLISHAS FILED — NO GATESOURCESyndicatedAI-GENERATEDNot disclosedNO GATEPUBLISHEDRan as filed10 OF 15 FABRICATEDSAME INPUT — ONE GATESOURCESyndicatedAI-GENERATEDNot disclosedGATEVerifiesPUBLISHEDVerified✕ caught, never shippedMECHANISM SHOWN SCHEMATICALLY — SEE CAPTION FOR THE REAL FIGURENo disclosure, no passSame risk either wayVerify before it ships
Same input, one inserted step: in May 2025, a syndicated summer-reading section ran in the Chicago Sun-Times with 10 of 15 recommended books fabricated. It was AI-generated by a freelancer for a licensed content partner, undisclosed, and never reviewed by the newsroom before it printed.

Atanas Mihov and Ping McLemore, writing for the Federal Reserve Bank of Richmond, found a related cost in banking, measured in money. At large US banks, a 10% increase in a bank’s AI investment was linked to roughly a 4% rise in its quarterly operational losses. Those losses concentrated in outside fraud, problems with clients and system failures, and the effect was strongest at banks that lack strong risk management.

Atanas Mihov and Ping McLemore, “AI and Operational Losses: Evidence from U.S. Bank Holding Companies,” Federal Reserve Bank of Richmond. richmondfed.org

Most people using AI at work have acted on it unchecked

A 2025 global study by KPMG and the University of Melbourne surveyed 48,340 people across 47 countries. Among those using AI at work, 66% said they had at least once relied on its output without checking it. 56% said AI had caused them to make a mistake on the job.

Gillespie, Lockey et al., “Trust, attitudes and use of AI: A global study 2025,” KPMG and the University of Melbourne. kpmg.com

In practice the check has two parts. The system must pass a fixed set of real cases, with known right answers, before it takes on more work. And a person sees any doubtful answer before a customer does, with the doubt flagged by a separate check, not by the AI’s own say-so.

For engineersA fixed, labelled test set run before every expansion of scope, plus a confidence threshold, based on an outside signal rather than the model's own stated confidence, that sends uncertain answers to a person.

An evaluation harness is the code that runs the system against a golden set: a fixed, labelled set of real cases with known-good answers, drawn from what the system will see in production. It runs before every expansion of scope. Scope grows only once the system clears the bar on cases from the next scope, not just the current one.

A confidence threshold routes anything below a stated bar to a person before it ships. The bar should come from an external signal — agreement between independent runs, schema validation, a cheap grading pass — rather than the model’s own stated confidence, which is not well calibrated. Schema validation checks that an answer has the shape it is supposed to have, the right fields and the right types, before anything else runs on it, and it catches a badly broken answer for a fraction of the cost of judging whether the answer is actually right. Sampling review then audits a fixed share of everything that already cleared the bar, on a schedule, because a threshold tuned once drifts as the input distribution does.

Whether you build the AI or buy it, the risk sits in what reaches a customer unchecked. The decision for anyone running an AI integration project is where the check sits, and whether it covers every output that can reach a customer or a regulator: your own system’s, a vendor’s and a freelancer’s.

Updated 26 Sep 2026: corrected the MIT NANDA figure to match the report’s own wording, 95% of organisations rather than 95% of pilots. Fixed the Richmond Fed byline order, and dropped an unsourced claim that the 95% figure had been publicly disputed. Also removed Richmond Fed sample and journal details we couldn’t verify, and a claim about execution and processing errors that the Richmond Fed page doesn’t make.

AGNIZAR
Production AI · System architecture · Fractional CTO

Agnizar builds AI into your core systems, then hands it over or keeps it running. Every job starts small: one bounded piece of work, one named result, one clear decision. Book an AI Architecture Review; a senior engineer replies within one business day.

Book an AI Architecture Review