Skip to main content

US Insider

Thursday, July 16, 2026

Why Enterprise AI Pilots Stall Before Production, and What the Successful 5% Do Differently

US Insider
Why Enterprise AI Pilots Stall Before Production, and What the Successful 5% Do Differently
Photo Courtesy: Unsplash.com

Why Enterprise AI Pilots Stall Before Production, and What the Successful 5% Do Differently

By: Audrey Denise B. Cachuela

Many enterprise AI pilots pass the demo. The system answers correctly, and the numbers check out. Months later the same system still has not touched a live record, because reading information and changing it are different problems and only the first one shows up in a demo.

The hard part of enterprise AI starts after the answer arrives, with getting a system permission to act, getting a person to approve that action, and keeping a record clean enough to survive an audit. Rajeev Jaswal, co-founder and CEO of Argo Intelligence and a former CIO, built the company’s platform around that sequence.

The widely quoted figure, that 95% of enterprise AI pilots fail, usually arrives without its context. MIT NANDA’s 2025 research found something more specific: only 5% of integrated enterprise AI pilots were extracting millions of dollars in value, while the other 95% showed no measurable effect on profit and loss. Researchers reached that figure after reviewing more than 300 public AI initiatives, running 52 structured interviews, and surveying 153 senior leaders in the first half of 2025 (Source: MIT NANDA, 2025).

The failure rate landed unevenly across company sizes, which says something about where the obstacle actually sits. Large enterprises, defined in the study as firms above $100 million in annual revenue, ran the most pilots and assigned the most staff to AI work, yet converted pilots into live deployments at the lowest rate of any size band. Top mid-market performers went from pilot to production in about 90 days, while enterprises typically took nine months or longer (Source: MIT NANDA, 2025). More staff and more pilots produced fewer finished deployments, so effort and budget are not the main constraints.

Why Enterprise AI Pilots Look Better Than They Perform

A pilot usually runs under controlled conditions. A team picks a narrow task, cleans the data, limits the number of users, and keeps a person close enough to catch mistakes before they spread. Inside that room an AI assistant summarizes documents, drafts responses, and finds patterns in minutes that would take an analyst all afternoon.

Production removes those conditions. The system meets messy data, undocumented processes, conflicting permissions, applications that go down without warning, and employees who cannot stop work every time the AI behaves unexpectedly. A plausible answer stops being enough once somebody has to act on it.

The mismatch shows up in how people use these tools. Only 40% of companies in the same research had purchased an official AI subscription, and employees at more than 90% of surveyed companies used personal AI accounts for work anyway (Source: MIT NANDA, 2025). Employees were already getting value from AI on their own terms, while official pilots struggled to turn that into something the company could standardize and trust.

MIT NANDA describes generic AI tools that do not learn on the job, performing fine for one person asking one-off questions while failing to retain feedback or adjust to a company’s context inside a workflow that’s already running. Tools bought from or built with outside partners reached deployment about 67% of the time, roughly double the rate of tools built internally (Source: MIT NANDA, 2025). That split points to an organizational problem. A company can prove a model writes convincing text long before it can prove the surrounding system runs safely enough for the business to stand behind it.

The Real Test Is Moving From an Answer to an Action

Asking an AI system to summarize last quarter’s numbers carries little risk. Letting it change a forecast, update a customer account, approve an invoice, or edit a production schedule changes what the company does.

Organizations cross that line more casually than they should, and NIST’s generative AI guidance names the reason as automation bias, where people defer to a system’s output beyond what the evidence supports because the answer arrived quickly and sounded certain (Source: NIST AI Risk Management Framework, 2024). A fast, confident answer is not necessarily correct.

Crossing that line also moves who owns the consequences. A person can still catch a bad recommendation before it does damage. Once an AI system executes the change, the moment for catching it has passed, and whoever approved that authority owns whatever happens next. Most pilots are built without anyone having answered that accountability question.

Jaswal watched this pattern across 25 years in enterprise IT, including CIO roles at Rapid7 and Red Hat, where promising projects stalled when they reached production. The delay came from questions nobody had answered yet: who can request the action, which records the system can touch, what it can change, whether a person signs off first, and whether the company could reconstruct the event if something went wrong.

“I rarely saw an AI project stall because the model wasn’t good enough,” Jaswal says. “They stalled when someone from infosec asked who approved a change, and nobody had an answer. The technology was ready. The accountability wasn’t.” NIST treats these questions as central to managing generative AI risk, calling for documented decision ownership, explicit human oversight, retained records, and procedures for when something breaks (Source: NIST AI Risk Management Framework, 2024). Companies that defer this work until the pilot is wrapping up do the hardest part of the project under the worst time pressure.

Enterprise AI Governance Starts With Access Control

The first question is what the system is allowed to touch. An employee typing a plain-English request should get no broader authority through the AI than they already hold, so the system has to know who is asking and whether that person can already see or change the information involved. A request from finance should never give access to personnel files.

This gets harder the moment one request spans several applications. A revenue question might pull from a CRM, an ERP system, a data warehouse, and a spreadsheet one department has maintained by hand for years, each with different owners, permissions, and update schedules. The AI layer has to carry the person’s identity and authority through the whole transaction rather than smoothing over those differences.

The risk compounds when one agent is wired into many systems. NIST flags unvetted third-party components as their own risk category, because they strip visibility from whoever downstream depends on that system’s output (Source: NIST AI Risk Management Framework, 2024). An agent connected to a dozen applications adds a dozen places a permissions mistake can hide.

Human Approval Workflows and Audit Trails Turn Oversight Into Evidence

Human oversight often gets treated as a single setting, where either a person reviews everything or nobody does. Reviewing everything becomes a formality, and reviewing nothing removes the check entirely. Workable approval scales to the risk. Already-approved data pulls run unreviewed, low-stakes reversible actions run automatically inside defined limits, and anything expensive, sensitive, or hard to undo requires explicit approval.

The question worth settling is which named person approves what, at which moment, with what information in front of them, and with what power to reverse it afterward. The EU AI Act’s requirements for high-risk AI systems follow a similar structure of logging for traceability, documentation, risk controls, and human oversight (Source: European Commission AI Act, 2026). Companies outside those rules can still borrow the structure, since oversight works only when a specific person is responsible for it.

Approval settles who authorized an action. An audit trail settles what happened. For any transaction that matters, a company should be able to reconstruct the original request, the identity behind it, which systems it touched, what came back, who approved it, and the outcome, with a timestamp and the system version.

Those records let engineers investigate failures instead of guessing, let managers catch unusual behavior early, and let business owners separate a model mistake from a permissions failure, a stale data source, or a human decision that turned out wrong. NIST recommends tying these histories to testing and validation records and including data sources, known issues, and access details in an AI system inventory (Source: NIST AI Risk Management Framework, 2024).

A fourth control rarely appears in pilots. It checks the system’s account of what it did against what it actually did. Argo runs that verification step between approval and the record, so the audit trail reflects the transaction rather than the system’s summary of it. Without it, a company is trusting the same system to both act and report on its own actions.

Together, these controls let a company expand AI’s authority in stages. Hand over limited authority, watch what the system does, measure the outcome, and expand access once the evidence supports it.

What the Successful 5% Do, and How They Measure Enterprise AI ROI

The strongest deployments start from a recurring business process. Teams define success and map every system and person the work touches before deciding where AI fits.

They also design for the whole transaction. If the outcome requires updating a system of record, the pilot includes that update under real controls from day one, because a recommendation an employee still has to carry by hand through five applications has not proven anything.

Ownership gets split on purpose. Security defines what the system can access, legal and compliance interpret which rules apply, business leaders decide what a good outcome looks like, technical teams watch performance, and the people doing the work flag where the plan misses daily reality. Without that, an innovation team hands off an unresolved pile of security, integration, and accountability questions once the demo ends.

Measurement follows the job. A pilot ready for production should show whether it finished the intended workflow, stayed inside its permissions, routed exceptions correctly, logged the transaction, and cut the time or cost tied to the original process. Systems answering internal data questions get measured on response time, hours saved, correction rates, and whether the answer arrived before the decision closed. Systems updating records get measured on completions, rejections, reversals, approval time, and policy violations.

Organizations that got this right reported eliminating $2 million to $10 million a year in outsourced customer service and document processing, cutting outside agency spending by roughly 30%, and saving about $1 million a year on outsourced risk checks in financial services, all without shrinking the internal teams that used to do that work by hand (Source: MIT NANDA, 2025). Those savings came from back-office functions, which most pilots skip in favor of customer-facing work.

The Path From Pilot to Production Runs Through Trust

Companies that reach production hand AI real authority while keeping control of their data and decisions. That depends on settling who can start an action, who approves it, whether the system’s account gets checked, and what record survives before any pilot scales.

Those controls only work together. Access control without approval limits who can act, but not what they do. Approval without an audit trail leaves no way to prove the sign-off happened as intended. Failures happen when any one of them is missing.

Argo Intelligence built its platform in that order, putting access control, human approval, verification, and the audit record in place before adding automation. The same test applies to any platform a company is weighing. Ask which of the four already exist in the product and which ones the buyer is expected to assemble afterward.

Timing matters. MIT NANDA found many enterprises on track to lock in vendor relationships over the next 18 months that will be difficult to reverse, with switching costs compounding the longer a system trained on a company’s workflows stays in place (Source: MIT NANDA, 2025). Waiting for a cleaner moment means building governance after the decision has effectively been made.

For any organization whose AI pilots keep producing polished answers without changing how work gets done, the next step is the same. Map one high-value process, define where its access boundaries sit, assign every approval decision to a named person, and require a full audit trail before it touches anything live.

US Insider

This article features branded content from a third party. Opinions in this article do not reflect the opinions and beliefs of US Insider.