# Agent ROI: why fewer than 25% of AI agent pilots reach production

> Most AI agent pilots stall before production: no owned metric, no workflow redesign, no kill criterion. McKinsey shows 62% experimenting but 23% scaling. The five-question gate to run first.

- **Pillar:** Explainers
- **Author:** Nishtha Gupta (Contributor · Operations Lead, Demand Nexus)
- **Published:** 2026-06-16T15:14:00.000Z
- **Tags:** agents, enterprise-ai

## TL;DR

Most AI agent pilots stall before production because they launch without an owned business metric, a workflow redesign, or a kill criterion. McKinsey's data shows 62 percent of organizations experimenting with agents but only 23 percent scaling one anywhere. The fix is a gate you run before the pilot starts, not a postmortem after it dies.

import SignalChart from '~/components/viz/SignalChart.astro';

Somewhere in your stack right now there's an agent pilot with no owner, no target metric, and a renewal date in Q3. That pilot is already dead; the invoice just hasn't noticed. The numbers say this is the normal case, not the exception: McKinsey finds [62 percent of organizations experimenting with AI agents while 23 percent report scaling an agentic system anywhere in the enterprise](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai), and in any single business function, [fewer than 10 percent are scaling agents](https://www.colabsoftware.com/post/mckinseys-state-of-ai-2025-what-separates-high-performers-from-the-rest).

Your CFO will eventually ask which side of that funnel your spend is on. Better to ask it yourself first.

## The funnel, with numbers on it

Walk the stages and watch the drop-off:

- **88 percent** of organizations [use AI in at least one business function](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai).
- **62 percent** are [experimenting with AI agents specifically](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai).
- **23 percent** are [scaling an agentic system somewhere](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai).
- **Under 10 percent** are [scaling agents in any one function](https://www.colabsoftware.com/post/mckinseys-state-of-ai-2025-what-separates-high-performers-from-the-rest).

And the forward-looking number that should shape your contracts: Gartner predicts [over 40 percent of agentic AI projects will be canceled by the end of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027), citing escalating costs, unclear business value, and inadequate risk controls. S&P Global data points the same direction: the share of enterprises [abandoning most of their AI initiatives jumped from 17 percent in 2024 to 42 percent in 2025](https://www.softwareseni.com/why-88-to-95-percent-of-enterprise-ai-pilots-never-reach-production/), as reported in industry analyses.

The money is flowing anyway. Gartner forecasts [agentic AI](/what-is-agentic-ai) software spending [up 141 percent in 2026 to nearly $202 billion](https://dataconomy.com/2026/05/28/global-ai-spending-259t-2026/). That combination, surging spend plus a 40 percent cancellation forecast, is not a contradiction. It's a market paying tuition.

<SignalChart
  type="bar"
  orientation="horizontal"
  pillar="explainers"
  title="The agent pilot funnel, 2025-26 survey data"
  caption="Each value traces to the cited source in the text. The final stage is reported as under 10 percent."
  suffix="%"
  max={100}
  categories={["Using AI in ≥1 function", "Experimenting with agents", "Scaling agents anywhere", "Scaling in a given function"]}
  series={[{ name: "Share of organizations", data: [88, 62, 23, 10] }]}
  source={{ label: "McKinsey, The State of AI", href: "https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai" }}
/>

## Why pilots die: it's the operating model, not the model

The post-mortems are remarkably consistent, and almost none of them blame the AI.

**No owned metric.** "Deploy an agent" is not an objective. The organizations getting value set outcome targets first. McKinsey's high performers, the [5.5 percent of respondents attributing more than 5 percent of EBIT to AI](https://www.colabsoftware.com/post/mckinseys-state-of-ai-2025-what-separates-high-performers-from-the-rest), are the ones who tied AI work to a number someone owns.

**No workflow redesign.** Agents bolted onto an unchanged process automate the old process's problems. High performers are [2.8x more likely to report fundamental workflow redesign, 55 percent vs 20 percent of others](https://www.cxtoday.com/ai-automation-in-cx/mckinseys-state-of-ai-the-scaling-gap-is-now-cxs-problem/). That's the single cleanest "do this" signal in the dataset.

**Costs counted at the token, not the project.** The pilot budget covers API calls. Production costs include integration, evaluation, monitoring, and the people reviewing outputs. Gartner's cancellation drivers lead with [escalating costs](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) for a reason. We're publishing a full cost breakdown of the agent stack on June 29; the short version is that tokens are the visible minority of the bill.

**No kill criterion.** Pilots without a pre-agreed failure condition don't fail; they linger. Lingering pilots are the most expensive kind, because they consume the team's credibility along with the budget.

## The gate: five questions before any agent pilot starts

Run this before you sign, not at renewal. A "no" on any of the first four means don't start yet.

1. **Which number moves?** Name the metric (deflection rate, cycle time, cost per ticket, pipeline coverage) and the owner whose comp feels it.
2. **What's the baseline?** If you can't measure the process today, you can't prove the agent improved it. Instrument first.
3. **What changes about the workflow?** If the answer is "nothing, the agent just helps," you're buying a 23-percent-survival-rate lottery ticket.
4. **What kills it?** A date and a threshold. "If audited accuracy is under X by week 8, we stop."
5. **What does production cost?** Not the pilot. Production: integration, eval suite, review labor, monitoring. Ranges are fine; silence is not.

Where this framework breaks, and you should know its edges: it under-serves genuine R&D. If you're deliberately exploring capability with no near-term metric, exempt that work explicitly, cap it as a research budget, and don't let it masquerade as an ROI pilot. The framework also can't rescue a pilot built on a vendor that isn't real; Gartner estimates [only around 130 of thousands of self-described agentic vendors](https://www.reuters.com/business/over-40-agentic-ai-projects-will-be-scrapped-by-2027-gartner-says-2025-06-25/) actually are. Vendor diligence is a separate exercise (our June 20 piece covers it).

## What to do Monday

**If you run a team under ~200 people:** inventory your agent pilots. For each, write the metric, owner, baseline, and kill date on one page. Any pilot where you can't fill all four fields gets 30 days to earn them or it's cut. You likely have one or two pilots; this takes an afternoon.

**If you run revenue ops at scale:** same inventory, plus two additions. Move every agent contract you can to outcome-aligned or short-cycle terms ahead of the renewal wave, and pick one workflow (not one tool) for genuine redesign in H2. The 2.8x redesign signal is the highest-value fact in this piece; resource it like you believe it.

Either way, the position to be in by Q4 is simple: every agent dollar in your budget maps to a metric, an owner, and a kill switch. The 40 percent that get canceled will mostly be the dollars that didn't.

## FAQ

**What percentage of AI agent pilots reach production?**
McKinsey's State of AI survey shows 62 percent of organizations experimenting with AI agents while 23 percent report scaling an agentic system anywhere in the enterprise, and fewer than 10 percent are scaling agents within any single function. Separately, Gartner predicts over 40 percent of agentic AI projects will be canceled by the end of 2027. Roughly one in four experimenters reaching scale is the honest planning number.

**Why do most AI agent projects fail?**
The dominant causes are organizational, not technical: no owned business metric, no baseline measurement, no workflow redesign, and costs that escalate beyond the pilot budget. Gartner cites escalating costs, unclear business value, and inadequate risk controls as the drivers behind its 40 percent cancellation prediction. McKinsey's high performers are 2.8x more likely to have fundamentally redesigned workflows around AI rather than bolting agents onto existing processes.

**How do I measure ROI on an AI agent?**
Define one target metric and its owner before the pilot starts, measure the baseline first, and set a kill threshold with a date. Count full production costs (integration, evaluation, human review, and monitoring, and not tokens alone) against the metric's movement. If you cannot name the metric, the baseline, and the kill criterion on one page, the pilot is not measurable and should not start.