The last article looked at why most AI agent projects fail. The natural next question: how do you even measure that in the first place? Anyone judging whether an agent is working, as a CFO or as a founder, needs more than a gut feeling that things "feel more productive." And once you look at the current surveys, that turns out to be a separate problem of its own, independent of whether the agent itself performs well.
The number you have to read twice
PwC surveyed 4,454 CEOs across 95 countries for its 29th Global CEO Survey (released January 2026 in Davos). 56% report neither increased revenue nor lower costs from AI over the past year. At first glance, that reads like confirmation that "AI doesn't pay off." Look again and a different picture emerges: 30% of CEOs do report revenue gains, 26% report cost reductions, just rarely the same company at the same time. Only 12% achieve both. PwC calls this group the "vanguard," and it isn't distinguished by better models but by defined roadmaps and a technology environment built for integration.
That's the first thing worth noticing: AI isn't worthless at these 56% of companies. Value and cost savings simply tend to land in different departments, and almost nobody tracks both together.
The metric that's actually being tracked
A survey commissioned by Appian but editorially run by Harvard Business Review Analytic Services, covering 385 decision-makers (spring 2026), supplies the explanation. 59% of organizations already have AI in production. But look at what's actually being measured, and the picture flips: 64% track productivity gains, 58% operational efficiency, but only 30% new revenue streams and only 35% ROI. (Worth flagging: the study was commissioned by Appian, an automation vendor. That doesn't necessarily undercut the numbers, but it's always worth noting for sponsored research.)
Most companies, in other words, measure exactly the metrics that are easiest to spin positively ("feels more productive") and avoid the metrics that would actually show whether the investment paid off.
The mistake Gartner names outright
Perhaps the sharpest diagnosis comes from Gartner. Twisha Sharma, Senior Principal Research in Gartner's finance practice, puts it this way in a recent analysis: "AI does not follow one cost curve, and it does not produce one uniform type of value. CFOs need to stop looking for a single ROI formula and instead build a balanced portfolio that includes productivity use cases, targeted process improvements, and selective transformational bets." — Gartner Says CFOs Need to Rethink the ROI of AI Investments
A customer-support agent that pre-sorts tickets, an analysis agent that prepares decision material, and an agent that unlocks an entirely new product capability: these are three different kinds of bets, with three different time horizons and three different definitions of success. Measuring all three with the same ROI formula systematically underrates some and overrates others. Fittingly, a separate, more recent Gartner survey of 204 finance leaders found that 45% of finance AI investments lean toward productivity and only 20% toward better decision quality, further evidence that most organizations stay stuck in the "easily measurable" zone instead of investing where the real leverage is.
The connection to the last article
This is where it loops back to the previous piece: MIT and BCG independently found the same 5% success rate for substantial AI value in the enterprise. Both studies show something else too: the successful 5% aren't distinguished by better models. They're distinguished by narrowly scoped use cases embedded in real workflows. Measurement discipline and narrow scope are two sides of the same coin. Tie a single, clearly defined metric to a single, narrowly scoped task from the start, and you'll find out early whether the agent is carrying its weight almost by default. Start broad and measure "productivity," and you can burn months without ever knowing if it paid off.
A framework that works with a finance lens
From my experience with 30+ founders, and from the research above, a simple principle falls out, deliberately not a 7-point framework, but three questions that should be answered before the first prompt:
- What kind of bet is this? Routine automation (hard cost/time savings, measurable in weeks), process improvement (quality/error rate, measurable in months), or a transformational bet (new revenue stream or capability, only visible in quarters)? Each needs its own metric, Gartner's core point above, broken down one level further.
- What single number decides whether it worked? Not three, not "felt productivity," but one concrete figure tied to a P&L line (cost per transaction, hours per week, conversion rate), fixed before launch. Exactly the point where, per the HBR/Appian data, most companies dodge.
- Who decides when it counts as "failed"? Without pre-defined kill criteria, a project usually keeps running until the "feeling" turns, months after the numbers would already have shown it.
That sounds less exciting than "we're building an AI agent for X." But it's the difference between a project that delivers a clear answer after one quarter, yes or no, and one that just keeps running because nobody defined what a no would look like.
Sources:
Analysts
- Gartner: Gartner Says CFOs Need to Rethink the ROI of AI Investments (March 2026)
- Gartner: Gartner Survey Shows 45% of CFOs Say Their AI Investments Lean Toward Productivity, While 20% Say These Investments Lean Toward Decision Quality (July 2026, n = 204 finance leaders)
Professional Services
- PwC: 29th Global CEO Survey (n = 4,454 CEOs across 95 countries, January 2026)
Research (sponsored — see note in text)
- Harvard Business Review Analytic Services, commissioned by Appian: cited via Appian press release (n = 385 decision-makers, 2026)
Direct follow-on
- Why Most AI Agent Projects Fail — the previous article in this series
- What an AI Agent Actually Costs — the next article in this series