AIAI for Business

The spreadsheet that embarrassed a CFO and changed how we measure AI

A major retailer celebrated millions in projected AI savings, then watched the number quietly shrink to almost nothing once someone counted the full cost. That moment, repeated across industries throughout the early 2020s, explains why measuring AI returns remains the most underrated skill in enterprise technology.

🎙️

Listen to the podcast

4 min

In late 2022, Walmart's finance team was in an uncomfortable position. The company had invested heavily in AI-powered demand forecasting tools, and the early internal projections were, by any measure, striking. Analysts estimated the systems would cut inventory waste by hundreds of millions of dollars annually. Slides went to leadership. Timelines were set. Then someone asked a question that no one had answered clearly: what are we actually counting as the cost?

What actually happened

Walmart has been transparent in earnings calls and press interviews about its AI investments in supply chain, though the specific internal projections described above reflect reported industry patterns rather than a single disclosed internal document. What is well-documented is the broader dynamic: by 2023, Walmart confirmed meaningful efficiency gains from AI-assisted inventory management, but company leadership also acknowledged that calculating the true return required factoring in infrastructure upgrades, model retraining cycles, data quality projects, and the workforce hours redirected toward AI governance rather than traditional operations.

This is not a story of failure. Walmart's supply chain AI genuinely works. But the company's experience illustrates a pattern that McKinsey documented in its 2023 State of AI report (an independent research publication): the majority of organizations that reported measuring AI value were counting only the direct output savings, not the full input costs. The gap between "AI saved us X" and "AI returned X net of everything we spent to make it work" turned out to be wide enough to embarrass more than a few CFOs when anyone looked closely.

The classic mistake follows a predictable shape. A business unit deploys a tool, perhaps an AI assistant that automates contract review or a model that flags customer churn risk. Someone counts the time saved per task and multiplies it by volume. The number looks compelling. What does not appear in that calculation: the internal data engineering hours spent cleaning the training data, the IT infrastructure scaling costs, the weeks of a product manager's time integrating the output into existing workflows, the ongoing cost of a data scientist reviewing model drift, and the legal review of AI outputs before they go to clients. A Deloitte survey from 2022 (independent research) found that fewer than one-third of enterprises had a formal methodology for capturing total cost of AI ownership. Most were measuring gross savings against purchase price, which is roughly the equivalent of measuring a factory's profitability by comparing product revenue to the cost of the machinery while ignoring labor, energy, and maintenance.

The issue compounds when companies use vendor-provided ROI calculators. Tools from large AI platform vendors, including Salesforce, Microsoft, and Google (all of whom sell AI products and have a clear commercial interest in favorable ROI figures), typically model time savings and productivity lift using their own benchmark studies. These figures are worth examining, but they should not be taken as independent validation. That point deserves repeating whenever a vendor deck lands in your inbox.

Why it still matters

The measurement problem has not gone away. If anything, the proliferation of AI tools between 2023 and today has made it harder to track. By 2026, most enterprises run dozens of AI-assisted processes simultaneously, which makes attribution genuinely difficult. Did customer retention improve because of the AI churn model, or because the sales team changed its outreach cadence at the same time? Did legal review costs fall because of AI contract analysis, or because the company renegotiated its outside counsel rates?

What the Walmart-era experience revealed is that AI ROI is not a technical measurement problem. It is an organizational design problem. You need someone with actual authority to define what counts as a cost and what counts as a benefit, to enforce that definition across business units, and to hold measurement cadences to a fixed timeline rather than cherry-picking the moment when the numbers look best.

The companies that solved this earliest, including a handful of financial services firms and large logistics operators, did so by assigning ROI accountability to finance, not to the AI team. When the people who want the investment to succeed are also the people measuring it, the numbers will tend toward optimism. This is not cynicism; it is standard organizational behavior.

The takeaway for you

If you are involved in any AI investment decision, from approving a departmental tool to sitting on a board reviewing an enterprise AI roadmap, one question will protect you from the spreadsheet problem: who owns the cost baseline, and are they independent of the team that proposed the investment?

The methodology itself does not need to be exotic. A solid AI ROI framework covers four things: the full-loaded cost of deployment (infrastructure, integration, and staff time, not just licensing), a defined productivity baseline measured before the tool goes live, a realistic attribution model that accounts for other changes happening simultaneously, and a measurement window long enough to capture model degradation and retraining costs, which rarely appear in year-one projections.

Walmart's AI investments are genuinely productive. The lesson from their experience is that even a successful deployment can look misleading on paper if the measurement discipline is not in place from the start. The tools are not the hard part. Counting honestly is.

Finished reading?

Validate your read to earn XP and feed your radar.