Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

McKinsey's research indicates that 93% of enterprise AI teams are over budget, primarily due to costs related to response refinement in agentic AI, which accounts for 60% of the total AI expenditure. This highlights the financial strain organizations face while trying to enhance the effectiveness of their AI systems.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Learn this in 60 seconds

Key facts, context, and what it means, in one minute.

:60
0:001:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprise AI teams exceed their budgets.

02

Response refinement consumes 60% of agentic AI spending.

03

Agentic AI cost overruns are a common issue among enterprises.

Token prices have fallen more than 99% in roughly two years. Enterprise AI bills have tripled anyway. That contradiction sits at the center of a July 2026 McKinsey report that maps, in unusually precise terms, where agentic AI money actually goes and why most organizations cannot yet account for it.

The numbers are stark. According to McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 across 75 qualified respondents spanning five major industries, 93% of enterprise participants report exceeding their AI budgets. Separately, data from Menlo Ventures cited in the report showed that enterprise large language model spending tripled over a 12-month period by the end of 2025. Meanwhile, Stanford's HAI 2025 AI Index documented that inference cost for GPT-3.5-level capability collapsed from $20 per million tokens to $0.07 through 2024, a drop of more than 99%.

The divergence between falling unit costs and rising total bills is not a paradox. It is a volume and architecture problem, and McKinsey's QuantumBlack team lays out why.

Where the money actually goes

The single largest cost center in an agentic AI deployment is response refinement, the iterative loop in which an agent checks, revises, and regenerates its own outputs before returning a final answer. According to McKinsey, that process consumes 60% of total agentic AI costs. For enterprise teams that assumed their spend was split evenly across retrieval, reasoning, and generation, that concentration is a significant recalibration.

Share of agentic AI costs by category60Response refinement40Other agentic operations
McKinsey & Company, July 2026 · © MarketScaleDownload chart

Three structural forces compound the problem, according to the McKinsey report. First, enterprises are scaling AI efforts at pace, so more workloads are generating token consumption. Second, LLM providers have broadly shifted from flat subscription pricing to consumption-based models, which creates an incentive for longer, more elaborate outputs. Third, and perhaps most correctable, expensive frontier models are frequently deployed for routine tasks that cheaper, smaller models could handle.

The Economic Times also reported on the McKinsey findings, noting that enterprise leaders are increasingly focused on the economics of operating AI agents rather than the underlying technology, a shift that reflects how quickly agentic deployments have moved from pilot to production.

The question is no longer whether you can deploy an AI agent. It is whether the value of what that agent produces justifies every dollar it costs to run it.

The constraint is already real

Budget pressure is not a future concern. McKinsey's forthcoming 2026 State of AI global survey, fielded between May 4 and June 8, 2026, with 1,719 participants, found that one in five organizations has already constrained AI use specifically because of AI-related operating costs. That is a material brake on adoption, one that shows up in deployment decisions, model selection, and the scope of workflows that get assigned to agents.

The pattern reflects a broader tension in enterprise AI right now. Boards and executive teams approved AI investment expecting that declining model costs would keep total spend manageable. What they underestimated was the multiplicative effect of scale and architecture: more agents, more tasks, more refinement loops, and consumption pricing that rewards output volume.

McKinsey describes the core executive question as deceptively simple: are the AI agent capabilities being built and run worth the value being extracted from them? But answering it requires instrumentation most enterprises do not yet have, specifically the ability to track whether agent output is correct, how much human supervision or repair it requires, and whether the completed work's value actually exceeds its full operational cost.

Tokens are the bill, not the value

The McKinsey report draws a pointed distinction between token spend and business value, attributing a framing to David Tepper, CEO of Pay-i: tokens are not value, tokens are the bill. That reframe matters operationally. Cost-reduction conversations anchored to per-token pricing miss the actual lever, which is the ratio of agent output quality and business impact to total operational cost.

For a VP of Operations or a CIO evaluating agentic AI portfolios, this means the relevant metric is not cost per million tokens. It is cost per completed, accurate, human-review-free task. An agent that runs more refinement loops but delivers output that never requires correction may be cheaper in practice than a cheaper-per-token agent whose outputs routinely need human repair.

The McKinsey authors, including Lari Hämäläinen, Mark Patel, Sven Blumberg, Tanguy Catlin, and Wasim Lala from QuantumBlack, position agentic economics as the defining operational challenge for enterprise AI in this phase of adoption. With 93% of surveyed enterprises already over budget and one-fifth actively pulling back on deployment scope, the window to build proper cost-value instrumentation is narrowing.

What this means for your team

  • Audit where your agentic token spend actually concentrates. If you cannot attribute costs by workload phase, including retrieval, reasoning, and refinement, you cannot target the 60% that McKinsey identifies as the dominant cost driver.
  • Evaluate model-task fit across your agent stack. Deploying frontier models on routine tasks is a documented cost driver; map each agent use case to the minimum capable model tier and quantify the savings before your next budget cycle.
  • Build output-quality tracking alongside cost tracking. Measuring tokens spent without measuring accuracy, human correction rates, and task completion rates makes cost-value comparison impossible.
  • Use the McKinsey framing as an executive forcing function: for each active agent deployment, require a documented answer to whether the value of completed work exceeds the full operational cost of generating it.

Featured companies

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. See how AI describes your company today, and where competitors show up instead.

Free workspace

You just read one Software & Technology expert. Imagine publishing your whole team.

This article was produced through MarketScale. Create a free workspace and turn your own team's Software & Technology expertise into the articles, video, and social content B2B marketing buyers in your industry are searching for. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale Studio workspace
One video edit a month, on us
AI writing, editing, and publishing tools
In-platform coaching to learn the system

More Software & Technology Insights

Enterprise AI hits an inflection point: governance, agentic systems, and the ROI reckoning

Enterprise AI hits an inflection point: governance, agentic systems, and the ROI reckoning

Enterprise AI is transitioning from experimentation to a focus on accountability. Key areas now influencing success include agentic systems, budget scrutiny by CFOs, and robust data governance initiatives. These factors play a critical role in determining the efficacy and ROI of AI implementations in businesses.

  • 01Agentic systems are becoming crucial in enterprise AI for ensuring efficient, autonomous decision-making.
  • 02CFOs are scrutinizing AI investments more closely to ensure their alignment with budget constraints and ROI goals.
  • 03Robust data governance is essential in capturing the full potential of enterprise AI.

Jul 21, 2026

AI governance gaps are blocking enterprise scale-up across the Middle East

AI governance gaps are blocking enterprise scale-up across the Middle East

AI governance challenges are preventing many businesses in the Middle East from scaling up their AI deployment. While technology advancements are not a barrier, the lag in internal governance frameworks is slowing down AI adoption. Addressing these governance gaps can accelerate the integration of AI within enterprises.

  • 01AI adoption in the Middle East is hindered by outdated internal governance frameworks.
  • 02Technology is not the limiting factor for AI deployment; governance is.
  • 03Improving governance frameworks can accelerate AI scale-up in enterprises.

Jul 21, 2026

Palo Alto Networks CEO puts a number on the AI cost problem: 90% token price drop needed

Palo Alto Networks CEO puts a number on the AI cost problem: 90% token price drop needed

Nikesh Arora, CEO of Palo Alto Networks, stated that for enterprise AI to scale, token costs must decrease by 90% within two years. He highlighted that high costs have already impacted companies like Uber, which spent its full-year AI budget by April.

  • 01Token costs for AI need to decline by 90% in two years for scalability.
  • 02Uber exhausted its annual AI budget by April due to high costs.

Jul 20, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512