Skip to content
MarketScale
‹ Back to IndustriesSoftware & Technology

60% of agentic AI costs go to response refinement, and most enterprises are already over budget

A McKinsey study reveals that 93% of enterprises exceed their AI budgets as agentic AI systems expand. A significant portion, 60%, of AI costs are directed towards refining responses. The cost structures for these systems are often not fully understood by many operators.

This story was produced through MarketScale. See how Software & Technology teams put it to work with Executive Thought Leadership.

By MarketScale Newsroom · MckinseyAgentic AiAi AgentsEnterprise Ai
Share
Listen to the audio brief

Key facts, context, and what it means.

AUDIO
0:00
60% of agentic AI costs go to response refinement, and most enterprises are already over budget

Key takeaways

01

93% of enterprises exceed AI budgets as agentic systems scale.

02

60% of agentic AI costs are allocated to response refinement.

03

Many operators have not fully mapped the cost structures of agentic systems.

Get featured

Want to get featured in MarketScale Software & Technology?

Create a free MarketScale workspace and get your company's expertise featured across our Software & Technology coverage. No credit card, no demo required.

Start free

Token prices have collapsed. Inference on GPT-3.5-level capability cost $20 per million tokens in early 2024 and fell to $0.07 by year's end, according to the Stanford HAI 2025 AI Index cited in McKinsey's new report. Enterprise AI spending is accelerating anyway. LLM expenditures tripled over a 12-month period by the end of 2025, according to a Menlo Ventures study also cited by McKinsey, and 93 percent of organizations surveyed by McKinsey in May 2026 reported exceeding their AI budgets. Cheaper tokens, it turns out, do not automatically mean cheaper AI programs.

The explanation sits inside the cost structure of agentic AI, which differs fundamentally from the simple prompt-and-response model most budget forecasts were built around. McKinsey's July 2026 report, published in the McKinsey Quarterly and authored by researchers from QuantumBlack, AI by McKinsey, pinpoints response refinement as the dominant cost driver: 60 percent of total agentic AI spend goes not to the first inference call but to the iterative cycles of checking, correcting, and improving that agents run before delivering a usable output.

Why the cost structure of agents caught enterprises off guard

The shift from subscription to consumption pricing by major LLM providers is one structural cause McKinsey identifies. Under consumption models, answer length directly affects the bill, creating an incentive architecture that rewards verbosity. Enterprises are also routinely routing straightforward tasks through frontier models that are priced for complexity, compounding the overspend. Neither dynamic was fully visible when organizations set their 2025 and 2026 AI budgets.

The scale-up effect amplifies both problems. Companies that began with pilots at controlled token volumes are now running agents across production workflows, and the cost curves are non-linear. A single agent orchestrating multiple sub-agents, each refining its outputs before passing results upstream, can generate a token bill that is orders of magnitude larger than a direct LLM query producing the same end result.

Sixty percent of agentic AI costs sit in response refinement, a cost center that most enterprise budget models never built a line item for.

The breadth of the budget problem is striking. McKinsey's Enterprise AI FinOps Survey, conducted in May 2026 with 75 qualified respondents across five major industries, found that 93 percent have already blown past their AI budgets. Separately, one in five participants in McKinsey's forthcoming 2026 State of AI survey, which gathered responses from 1,719 participants between May and June 2026, said their organizations have actively constrained AI use because of operating costs. That is a notable reversal: AI adoption being slowed not by capability gaps or organizational resistance but by economics.

Token cost is the wrong target metric

McKinsey's central argument is that obsessing over token price reduction is a category error. The report cites David Tepper, CEO of Pay-i, making the distinction plainly: tokens are not value, tokens are the bill. The relevant question for enterprise operators is whether the output an agent produces is worth more than the full cost of producing it, including refinement cycles, human supervision time, and the cost of correcting errors downstream.

That framing shifts the evaluation criteria considerably. An agent that costs three times as much per task but requires no human review and produces zero rework may be the cheaper option in total. Conversely, an agent running on a cheaper model but requiring frequent human correction may cost more in fully loaded terms. Neither the token price nor the model tier alone tells the story, according to McKinsey's analysis.

The report lays out several variables that determine real agent value: output correctness, human supervision and repair burden, compute consumed during reasoning, and whether the completed work exceeds the full operational cost of generation. McKinsey frames these as the levers a CEO needs to understand as AI systems evolve from agentic coworkers toward more autonomous multi-agent architectures.

What this means for enterprise FinOps and procurement teams

For operations and IT leaders, the practical implication is that current AI cost models are almost certainly undercounting. If 60 percent of the bill sits in refinement loops rather than primary inference, and most budget frameworks were built around inference costs alone, the gap between forecast and actual spend will widen as agentic deployments scale. The McKinsey findings, reported by the Economic Times, suggest that the next phase of enterprise GenAI adoption will be shaped less by which models organizations choose and more by how well they instrument and govern agent behavior end to end.

Model selection, routing logic, and output validation are no longer just engineering decisions. They carry direct budget consequences that procurement and finance teams need to co-own with their technology counterparts. Organizations that build FinOps disciplines around agent-level cost attribution, rather than aggregate LLM spend, will be better positioned to scale without the budget overruns that are already affecting the majority of enterprise AI programs.

Where agentic AI costs go
McKinsey & Company, July 2026 · © MarketScaleDownload chart

McKinsey's next data point to watch is its full 2026 State of AI report, which will carry responses from 1,719 global participants and is expected to detail how constrained AI budgets are reshaping deployment priorities across industries. The findings from its May 2026 FinOps survey already suggest that the budget conversation has moved from the IT team to the C-suite, and it is unlikely to move back.

Featured companies

Your experts belong here

Every story in MarketScale Software & Technology starts with a company putting its solutions engineers, product teams, and customer engineers on the record. Buyers are already reading this topic. The only question is whose experts they find.

Buyers ask AI engines who to consider, and published expert answers are what those engines cite.

Get your team featuredSee how it works15 minutes, straight to a calendar.

About the author

MarketScale Newsroom
MarketScale NewsroomEditorial Team, MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

Follow Software & Technology Insights

Get new expert content in your inbox.

Software & Technology: are you visible to AI?

Before they reach out, Software & Technology buyers ask AI engines which vendors to trust. Explore how your experts, customers, and partners can become useful content for buyers and AI search.

Free plan

You just read one Software & Technology expert. Your company is full of them.

This article was produced through MarketScale. The same platform turns your solutions engineers, product teams, and customer engineers into the articles, video, and social content Software & Technology buyers are searching for. Create a free workspace and see it with your own people. No credit card, no demo required.

NPS +73 · 1,000+ creators · 38+ countries

What you get, free

Your own MarketScale workspace, up to 10 people
One professional video edit a month for qualifying companies
Media requests to your crowd, remote recording, AI writing tools
$0, no credit card, nothing that expires

More Software & Technology Insights

Nvidia CEO Jensen Huang Says Chip Sales Could Double Next Year—If the Supply Chain Can Keep Up

Nvidia CEO Jensen Huang Says Chip Sales Could Double Next Year—If the Supply Chain Can Keep Up

Nvidia CEO Jensen Huang forecasted doubling chip sales next year, but the company's CFO frames this as the supply-unconstrained scenario, signaling that supply chain capacity—not demand—is the real constraint. Nvidia and Palantir announced a collaboration to apply AI to Nvidia's own supply chain operations to identify bottlenecks and allocate materials more effectively.

  • 01Nvidia's doubling forecast depends on supply chain throughput, not demand—the company itself is supply constrained according to CFO Colette Kress.
  • 02Nvidia and Palantir said their first sovereign AI deployment for Nvidia’s operations is designed to spot supply constraints earlier and improve how materials are allocated across production.
  • 03Enterprise buyers should plan for competitive allocation pressure, higher networking and infrastructure costs alongside GPU spending, and the emergence of on-premises architectures as first-class options.

Sep 20, 2026

Fifth Third, Priority and CSI deals put a premium on payments built into software

Fifth Third, Priority and CSI deals put a premium on payments built into software

Fifth Third led a strategic investment in Payload, Priority Commerce agreed to acquire IntelliPay, and CSI acquired Qolo in a series of summer transactions, PYMNTS reported. Together, the deals point to buyers valuing payments technology already integrated into the software customers use, not just standalone processing capacity. For operators, that means the entity holding payment data can change hands without the front-end software changing.

  • 01BCG says SaaS providers with integrated payments accounted for 36% of small and midsize business acquiring revenue in 2024 and projects that share will reach 45% by 2028.
  • 02Finance and IT leaders using property, practice management or utility billing platforms should reread payments and data clauses, since the entity holding payment data can change hands even if the software front end does not.

Sep 19, 2026

System integrators shape whether factory tech pays off, Smart Industry argues

System integrators shape whether factory tech pays off, Smart Industry argues

Smart Industry’s Sept. 9, 2026 piece argues plant technology creates no business value until system integrators fit it into existing systems, operations and workflows. Related summer coverage highlights upskilling, institutional knowledge and technician demand alongside the same integration-and-deployment theme. The framing shifts attention from which platform to buy to who implements it and how the engagement is scoped.

  • 01Smart Industry's framing moves the buying question from which platform to license to who integrates it and how that engagement is scoped, which puts the system integrator line item at the center of the return rather than in implementation overhead.
  • 02Gartner figures cited by Quality Magazine show 24% of industrial enterprises using IoT have implemented digital twins and 42% plan to, suggesting most IoT-using plants still have digital twin integration work ahead.
  • 03The Deloitte and Manufacturing Institute report, as covered by Smart Industry, says AI can embed skills into workflows to address technician demand; the sharper question for a plant manager is whether that changes headcount or changes what each technician can cover.

Sep 18, 2026

Explore More Software & Technology Insights

Read more expert perspectives from across Software & Technology.

Browse Software & Technology Hub

About the Expert

MarketScale Newsroom
MarketScale Newsroom

Editorial Team

MarketScale

The MarketScale Newsroom reports on the companies, technologies, and trends shaping 16 B2B industries. It turns primary sources and expert commentary into clear, useful coverage for the people doing the work.

For B2B teams

Your experts could be publishing here

Stories like this one run on content MarketScale captures from real practitioners. See how your team's expertise becomes coverage in Software & Technology and beyond.

Book a 15-minute demo

Or call us. No forms required. We pick up. 214-945-2512