01.08.2026
7 min read

98 percent of FinOps practitioners now manage AI spend-up from 31 percent two years ago. Tokenomics has shifted from an engineering detail to a boardroom topic for CIOs, CTOs, and executives.

Key Takeaways

  • AI spend is now standard: According to the State of FinOps 2026 report, 98 percent of practitioners manage AI costs. AI cost management is the most frequently cited skill requirement.
  • C-level influence grows: 78 percent of FinOps practices report to CTOs or CIOs. On June 3, 2026, the Linux Foundation announced the Tokenomics Foundation.
  • Maturity and impact: McKinsey estimates only 20 to 25 percent of companies have mature AI FinOps practices. Thoughtful consumption strategies can save 20 to 30 percent of AI costs.

RelatedThree AI budgets, no unified accounting  /  Quantum computing: The answer to AI’s energy hunger?

FinOps takes charge of AI costs-and sits closer to the tech leader

The State of FinOps 2026 survey by the FinOps Foundation draws on responses from 1,192 participants and represents an annual cloud spend exceeding €76 billion. The findings are clear: AI cost management has moved from a fringe concern to a core responsibility. Ninety-eight percent of FinOps practitioners now manage AI spend, compared to just 31 percent two years ago. AI cost management is also the most frequently cited skill gap teams say they need to fill.

This shift is reshaping organizational structures. Seventy-eight percent of FinOps practices now report to CTOs or CIOs. With C-level and VP engagement, influence over technology selection is growing significantly. For digital leaders in DACH organizations, this means AI budgets and model decisions must be integrated into the same governance loop as cloud portfolios and vendor governance-once FinOps is visibly aligned with the technology executive.

The survey launches on February 19, 2026, marking the empirical starting point for a broader debate that will expand throughout the year to include tokenomics and open standards. Those who treat AI merely as a feature pipeline-and clean up costs afterward-risk falling behind the new norm of the discipline.

Value, self-funding, and the pressure behind AI growth

Pure waste reduction is no longer enough in FinOps practice. Scopes, governance, alignment, and forecasting now carry more weight than optimization alone. The goal is value-the contribution of cloud and AI investments to business outcomes-not just another percentage point of cost reduction.

Many organizations are expected to self-fund AI investments from optimization savings. This self-funding model ties efficiency pressure directly to AI growth: every euro freed up through disciplined cloud and inference practices can fuel new models, agents, or use cases. For executives, this creates a dual mandate. Efficiency programs finance innovation, and innovation in turn drives consumption-unless governance and consumption strategies scale in tandem.

McKinsey’s findings underscore the maturity gap. In its Enterprise AI FinOps Survey from May 2026 (referenced in July 2026), only 20 to 25 percent of companies have mature AI FinOps practices. At the same time, 20 to 30 percent of AI spend often remains unaccounted for. McKinsey estimates that thoughtful consumption design can save 20 to 30 percent of AI costs. The strategic implication is clear: without proper allocation and consumption design, self-funding remains a hopeful aspiration rather than a controllable cycle.

Tokenomics turns the token into the atomic control unit

Tokenomics, according to J.R. Storment, Executive Director of the FinOps Foundation, is FinOps applied to AI. The token is the atomic billing and control unit for inference-not an ownership token or a speculative metric. Storment framed this in May 2026: token costs and efficiency now belong on the CEO agenda.

Tokens behave differently from traditional cloud resources. User-input tokens and the prompt tokens actually billed at the API endpoint can diverge. Real-time monitoring of token and API metrics is essential, the FinOps Foundation states, to avoid budget overruns. Reading monthly invoices alone means steering too late.

At the same time, tokens are only one layer in the AI-cost stack. Beneath them sit GPU and compute, SaaS embedding, shadow AI, and engineering and governance overhead. A purely token-centric view remains incomplete. CIO Dive summed it up in June 2026: tokens are only one slice of AI spend, but the easiest to meter. That’s precisely why they’re a good entry point-without replacing the rest of the stack.

Five consumption drivers and falling unit prices amid rising aggregate spend

Five factors govern token consumption per request: system-prompt overhead, context and memory, model choice, output length, and retries plus orchestration. Agents and RAG can inflate consumption by one to two orders of magnitude. For portfolio management, that means the single chat is rarely the most expensive unit; instead, it’s the orchestrated chain of retrieval, tools, and multiple calls.

Model choice remains a direct lever. Using the most expensive and complex models for every task drives unnecessary costs. The model must match context and quality needs. FrugalGPT, the 2023 research by Lingjiao Chen, Matei Zaharia, and James Zou, provides academic backing: model cascading and routing can cut up to 98 percent of costs versus a GPT-4 baseline while maintaining comparable performance. For CIOs, this proves routing and cascading belong in standard architecture.

Unit prices per token may fall even as aggregate spend rises. Demand is elastic: cheaper units generate more usage, longer contexts, and more agent runs. Storment crystallizes the paradox: a token can become cheaper, yet total token spend does not. Budget planning that merely extrapolates the list price per million tokens misses this volume expansion.

McKinsey defines tokenomics more broadly to include prompt and inference costs: model selection, routing, orchestration, agent behavior, workflow design, infrastructure utilization, and waste elimination. This elevates architecture and operating model into the same control loop as the token counter.

Standards, quotas, and stakeholder governance for AI at scale

On 3 June 2026, the Linux Foundation announced plans to launch the Tokenomics Foundation for open standards, benchmarks, and best practices in the AI-cost economy, in partnership with the FinOps Foundation. Jim Zemlin, CEO of the Linux Foundation, calls tokens the new unit of technology spend and notes that neutral standards for token efficiency across models and vendors are missing. J.R. Storment reiterates that token costs and efficiency must be negotiated at CEO level.

FinOps Foundation best practices include quotas, tagging, GPU allocation, and stakeholder governance involving data science, procurement, and finance. CIO Dive reported in June 2026 that the Tokenomics Foundation will collaborate with the FinOps Foundation to develop enterprise AI-at-scale best practices. For DACH organizations juggling multiple cloud and model vendors, the message is clear: build internal measurement and tagging systems that can plug into future industry benchmarks.

No primary sources argue that tokenomics is unnecessary. The counterweights lie within the discipline itself: optimization alone is insufficient, the token is not the entire stack, and maturity levels in most companies remain low. These internal tensions make governance both challenging and board-ready.

What CIOs and CDOs Must Lock Down in Their AI Portfolio Now

First: route AI spend into the FinOps reporting line to the CTO or CIO and assign clear ownership for quotas, tags, and budget alerts. Second: embed consumption design and model routing as architectural principles instead of merely auditing invoices. Third: make self-funding transparent-show which savings feed which AI investments and what risks those investments create for aggregate spend.

Fourth: map the AI cost stack-tokens, GPU and compute, embedding SaaS, shadow AI, and governance overhead. Fifth: institutionalize stakeholder rounds with data science, procurement, and finance. Sixth: demand real-time monitoring of token and API metrics so budget overruns become visible before month-end close.

Teams that pull these levers manage AI as a portfolio with measurable value logic. Data from the State of FinOps 2026 report, the FinOps Foundation’s tokenomics framework, and the upcoming Tokenomics Foundation from the Linux Foundation all point the same way: AI cost economics is shifting from the edge of the cloud bill to the center of the boardroom and becoming a leadership tool for enterprise technology decisions.

Frequently Asked Questions

What is tokenomics in the FinOps context?

Tokenomics applies FinOps principles to AI. The token serves as the atomic unit of billing and control for inference. McKinsey also includes model selection, routing, orchestration, agent behavior, workflow design, infrastructure utilization, and waste elimination under this umbrella.

Why are AI costs rising despite falling token prices?

While per-token prices may decline, demand can grow elastically. Increased usage, longer contexts, agents, and RAG drive aggregate spending. Tokens become cheaper individually, but not automatically for the business in total.

Which roles should be part of AI cost governance?

The FinOps Foundation highlights quotas, tagging, GPU allocation, and stakeholder governance involving Data Science, Procurement, and Finance. Already 78 percent of FinOps practices report to CTOs or CIOs, increasing influence over technology choices.

Are token metrics sufficient for managing AI spend?

Tokens represent the most measurable layer and suit real-time monitoring. Yet they cover only part of AI spend. GPU and compute costs, SaaS embedding, shadow AI, and engineering and governance expenses belong in the same stack.

What proven savings potentials exist?

McKinsey cites 20 to 30 percent savings through thoughtful consumption. FrugalGPT reports up to 98 percent cost reduction via model cascading and routing while matching GPT-4 baseline performance. The exact levers depend on architecture and use-case mix.

Read more on Digital Chiefs

Digital ChiefsAI Shifts Development into Review Work. Who Owns Sign-off?Digital ChiefsYou are paying for the R&D of the next competitorDigital ChiefsFalse Sense of Security: When Cyber Police Fail to Act in a Real Emergency

More from the MBF Media Network

cloudmagazinFrontier Labs are eating their best customers mybusinessfutureSamsung Q2: Memory remains tighter than expected securitytodayAnthropic: Claude cracked three companies

Image source: AI-generated (August 2026)

Share this article:

Also available in

More Articles

09.09.2026

Nvidia Buys Hugging Face for Over 11 Billion Euros

Eva Mickler

4 min read Nvidia is acquiring Hugging Face for around 11.1 billion euros; the contract was signed on ...

Read Article
08.09.2026

SAP Lets Joule Steer Robots Directly, Liability Still Open

Bernhard Liebl

4 min read SAP has documented the first Embodied AI Jam at the Swiss Smart Factory in Biel. Inspection ...

Read Article
15.08.2026

ChatGPT wants to read the Mac

Eva Mickler

6 min read On 13 August 2026, OpenAI described Computer History for the ChatGPT Mac app in its release ...

Read Article
14.08.2026

SpaceX acquires Cursor: EU clauses stay open

Eva Mickler

5 min read The purchase agreement was finalized on 14 August 2026. Any company using the tool now has ...

Read Article
13.08.2026

CRA forces manufacturers to report within 24 hours

Bernhard Liebl

9 min read On 11 September 2026, Article 14 of the Cyber Resilience Act comes into force. From that ...

Read Article
11.08.2026

NVIDIA capital plans and what operators must check now

Bernhard Liebl

7 min read On 10 August 2026, NVIDIA announced it will partner with six capital partners to build financing ...

Read Article
A magazine by Evernine Media GmbH