Local AI: Governance Before Hardware Purchase
Benedikt Langer
10 min readFour developments over two weeks show that locally operated AI goes far beyond the tech stack. ...
5 Min. read time
Hyperscalers continue to expand. Yet analysts and earnings calls point to a slower growth rate in capital expenditures. For multi-year AI cloud contracts, the commitment scenario is key: what happens if capacity, pricing, or exit terms no longer align with strategic decisions?
Key Takeaways
RelatedHardware beats software deals – rethinking capex / Token OPEX: Inference drives costs, not seat budgets
AI cloud contracts here refer to multi-year agreements covering compute capacity, models, or inference performance: volume, pricing, reserved capacity, and exit terms. They tie budgets to a usage scenario and thus require tiered commitments, re-openers, portability, and exit cost structures spanning multiple operational years.
The public debate fixates on stock prices and individual earnings weeks. Procurement teams focus on something else: multi-year commitments for cloud and AI capacity. Those securing 2027 and 2028 are buying into a world where hyperscalers’ capital expenditures may remain absolutely high-while capex growth rates slow across multiple market scenarios.
This isn’t a contradiction. It’s the difference between level and slope. A contract that only maximizes today’s discount ignores the slope. One that fears the slope alone forfeits capacity. What works is a scenario matrix that translates both dimensions into contractual clauses.
Absolute capex levels of major cloud and AI infrastructure providers remain in a league that shapes regional capacity, GPU availability, and enterprise discount margins. At the same time, analyst projections and earnings guidance increasingly describe a flattening growth rate-not a halt in expansion.
For DACH procurement teams, this doesn’t spell panic or a blank check. It raises a planning question: What pace is baked into your internal AI TCO model? If you price in linear growth and hit a plateau, you’re stuck with commitments you can’t use. If you price in a plateau and face a growth scenario, you’re left without capacity-and paying spot prices.
Decision Unit
The unit is commitment per scenario. Plus exit per year and the threshold where in-house inference becomes cheaper and more controllable than the default API route.
Token and runtime costs (see related article on token OPEX) remain the variable layer. The providers’ capex pace drives the layer beneath: how expensive and how scarce physical and contractual capacity becomes for running inference. Both layers belong in the same decision template-measured separately, negotiated together.
Instead of a single forecast, three robust scenarios suffice. Each requires the same KPIs: GPU/TPU hours used, cost-per-outcome for core workflows, share of reserved vs. on-demand capacity, exit costs, and time-to-replatform in months.
| Scenario | What’s Happening in the Market | What the Contract Must Deliver |
|---|---|---|
| Growth | Capex and regional capacity continue to expand rapidly; reserved capacity becomes scarcer. | Early option windows for additional quotas; price caps on overruns; multi-region fallback. |
| Plateau | Levels remain high, but growth slows; discounts and availability stabilize unevenly. | Tiered commitments instead of one-time pledges; annual re-openers; clear definition of “unused commit.” |
| Slowdown | Growth decelerates noticeably; providers prioritize margins and utilization of existing assets. | Exit without automatic penalties; workload portability; on-prem/private inference as a negotiated path. |
The scenarios are planning frameworks-not market forecasts. Earnings and research data serve as inputs, never as contractual gospel.
When you populate all three scenarios with the same KPIs, it becomes clear where the planned framework agreement only holds up in a growth phase. That’s the moment procurement and architecture need to speak the same language: less “more discount,” more “which clause survives the plateau.”
Discounts are visible. The costly mistakes lurk elsewhere. Four types of clauses separate a negotiable AI cloud framework from wishful thinking.
1. Tiered Commitments Instead of Front-Loading. Instead of locking in 100 percent of the three-year volume in year one, tiered commitments tie tranches to usage milestones. Each tranche includes an opt-out with a deadline. This often costs a bit of discount-but it avoids the more expensive mistake: committing without workload.
2. Exit and Portability with a Checkpoint. An exit without data and model portability is just theater. The contract specifies formats, export deadlines, and one test run per year. If a provider dodges the test run, they know the exit doesn’t really exist.
3. Reserved Capacity with Expiry and Reallocation. Reservations without reallocation to related SKUs and without an expiry mechanism create dead quotas. The reallocation rule belongs in the main contract, not buried in a provider’s FAQ.
4. Price and Capacity Triggers. If regional capacity drops below an agreed threshold or the reference price of a defined SKU rises by more than X percent, a renegotiation window opens. X is set internally-not by the vendor’s default.
The cheapest cloud discount is expensive if it assumes a scenario your AI rollout can’t sustain.
On-prem or private inference isn’t an alternative to the cloud-it’s the switch in plateau and ramp-up scenarios with tight capacity. The decision hinges on three measurable factors: stable workloads with high token density, predictable latency requirements, and the ability to run eval suites and routing in-house.
If you can’t measure these three factors, you’re still buying the API default-and you should call it what it is. If you can measure them, negotiate a smaller cloud tranche and keep an internal path open. Make-or-buy for models is the second axis: a German or European model often only makes sense when inference costs and data residency are part of the same equation.
Hardware capex in-house (servers, cooling, power contracts) shifts the risk-it doesn’t eliminate it. The related discussion on hardware vs. software capex prioritization remains the sister debate. Here, the only thing that matters is the interface: Which workloads have a solid break-even case against the hyperscaler path-and are they in the contract as an option, not just a PowerPoint?
The test is simple: Would you still sign if the provider’s capex growth in 2027 turns out significantly flatter than in 2025/26? If the answer depends on a single discount, the contract is too weak.
Only if workloads are truly portable and the second provider has the capacity. Multi-cloud without an exit test and data path is just double commitment-hardly a hedge.
As a starting assumption, yes-but not as the controlling factor. Contracts should be governed by usage, tiered pricing, and triggers, not the next earnings slide.
When load is stable, evaluation criteria are clear, and cost-per-outcome is measurable. Without all three, the API default stays the more honest-and cheaper-option for governance.
Front-loaded commitments without tiered pricing or annual portability tests. The discount looks great-until you see the unused tranche.
Token OPEX drives the variable usage layer, while provider capex volatility dictates availability and price discipline for underlying capacity. Both belong in the same decision framework.
Read more on Digital Chiefs
Digital ChiefsRigid RZ Contracts Meet the Flexible EnEfGDigital ChiefsWhen the Cloud Is the Wrong ChoiceDigital ChiefsHow to Stifle Open Source Without Banning ItMore from the MBF Media Network
cloudmagazinWhen GPUs eat into the SaaS budget mybusinessfutureAI make-or-buy: Build in-house or outsource? securitytodayNIS2 patchwork: Four countries face ECJ rulingImage source: AI-generated (July 2026)