Local AI: Governance Before Hardware Purchase
Benedikt Langer
10 min readFour developments over two weeks show that locally operated AI goes far beyond the tech stack. ...
European companies rely almost exclusively on U.S. or Chinese providers for their AI infrastructure. Google’s Gemma 4-a high-performance open-source model running on proprietary hardware that matches frontier models in benchmarks-is reshaping the equation. For CIOs, the question is no longer whether local AI is powerful enough. The real question is how quickly they can scale up their own AI capabilities.
When IT organizations today integrate AI into business processes, they typically do so via APIs-OpenAI, Google, Anthropic, and increasingly DeepSeek and Alibaba. The integration is fast, the results are strong, and initial costs appear manageable. More on this in our article on Digital Sovereignty.
What often gets lost in the excitement: each of these integrations creates a strategic dependency. And strategic dependencies have a tendency to deepen over time-the more deeply embedded the integration, the harder the exit becomes.
Three scenarios that are not hypothetical:
Pricing control: OpenAI has adjusted its API pricing multiple times over the past two years-both up and down. Organizations that have built business processes around a specific cost-per-token model are exposed to these fluctuations. As usage scales, costs can grow faster than the value delivered.
Geopolitical risk: Export restrictions on AI technology are already a reality. The U.S. has significantly curtailed chip exports to China, while China regulates foreign access to its AI models. Europe finds itself in between-primarily a consumer, not a producer. What happens if a trade conflict restricts access to U.S.-based AI APIs?
Regulatory misalignment: The EU AI Act imposes requirements on transparency and documentation for AI systems. With cloud-based services, control over model behavior and training data lies with the provider, not the user. This creates a compliance gap-one that widens with every new regulatory development.
“Gemma delivers unprecedented performance per parameter. These models aren’t massive-they’re relatively small, ideally suited to run on your own GPU.”
– Google, Gemma 4 Announcement (April 2026, paraphrased)
The main objection to local AI so far has been: too weak, too complex, too expensive. With Gemma 4, Google invalidates all three arguments:
For comparison: Alibaba’s Qwen 3.5 delivers similar benchmark results but requires 397 billion parameters to do so-making it strictly cloud-dependent. Gemma 4 31B, by contrast, runs on a single machine that fits in any office.
Four model sizes (ranging from 2B to 31B parameters) cover the full spectrum from smartphones to workstations. The smallest variants process audio, video, and images directly on the end device. Larger models support function calling and structured outputs-the foundation for automated workflows without human intervention.
For IT decision-makers, the question is not a technological one, but a matter of governance: How much control over its own AI infrastructure does a company want-and can it afford-to retain?
Entirely cloud-based (the current standard for most organizations): Maximum quality with minimal in-house effort. However, this comes with maximum dependency and limited control over costs, data flows, and availability. Best suited for companies with low AI workloads and non-critical use cases.
Entirely on-premises: Maximum control and data sovereignty. But it requires GPU infrastructure, MLOps expertise, and the acceptance that performance will fall short of frontier-level models. Ideal for highly regulated industries and applications involving sensitive data.
Hybrid (the rational middle ground): On-premises models handle 70-80% of standard inference tasks-classification, summarization, data extraction, and routine operations-while cloud-based frontier models manage the remaining 20-30% involving complex analysis, strategic tasks, and creative applications. Routing requests based on data sensitivity and task complexity becomes the new architectural challenge.
The hybrid approach has a clear investment trigger: the on-premises foundation must be built now. GPU procurement, MLOps pipelines, routing logic, and access controls are all essential. Organizations that delay will only deepen their reliance on cloud providers-making any future shift significantly more expensive.
The leading open-source models originate in the U.S. (Meta, Google, Mistral) and China (Alibaba, DeepSeek, Zhipu). Europe produces hardly any of its own foundation models with comparable performance. This structural weakness has not yet been offset by initiatives such as Gaia-X or individual European AI startups.
What Europe *can* do, however, is run available open-source models on its own infrastructure-preserving at least operational sovereignty. Models licensed under Apache 2.0, such as Gemma 4, enable precisely this, eliminating dependence on the goodwill of the original developer.
For CIOs in DACH-region enterprises, this is the pragmatic solution to the sovereignty question: rather than waiting for European frontier models-which may never materialize-deploy the best available open models on in-house infrastructure. The license allows it. The hardware exists. And the quality is sufficient.
Three key points for the next strategic planning cycle:
Budget for GPU capacity: Local AI inference requires dedicated GPU resources. This is a new line item in the IT budget-but one with a clear return on investment. A single GPU workstation (€3,000-5,000) can replace monthly API costs of €500-2,000. Payback periods range from three to twelve months, depending on usage volume.
Build MLOps expertise: Setting up, updating, integrating into existing systems, and monitoring local models requires know-how that many IT teams currently lack. The effort involved is manageable-comparable to establishing a new database infrastructure-but it must be planned and funded accordingly.
Define a routing architecture: Which tasks should run locally, and which should use cloud APIs? Decision criteria include data sensitivity, task complexity, latency requirements, and cost. This routing capability will become a core competency for IT organizations-much like hybrid cloud decisions a decade ago.
The comparison to cloud migration is deliberate: back then, the challenge wasn’t “all or nothing,” but rather about finding the right balance. And just as in the past, companies that develop a proactive strategy early on will gain a significant advantage-instead of reacting to market trends after the fact.
No. Cloud-based AI remains the best option for the most complex tasks. The point is: not everything needs to go to the cloud. For the majority of AI workloads, local models now offer sufficient quality-with better control and lower costs. The smart strategy is hybrid, not dogmatic.
No. Apache 2.0 is a perpetual license-software once released under this license remains permanently free to use. Google could release future versions under a different license, but Gemma 4, as already published, stays under Apache 2.0. This is a key difference from proprietary cloud services, whose terms of use can be changed at any time.
A dedicated AI team isn’t necessary to get started. Setting up a local model using frameworks like Ollama or vLLM is achievable for experienced IT administrators within a day. For integration into business processes and ongoing operations, it’s advisable to assign responsibility to an existing team-such as infrastructure or platform-without making it a full-time role, but rather as an extension of their current duties.
Europe has a relevant player in open-source AI with Mistral (France), but it still lags behind Google, Meta, and Alibaba in model performance. The EU’s strategy focuses more on regulation (the AI Act) than on developing its own foundation models. For businesses, this means a pragmatic approach: run the best available open models on your own infrastructure to secure operational sovereignty-without waiting for European frontier models.
Yes, in the medium term. If 70-80% of standard inference workloads are handled locally, API usage with cloud providers will decline accordingly. However, total AI costs must be assessed holistically: lower API expenses are offset by investments in hardware, skill development, and infrastructure. The break-even point typically falls between three and twelve months, depending on usage volume and previous cloud AI spending.
Image source: Pexels