04.08.2026
10 min read

Four developments over two weeks show that locally operated AI goes far beyond the tech stack. Anyone who runs models themselves is buying operations, liability and new dependencies. That needs to be decided as a governance choice before a tool pilot becomes an operating standard.

Key takeaways

  • Responsibility moves in-house. Downloading model weights shifts liability and operations from the vendor contract into your own operations handbook, where both need to be budgeted.
  • Three terms, three legal consequences. Open weights, open source under OSI (Open Source Initiative) and the exception in Recital 104 of the EU AI Act are different things that legal must keep separate.
  • The business case regularly underestimates itself. Raw GPU costs are 30 to 40 percent of the investment; real utilization sits at 40 to 65 percent instead of the planned 80 to 90.
  • Flexibility beats lock-in. Anyone who locks in pure cloud operations with multi-year commitments today will later negotiate on a cost base they themselves have tied down.

RelatedSupervisory Board Governance: Building Digital Competence Deliberately  /  Records Management as a CIO Topic: Why Governance Needs Ownership

TL;DR

  • Downloading model weights shifts responsibility from the vendor contract into your own operations handbook.
  • Open weights, open source under OSI and the exception in the EU AI Act are three different things. Legal must keep them separate.
  • The decision only becomes robust once data class, owner, license, operations and exit are settled before the hardware purchase.

Four signals that belong together

The debate is running ahead of availability. In early August, rankings of the best open models are circulating in which a model sits near the top whose weights have not yet been released. Alibaba itself has announced them for the following week. Experts have corrected this, but the rankings came first. Anyone who draws procurement pressure from such sources is planning against a state of play that does not exist.

A vendor can change its positioning while already in use. Ollama has established itself as a tool for local AI and is now shifting models into its own cloud offering. In the community, critics say popular new models are being held back to drive subscriptions; a tool that keeps models behind a paywall is, in the end, just another cloud provider with an intermediate step. That is a community view rather than a judgment on the company. The structural point remains: a sovereignty strategy that rests on a single tool has only shifted the dependency.

The vendor landscape is polarizing in public. In late July, Anthropic published a position paper on models with open weights that was widely and controversially debated. Almost simultaneously, Moonshot released the weights of Kimi K3. A specialist debate has hardened into camps, forcing procurement to take a side rather than invoke market consensus.

The geopolitical dimension has entered the mainstream. In public discussion, the claim is circulating that a substantial share of companies is now relying on cheaper Chinese models with open weights. This figure comes from an interview and is not substantiated. It fails as market statistics, but works as a mood check: the question of US cloud versus Chinese open weights is now being asked at the leadership level.

Open Weights Are Not a Free Pass

The positions can be sorted. Sebastian Raschka puts the value of open models in plain terms: they create auditability, you can verify claims, and you can run models on your own hardware without handing personal data and intellectual property to closed providers. Andrew Chen of a16z adds the market view and sees frontier models as oversized for the vast majority of use cases. The counterposition is equally clear: many of these models cannot be run locally at all for lack of resources.

For compliance, a clean distinction matters. Open Weights means the weights are downloadable and the license governs use. Training data, training code, and reproducibility are not included. Open Source under the OSI (Open Source Initiative) definition (OSAID 1.0) additionally requires Data Information: origin, scope, acquisition and selection methods, labeling, processing, and filtering.

GLM-5.2 and DeepSeek V4 Pro are under MIT, a true OSI license. Even so, they are not Open Source AI under OSAID because the training data is missing. A permissive license on the weights and open-source AI are two different things. This is exactly where legal needs to look closely.

Kimi K3 carries a license with a revenue-sharing clause: anyone who runs the model themselves pays nothing. Anyone who offers it as a paid service for third parties is expected to share revenues. For companies that embed AI features in their own products, that is a contractually relevant condition rather than a footnote.

The EU AI Act tightens the situation further. Recital 104 requires publicly available parameters, weights, and architecture and usage information for the free and open-source model exception. But the exception covers only transparency-related obligations. The summary of training content and copyright compliance remain. The interpretation question is open: read strictly by OSI, almost every model falls outside the exception. Read more loosely as Open Weights, even risky models slip through. Procurement therefore needs its own license and AI Act review. The download link alone is not enough.

The curve you should not lock yourself into

Between these signals lies a development that barely shows up in procurement templates yet matters more for architecture decisions than any daily headline: the class of models that run on your own hardware is catching up with the top tier at a pace that classical investment cycles cannot match.

A verifiable example: In April 2026, Alibaba released a dense model with 27 billion parameters that runs on a single professional graphics card. On one of the harder benchmarks for agentic programming, it scores 77.2 percent, ahead of its own previous flagship with 397 billion parameters, which reached 76.2 percent. Roughly fifteen times smaller, slightly better, three and a half months apart. What was data-center class in spring was running at the desk in summer.

This is driven by competition applying pressure from both sides. Open-model vendors have followed up within weeks: Moonshot released the largest open model to date, DeepSeek and Zhipu are placing their top models under the MIT license, and Alibaba is opening its top model class for the first time at all. On the other side, closed providers are cutting prices and publishing policy positions on open weights that nobody considered necessary a year ago. Both camps are working the same curve right now.

Analyst expectations point the same way. Gartner expects inference on a model with one trillion parameters in 2030 to cost more than 90 percent less than in 2025, versus comparable models from 2022 at up to a hundredfold cost efficiency. Drivers are more efficient semiconductors, better model design, higher utilization, specialized inference hardware, and the use of end devices for certain use cases. On the share of local processing, the expectation is similarly clear: more than two thirds of enterprises with AI at the edge by 2029, versus around ten percent in 2025.

One caveat belongs expressly with this. It comes from Gartner itself: the savings do not land one-to-one with the user organization. Agentic applications consume significantly more tokens per task, which eats a substantial part of the price advantage again. Cheaper per token does not mean cheaper per business transaction.

The governance consequence is nonetheless clear. It is not “move everything in-house immediately.” It is: do not make any architecture decision that makes switching expensive. Anyone who locks in pure cloud operations today-with hard-wired vendor interfaces, multi-year lock-in, and no evaluation cases of their own-is betting against their own cost curve and loses negotiating position at every contract renewal. Conversely, a rushed hardware purchase is equally expensive because it cements a snapshot in time. The durable position sits in between: interchangeable inference endpoints, own evaluation cases instead of third-party rankings, and a data classification that already knows which processes would move in-house first as soon as it pays off.

What is realistic is a growing share: a part of the process landscape that belongs in-house for reasons of data class, volume, or latency, and that will grow over the coming years. That share should be known before the next framework agreement is signed.

Costs missing from the business case

One line from the debate has become a catchphrase: Open Weights, closed by hardware. When operations demand hundreds of thousands of GPUs, it is not a local model. It is a cloud model with downloadable weights. At the governance level that means: Downloading the weights is an operational decision with full responsibility in-house. What used to sit in the vendor contract then sits in your own operating handbook. That is legitimate. But it must be taken and budgeted as such a decision.

Raw GPU costs make up only 30 to 40 percent of the true infrastructure investment. As a rule of thumb, apply a factor of 2.5 to 3 on the hardware price. Real utilization in production sits at 40 to 65 percent, while many calculations still assume 80 to 90 percent and thereby overstate the benefit. The productive inference stack adds another 10 to 20 hours per month for updates, troubleshooting, monitoring and capacity planning.

Break-even against frontier model API prices sits at about 2 to 5 million tokens per day over twelve months. Against low-cost open model providers only at 50 million tokens per day and more. Anyone who misses these thresholds is funding a sense of sovereignty with local AI rather than a solid cost advantage.

On top of that comes operational drift. One documented case saw a model update raise video memory demand by 10 gigabytes. An update can render hardware planning obsolete after the fact. Anyone who only approves the purchase price and ignores operations, updates and capacity risks is green-lighting a stack without a solid operating plan.

Between the CLOUD Act and Open Weights

German companies rarely weigh the USA against China. They stand between US hyperscalers with CLOUD Act exposure and Chinese models with open weights. Local operation is the option that sidesteps both dependencies. It costs exactly what is stated above.

The tiers by data sovereignty rise as follows: US cloud with residual CLOUD Act risk, EU cloud with better data residency but still exposed via a US parent company, sovereign cloud, and on-premises, where data never leave the premises.

Important for precision: the GDPR does not mandate local operation. It requires protection, contracts, and data residency. Anyone who sells local AI as a GDPR obligation is arguing incorrectly. Local is one of several architectures that can achieve compliance. The right question is which data class may leave the perimeter and which protection goals can be achieved reliably with which operating model.

Company size is not a threshold here. A tax advisory firm with twelve people holds client data that often must not leave the premises. A large corporation with a zero-trust cloud standard stays hybrid. Size correlates with budget, staff, and risk appetite. The driver of the decision lies in data class and operational capability.

Eight questions before the decision

A board does not need false precision. It needs eight questions that can be answered yes or no. If the no answers pile up, local AI is not ready as a standard stack. Pilot, hybrid or cloud then remains the honest path.

  1. May or should the data class concerned leave the perimeter?
  2. Has an owner been named who brings together IT, the business unit and security?
  3. Is the use case defined tightly enough? Or is a general-purpose tool being procured?
  4. Does the quality of mid-tier open models demonstrably suffice, measured against own cases?
  5. Have continuous operations, updates, monitoring and model updates been budgeted alongside the hardware?
  6. Has the license been checked, including commercial use and redistribution in your own product?
  7. Have power capacity, cooling and floor space been clarified?
  8. Is there an exit strategy in both directions, including back to the cloud?

The result is read as a go or stop signal, never as a score. Every no is an open residual decision. It remains open only until someone takes on responsibility and budget.

Key Takeaways

For the governance layer, three consequences matter-regardless of which camp ultimately proves right.

The procurement basis must be decoupled from the news cycle. If a ranking rates a model that does not yet exist, it is unfit as a decision template. What is binding is what the company has measured itself on its own use cases.

License review belongs before the hardware purchase, not after. A permissive license on weights is not a free pass. Profit-sharing clauses are contractually material. The exception in the EU AI Act is narrower than its name suggests. This review costs days. A wrongly procured GPU estate costs quarters.

Sovereignty is an operating model. A download is not enough. It begins where responsibility, budget, and exit are co-decided. Local AI pays off where data class, volume, and risk support in-house operation. Without that foundation, the local stack remains an expensive detour with the feeling of control.

And the question that will have a different answer in twelve months than it does today: What share of your own process landscape belongs inside? For most organizations today, the answer is still: a small one. It will not stay that way permanently. When a model on a single card reaches, within one quarter, the performance of a predecessor fifteen times larger, and two analyst forecasts independently point in the same direction, the relevant preparation is not the hardware purchase. It is the ability to shift that share without rebuilding the architecture. Those who write this flexibility into contracts and interfaces today will decide later from a position of strength. Those who lock themselves out of it will, in two years, negotiate over a cost base they themselves have bolted down.

Frequently Asked Questions

Does the GDPR require AI models to run locally?

No. The GDPR (General Data Protection Regulation) requires protection, contracts and data residency. Local operation is one of several architectures that can meet these requirements, alongside sovereign cloud and contractually secured EU processing. Anyone selling local AI as a data protection obligation is arguing incorrectly and thereby weakens their own position.

Is an MIT license on the weights enough for a compliance review?

It is a good starting point and does not answer the question alone. A permissive license on the weights does not make a model open source under the OSI definition, because details on the training data are still missing. For the AI Act exception, it also matters whether the model poses a systemic risk.

Which license clause is most often overlooked in procurement?

Revenue-sharing clauses. Kimi K3 is free for self-hosting, but requires a revenue share once the model is offered as a paid service to third parties. For any organization that embeds AI functions in its own products, this is a contractually material condition rather than a footnote.

Should a company buy GPU hardware now?

For most organizations, the answer is still no. It will not stay that way forever. The preparatory work lies elsewhere: interchangeable inference endpoints instead of fixed vendor interfaces, own evaluation cases instead of external rankings, and a data classification that knows which processes would move in-house first.

Read more on Digital Chiefs

Digital ChiefsAI Regulation: Up to 3 Percent of Corporate RevenueDigital ChiefsYou are paying for the R&D of the next competitorDigital ChiefsModel Harness Instead of Model Marriage: Who Controls the AI Chain?

More from the MBF Media Network

Image source: AI-generated (August 2026)

Share this article:

Also available in

More Articles

03.08.2026

AI Regulation: Up to 3 Percent of Corporate Revenue

Tobias Massow

5 min read Article 50 of the AI Act has bound providers and deployers to concrete transparency obligations ...

Read Article
31.07.2026

You are paying for the R&D of the next competitor

Benedikt Langer

4 min read You are funding the R&D of your next competitor and calling it AI transformation. Frontier ...

Read Article
29.07.2026

Model Harness Instead of Model Marriage: Who Controls the AI Chain?

Eva Mickler

6 min read The lock-in is shifting from the individual model to the orchestration layer. Those who don’t ...

Read Article
28.07.2026

Washington decides which AI is allowed to run here

Eva Mickler

6 Min. read time In just eight days, Washington has shifted the dispute over Chinese AI models from ...

Read Article
23.07.2026

Orphaned Access: The Silent Cybersecurity Gap

Benedikt Langer

5 Min. Read Time Service accounts, API keys, and AI agents often outnumber human accounts. Many of these ...

Read Article
A magazine by Evernine Media GmbH