07.01.2025
3 min read

Some experts predicted it two years ago, and now it seems to be happening: AI is running out of high-quality training data! Major providers are therefore increasingly looking for alternatives.

Wikipedia, as a vast online encyclopedia, has an equally large global following. Many of its users actively contribute to expanding the knowledge base of over 60 million entries (as of early 2024).

AI, unfortunately, cannot rely on such a human following to increase the data underlying its knowledge. Training AI requires a constant supply of new, high-quality data. As t3n reports, developers of AI systems have often turned to freely available magazines and academic publications on the internet in the past.This source is gradually drying up. Major providers have therefore already signed contracts with publishing houses such as Springer, Reuters, or The New York Times. Scientific archives and communities like Reddit and Stack Overflow are also being used to nourish and expand AI knowledge. However, according to t3n, the problem is that these sources grow too slowly to satisfy the training appetite of increasingly sophisticated AI models.As early as 2022, there were warnings that the knowledge sources for AI training would run dry by 2026 at the latest. Other experts had assumed this would happen two years later.Some developers, such as Alphabet, the parent company of Facebook, simply resort to less reliable sources for their training purposes.

Another approach, adopted by companies like Anthropic since the Opus version of the Claude model series, is to use so-called synthetic data for training. OpenAI, the creator of ChatGPT, is also reportedly using this trick for its new language model, Orion. Synthetic data are not human-created data that mimic real data.

They are generated through computational algorithms and simulations based on generative AI technologies. Instead of expensive conventional data acquisition, this allows data sets of any size to be generated quickly and easily and adapted to specific requirements.

Blessing and curse of synthetic data

The method of relying on synthetic or low-quality data from social media posts for AI training is not without controversy among researchers. This is because it can negatively impact the quality of the output.

With synthetic data, it remains unclear how AI should continue to learn if it only has self-generated data available for training. Moreover, there is a risk that AI models may start to limit themselves if they imitate self-generated training data.

Synthetic data can even render AI unusable, as an experiment by Stanford University demonstrates.

The resulting training can lead to both errors and, in the best-case scenario, artifacts in AI responses, ultimately resulting in completely unusable answers. Among researchers, this is also known as “mad cow disease.”

To address this, OpenAI has established a dedicated team to explore how future AI models can be improved despite limited training data. Whether and how this will succeed remains the big question.

Source title image: Adobe Stock / 沈军 贡

Read more

Share this article:

Also available in

More Articles

18.07.2026

How to Stifle Open Source Without Banning It

Benedikt Langer

6 Min. Read The sharpest argument against China’s top open AI comes from a man at OpenAI. Dean Ball, ...

Read Article
17.07.2026

The data claim already applies to existing fleets

Benedikt Langer

9 Min. read time The right to access readily available product data has been in effect since September ...

Read Article
17.07.2026

NIS2 liability for management boards applies despite registration

Tobias Massow

9 Min. read time As of late May 2026, around 11,000 affected companies in Germany still lack BSI registration-and ...

Read Article
15.07.2026

Token-OPEX: Inference Controls, Not the Seat Budget

Angelika Beierlein

9 Min. read time Token costs aren’t a line item in SaaS contracts. They’re variable OPEX per workflow-and ...

Read Article
15.07.2026

Hardware Outperforms Software Deals – Rethinking Capital Expenditure Priorities

Benedikt Langer

9 Min. read time IBM reports a 7% decline in infrastructure for Q2, while distributed infrastructure ...

Read Article
14.07.2026

The Bill for Ten Years of Island Solutions

Benedikt Langer

8 min read For a decade, industry has bought point solutions: one system per machine, one standard per ...

Read Article
A magazine by Evernine Media GmbH