07.01.2025
3 min read

Some experts predicted it two years ago, and now it seems to be happening: AI is running out of high-quality training data! Major providers are therefore increasingly looking for alternatives.

Wikipedia, as a vast online encyclopedia, has an equally large global following. Many of its users actively contribute to expanding the knowledge base of over 60 million entries (as of early 2024).

AI, unfortunately, cannot rely on such a human following to increase the data underlying its knowledge. Training AI requires a constant supply of new, high-quality data. As t3n reports, developers of AI systems have often turned to freely available magazines and academic publications on the internet in the past.This source is gradually drying up. Major providers have therefore already signed contracts with publishing houses such as Springer, Reuters, or The New York Times. Scientific archives and communities like Reddit and Stack Overflow are also being used to nourish and expand AI knowledge. However, according to t3n, the problem is that these sources grow too slowly to satisfy the training appetite of increasingly sophisticated AI models.As early as 2022, there were warnings that the knowledge sources for AI training would run dry by 2026 at the latest. Other experts had assumed this would happen two years later.Some developers, such as Alphabet, the parent company of Facebook, simply resort to less reliable sources for their training purposes.

Another approach, adopted by companies like Anthropic since the Opus version of the Claude model series, is to use so-called synthetic data for training. OpenAI, the creator of ChatGPT, is also reportedly using this trick for its new language model, Orion. Synthetic data are not human-created data that mimic real data.

They are generated through computational algorithms and simulations based on generative AI technologies. Instead of expensive conventional data acquisition, this allows data sets of any size to be generated quickly and easily and adapted to specific requirements.

Blessing and curse of synthetic data

The method of relying on synthetic or low-quality data from social media posts for AI training is not without controversy among researchers. This is because it can negatively impact the quality of the output.

With synthetic data, it remains unclear how AI should continue to learn if it only has self-generated data available for training. Moreover, there is a risk that AI models may start to limit themselves if they imitate self-generated training data.

Synthetic data can even render AI unusable, as an experiment by Stanford University demonstrates.

The resulting training can lead to both errors and, in the best-case scenario, artifacts in AI responses, ultimately resulting in completely unusable answers. Among researchers, this is also known as “mad cow disease.”

To address this, OpenAI has established a dedicated team to explore how future AI models can be improved despite limited training data. Whether and how this will succeed remains the big question.

Source title image: Adobe Stock / 沈军 贡

Read more

Share this article:

Also available in

More Articles

09.09.2026

Nvidia Buys Hugging Face for Over 11 Billion Euros

Eva Mickler

4 min read Nvidia is acquiring Hugging Face for around 11.1 billion euros; the contract was signed on ...

Read Article
08.09.2026

SAP Lets Joule Steer Robots Directly, Liability Still Open

Bernhard Liebl

4 min read SAP has documented the first Embodied AI Jam at the Swiss Smart Factory in Biel. Inspection ...

Read Article
15.08.2026

ChatGPT wants to read the Mac

Eva Mickler

6 min read On 13 August 2026, OpenAI described Computer History for the ChatGPT Mac app in its release ...

Read Article
14.08.2026

SpaceX acquires Cursor: EU clauses stay open

Eva Mickler

5 min read The purchase agreement was finalized on 14 August 2026. Any company using the tool now has ...

Read Article
13.08.2026

CRA forces manufacturers to report within 24 hours

Bernhard Liebl

9 min read On 11 September 2026, Article 14 of the Cyber Resilience Act comes into force. From that ...

Read Article
11.08.2026

NVIDIA capital plans and what operators must check now

Bernhard Liebl

7 min read On 10 August 2026, NVIDIA announced it will partner with six capital partners to build financing ...

Read Article
A magazine by Evernine Media GmbH