How to Stifle Open Source Without Banning It
Benedikt Langer
6 Min. Read The sharpest argument against China’s top open AI comes from a man at OpenAI. Dean Ball, ...
Some experts predicted it two years ago, and now it seems to be happening: AI is running out of high-quality training data! Major providers are therefore increasingly looking for alternatives.
Wikipedia, as a vast online encyclopedia, has an equally large global following. Many of its users actively contribute to expanding the knowledge base of over 60 million entries (as of early 2024).
AI, unfortunately, cannot rely on such a human following to increase the data underlying its knowledge. Training AI requires a constant supply of new, high-quality data. As t3n reports, developers of AI systems have often turned to freely available magazines and academic publications on the internet in the past.This source is gradually drying up. Major providers have therefore already signed contracts with publishing houses such as Springer, Reuters, or The New York Times. Scientific archives and communities like Reddit and Stack Overflow are also being used to nourish and expand AI knowledge. However, according to t3n, the problem is that these sources grow too slowly to satisfy the training appetite of increasingly sophisticated AI models.As early as 2022, there were warnings that the knowledge sources for AI training would run dry by 2026 at the latest. Other experts had assumed this would happen two years later.Some developers, such as Alphabet, the parent company of Facebook, simply resort to less reliable sources for their training purposes.
Another approach, adopted by companies like Anthropic since the Opus version of the Claude model series, is to use so-called synthetic data for training. OpenAI, the creator of ChatGPT, is also reportedly using this trick for its new language model, Orion. Synthetic data are not human-created data that mimic real data.
They are generated through computational algorithms and simulations based on generative AI technologies. Instead of expensive conventional data acquisition, this allows data sets of any size to be generated quickly and easily and adapted to specific requirements.
The method of relying on synthetic or low-quality data from social media posts for AI training is not without controversy among researchers. This is because it can negatively impact the quality of the output.
With synthetic data, it remains unclear how AI should continue to learn if it only has self-generated data available for training. Moreover, there is a risk that AI models may start to limit themselves if they imitate self-generated training data.
Synthetic data can even render AI unusable, as an experiment by Stanford University demonstrates.
The resulting training can lead to both errors and, in the best-case scenario, artifacts in AI responses, ultimately resulting in completely unusable answers. Among researchers, this is also known as “mad cow disease.”
To address this, OpenAI has established a dedicated team to explore how future AI models can be improved despite limited training data. Whether and how this will succeed remains the big question.
Source title image: Adobe Stock / 沈军 贡