Get the latest tech news

AI 'gold rush' for chatbot training data could run out of human-written text as early as 2026


A new study released Thursday by research group Epoch AI projects that tech companies will exhaust the supply of publicly available training data for AI language models by roughly the turn of the decade -- sometime between 2026 and 2032.

Comparing it to a “literal gold rush” that depletes finite natural resources, Tamay Besiroglu, an author of the study, said the AI field might face challenges in maintaining its current pace of progress once it drains the reserves of human-generated writing. In the short term, tech companies like ChatGPT-maker OpenAI and Google are racing to secure and sometimes pay for high-quality data sources to train their AI large language models – for instance, by signing deals to tap into the steady flow of sentences coming out of Reddit forums and news media outlets. Epoch is a nonprofit institute hosted by San Francisco-based Rethink Priorities and funded by proponents of effective altruism — a philanthropic movement that has poured money into mitigating AI’s worst-case risks.

Get the Android app

Or read this on r/technology

Read more on:

Photo of Human

Human

Photo of gold rush

gold rush

Photo of written text

written text

Related news:

News photo

World’s first 3D e-skin gives robots human-like touching sense | This electronic skin from China can decode pressure, friction, and strain in real time

News photo

Autonomous AI Robot Creates a Shock-Absorbing Shape No Human Ever Could

News photo

OpenAI claims that its free GPT-4o model can talk, laugh, sing and see like a human