Monday, August 10, 2026

"China faces new AI bottleneck as it runs out of Chinese-language training data"

As Grandmother used to say, if it's not one tham ding it's another.

From The South China Morning Post, August 8:

The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years 

China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data.

While the US chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to be the next major bottleneck to the nation’s tech ambitions – and one that hardware workarounds cannot easily solve.

It is a challenge confronting AI giants on both sides of the Pacific – and some US companies are already resorting to aggressive measures to stay ahead.

The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years, according to US-based research institute Epoch AI.

OpenAI co-founder Andrej Karpathy has also warned of a looming “data wall” by the end of this decade, beyond which model capabilities could hit a plateau unless they were fed fresh, reliable information.

Top American labs are spending lavishly to mine offline human knowledge, igniting a fierce ethical debate in the process.

Under an internal initiative code-named Project Panama, Amazon.com-backed Anthropic spent tens of millions of dollars acquiring millions of physical books, severing their bindings and scanning every page into digital form before discarding the originals, according to court filings reported by The Washington Post in January.

The disclosures drew fierce condemnation from authors, archivists and preservationists who accused tech firms of treating human heritage as disposable raw material. “None of our data-acquisition programmes buy and destroy ‘rare’ or ‘antiquarian’ books,” an Anthropic representative told fact-checking site Snopes earlier this month.

Still, the controversy highlighted how desperate frontier AI developers have become to secure vast volumes of human-written text as online web data runs dry.

For China, however, the impending data wall poses a unique threat.

While English accounts for nearly half of all content on the global web as of this month, Chinese represents a mere 1.3 per cent, according to internet tracker W3Techs, placing it far behind languages like Spanish at 6 per cent, German at 5.9 per cent and Japanese at 5 per cent.

In response, Beijing is moving aggressively to treat data as a core strategic asset. In June, the National Data Administration unveiled a sweeping nationwide plan to boost the supply, circulation and commercialisation of high-quality AI training data....

....MUCH MORE 

Also at the SCMP:

Don’t you dare come between Chinese women and our virtual boyfriends 

Probably not helpful re: the birth dearth but, as always, the heart wants what the heart wants.

If interested in more on the book destroyers we have on offer July 29's "AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material"