China’s new AI bottleneck isn’t chips. It’s running out of Chinese-language training data.
TL;DR
China faces a critical AI data shortage. Chinese is just 1.3% of web content vs 49% English. Beijing plans national datasets by 2028. WeChat and Douyin don’t share data externally. Publishers are adding AI training bans.
China’s AI race has a new constraint, and it is not chips. The country is running out of high-quality Chinese-language training data. While US export controls on advanced semiconductors have dominated the debate over China’s AI capabilities, Chinese experts are increasingly warning that data scarcity could prove equally limiting, and unlike chips, there is no hardware workaround. Chinese accounts for just 1.3% of global web content, according to internet tracker W3Techs, compared to nearly half for English, 6% for Spanish, and 5% for Japanese.
The problem is global but hits China harder. Epoch AI estimates the worldwide supply of high-quality, publicly available text could be fully exhausted within six years. OpenAI co-founder Andrej...
Copyright of this story solely belongs to thenextweb.com. To see the full text click HERE