
China’s artificial intelligence ambitions are confronting an unexpected bottleneck. While export controls on advanced semiconductors have long dominated discussions about the country’s AI capabilities, a growing number of experts now argue that a shortage of high-quality Chinese-language training data may prove just as limiting. Unlike hardware, this constraint cannot be solved with alternative supply chains or engineering ingenuity. Chinese is spoken by over a billion people, yet it accounts for only 1.3% of global web content, according to web technology tracker W3Techs. By contrast, English makes up nearly half of all online content, with Spanish at 6% and Japanese at 5%.
This disparity creates a fundamental problem for Chinese AI developers. Large language models rely on vast amounts of text to learn grammar, reasoning, and factual knowledge. When native-language data is scarce, models must either draw on lower-quality sources or depend on translated material, which can introduce cultural bias and factual errors. The situation is particularly acute in specialized domains like medicine, law, and engineering, where precise terminology and contextual understanding are essential.
A Global Problem, Amplified in China
Data scarcity is not unique to China. Epoch AI, a research organization that tracks AI trends, estimates that the worldwide supply of high-quality, publicly available text could be fully exhausted within six years. OpenAI co-founder Andrej Karpathy has warned of a “data wall” by the end of this decade, a point at which the growth of AI model performance will slow due to a lack of fresh training material. However, the problem hits China disproportionately hard.
Because Chinese is underrepresented online, developers in China often pay more per useful token—the basic units of text that models process—than their Western counterparts. A token in Chinese may carry less unique information or appear in lower-quality contexts, forcing models to work harder to achieve the same level of accuracy. This cost disadvantage compounds over time, especially as models become larger and require exponentially more data.
The digital ecosystem in China exacerbates the shortage. Super-apps like WeChat and Douyin, which generate enormous amounts of user-generated content, do not share their data with third-party developers. This walled-garden approach protects user privacy and corporate interests, but it leaves AI labs to train on publicly accessible sources that are often thinner and less representative of everyday Chinese life. In contrast, many Western tech companies have already established extensive data-sharing agreements or have built proprietary datasets through years of search engine and social media operations.
Beijing’s Strategic Response
The Chinese government has recognized the severity of the situation and is treating data as strategic infrastructure. In June, the National Data Administration unveiled a nationwide plan to construct validated AI training datasets by 2028. The initiative covers key sectors such as manufacturing, energy, healthcare, finance, agriculture, autonomous driving, and embodied AI. These datasets are intended to be authoritative, high-quality, and annotated according to national standards, giving Chinese AI developers a reliable foundation.
Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, stressed the importance of this approach. “Competition in the AI era is not only about models and computing power, but also about high-quality data supply systems,” he said. The statement reflects a broader shift in Beijing’s AI strategy—from hardware procurement to data sovereignty. By building its own data infrastructure, China hopes to reduce dependence on foreign sources and shield its AI industry from external disruptions.
Tsinghua University computer scientist Sun Maosong has urged authorities to go further. He has called for the digitization of historical archives, ancient manuscripts, scientific literature, and regional dialects. These repositories of knowledge are largely untapped and could provide a unique advantage for Chinese AI models, especially in cultural and historical contexts where Western models are weak. Sun’s proposal aligns with the government’s broader effort to leverage cultural heritage for technological development, but it also raises questions about the feasibility of such large-scale digitization projects.
Publishers Push Back
While Beijing is eager to expand the pool of Chinese-language data, not everyone is willing to contribute. Huaxia Publishing House recently added an explicit warning to one of its new translations: “It is prohibited to use the content of this book for artificial intelligence training. Violators will be held legally responsible.” This marks a growing trend among Chinese publishers, who worry that their intellectual property is being exploited without compensation.
The publishing industry’s resistance mirrors a global movement. In the United States, several authors and media companies have filed lawsuits against AI developers over the unauthorized use of copyrighted material. For Chinese publishers, the concern is particularly acute because they see AI companies as potential competitors. If a model can replicate a book’s content or style, why would readers buy the original? Some publishers are now demanding licensing agreements similar to those negotiated by Western media organizations, while others are simply locking their digital versions behind stricter access controls.
The conflict between data demand and data ownership is likely to intensify. China’s national dataset plan may eventually include provisions for compulsory licensing, but for now, many content owners are choosing to opt out. This creates a paradox: the more China needs high-quality data, the more reluctant data holders become to share it.
Lessons from the United States
The United States faces a similar data challenge, but it has a different set of tools to address it. Anthropic’s Project Panama, for instance, reportedly purchased and destroyed millions of physical books to scan their contents, creating a massive repository of high-quality text that is not publicly indexed. This aggressive approach highlights the lengths to which US companies will go to secure training data. China, with its 1.3% share of web content and a publishing industry that is already putting up barriers, does not currently have comparable projects in the open.
Chinese researchers are aware of this gap. Some have proposed crowdsourcing approaches, where citizens are incentivized to contribute original text or annotate existing data. Others have suggested using synthetic data—text generated by AI models themselves—to augment limited real-world resources. However, synthetic data can lead to model collapse, a phenomenon where an AI trained on its own output becomes progressively less accurate and diverse. This is not a viable long-term solution.
The data shortage also affects the types of AI applications China can pursue. Autonomous driving, for example, requires vast amounts of localized road data, including traffic signs, driving behaviors, and road conditions in Chinese cities. While this data is being collected through sensor-equipped vehicles, it is not always available in a format suitable for training large models. Similarly, healthcare AI depends on clinical records, medical imaging, and case histories, all of which are highly sensitive and often locked inside hospital systems. The government’s push to standardize datasets across industries is an attempt to break down these silos, but progress is slow.
The Global Race for Data
China is not alone in treating data as a strategic resource. The European Union is developing common data spaces for sectors like mobility and health, while Japan has created a Personal Data Trust that permits individuals to share information for AI training. These initiatives recognize that data, like computing power, is a critical input for the AI economy. Countries that fail to build robust data ecosystems risk falling behind.
For China, the immediate challenge is to increase the supply of Chinese-language data without compromising user privacy or intellectual property rights. The national dataset plan is a step in the right direction, but it must be implemented carefully. If the government coerces publishers or individuals into contributing data, it may face public backlash and legal challenges. If it fails to provide adequate incentives, the data drought will continue.
Some experts believe that China’s cultural and linguistic diversity could actually be an asset. The country has dozens of regional dialects, each with its own vocabulary and expressions. By digitizing these dialects, Chinese AI models could achieve a level of nuance that English-language models cannot match. This would give China a unique edge in serving domestic users, who often communicate in Mandarin mixed with regional phrases. Yet this effort requires extensive fieldwork and community cooperation, which cannot be achieved overnight.
Meanwhile, the global AI industry is racing toward the same data wall. If Epoch AI’s forecast is correct, all major AI developers will soon face a scarcity of high-quality text, regardless of language. The countries and companies that have accumulated the largest and most diverse datasets will have a lasting advantage. The United States, with its large English-language web presence and aggressive data acquisition strategies, appears better positioned than China. Chinese developers, however, may find unexpected value in their own cultural resources—if they can unlock them in time.
