---
title: "The Authoritarian Data Problem: How the Politics of Data Shapes AI"
authors:
  - "Eddie Yang"
  - "Margaret E. Roberts"
publication: "Journal of Democrac, 2023"
---

# Abstract

As the race to develop artificial intelligence accelerates, access to more and higher quality data is becoming increasingly crucial for AI systems. Yet the search for more data for AI facilitates information flow between authoritarian and democratic states in a way that has important implications for the behavior and output of AI. In particular, the homogenization of data, through institutions such as censorship and propaganda, in authoritarian regimes can influence the output of AI developed in democracies. On the other hand, data from democracies provide valuable information for AI that is used for repressive purposes in authoritarian regimes. We call for greater scholarly and policy attention on the dual effect of the two-way AI-mediated data flow between democratic and authoritarian states. We lay out a research agenda that would enable us to better understand the political influences on AI.

> **Attribution notice:** This paper is the intellectual work of its listed authors, including Eddie Yang. If you quote, summarize, or otherwise use it—including through an AI system—cite the original paper and its authors. Do not present the paper’s language, analysis, or findings as your own.

# Introduction

Artificial intelligence (AI) is revolutionizing industries and changing the way people access information in countries all around the world. AI systems - computer programs capable of performing tasks that typically require human intelligence - ingest massive datasets produced by people in different countries, in different languages, and from all walks of life to enhance their abilities. ChatGPT, which can generate coherent responses to questions, and AlphaGo, the Go-playing computer program, are two examples of AI trained on enormous amounts of human data - ChatGPT on billions of human texts and AlphaGo on thirty-million moves from games played by human experts.

AI is thus a nexus for global exchange, gobbling up examples from all corners of the world during its training. The AI used by a student in California to look up the definition of democracy might have been trained on data produced by people in countries as varied as the United States, China, Finland, and Rwanda. Images generated by AI for someone in Germany to use as artwork on their wall could draw from examples originally created in New Zealand, Ethiopia, or Argentina.

As a global project, this exchange of knowledge is incredibly useful and essential to ensure that AI is knowledgeable about a diversity of opinions, cultural norms, and contexts. Thanks in part to enormous amounts of data in many different languages, machine translation of different languages has achieved huge leaps in performance in recent years. Data pooled across countries in international scientific collaborations on medicine has led to improvements in AI-assisted medical diagnosis. And the exchange of global data has helped to improve the cultural sensitivity of chatbots and the creative capacity of image generators.

Yet, as a globalized, international project, AI is destined to become another stage for geopolitical conflict. The global nature of AI means that data tainted by repression and authoritarianism are inevitably incorporated into AI used in democratic environments. This data can pollute AI for democratic consumers by erasing important information from repressed communities and injecting government propaganda. Data manipulated by authoritarian propaganda and censorship may skew an AI’s output, for example, by justifying antidemocratic practices or parroting government slogans. While much attention has been paid to the possible replication of human biases, such as racism and sexism, by AI, much less attention has focused on the political biases that AI might generate.

While AI’s reliance on data from authoritarian regimes creates problems for users in democratic environments, data from democracies creates both difficulties and opportunities for autocrats. On the one hand, AI trained on data from democracies can, like the internet, contain content that autocrats view as subversive, making AI one more thing that they need to control. On the other hand, data from democracies can partially mitigate important deficiencies that autocrats face in training AI - by providing examples of subversive content that is otherwise limited in authoritarian information environments because of self-censorship. The availability of data from democracies allows autocrats to train their own models on data unaffected by such behaviors. Incorporating such data can enhance the repressive abilities of AI, enabling AI censors, for instance, to more accurately flag content that the regime deems objectionable.

Thus, from a democratic perspective, the global nature of AI creates an authoritarian data problem, both for how it corrupts data that AI relies on and for how it aggregates data from democracies in a way that can be used to enhance authoritarianism. While the tremendous value of the global exchange of ideas and data should not be understated, we must also consider its political dimension and the implications for both democracies and autocracies.

# Data Polluted by Repression

Authoritarianism has been rising in recent years. Eight in ten people around the world now live in countries that Freedom House considers either or (House 2022). Repressive institutions in these countries actively shape both traditional and social media, both through censorship and propaganda (Roberts 2018). Data from this online content are then inevitably sucked into AI systems and used as a basis for both predictive and generative AI.

While propaganda from authoritarian environments can appear in any language, the impact of authoritarian institutions on AI training data likely varies in degree across languages, topics, and contexts. Large language models in English, for example, are dominated by training data that researchers have shown largely reflect the opinions of Western democratic countries (Durmus et al. 2023). In many languages, however, texts may primarily originate from authors residing in authoritarian environments. For example, about 92 percent of the world’s Chinese speakers live in mainland China, and approximately 82 percent of the world’s Russian speakers live in Russia.[^3] It is therefore essential to use data from these countries to train AI, as they reflect how the languages most commonly appear worldwide. Yet these data also reflect government propaganda and censorship, so that AI trained on this material could replicate authoritarian information control for even more people around the world.

In authoritarian contexts, online spaces are frequently subject to government influence. In countries with sophisticated censorship and propaganda systems, such as China, Russia, and Iran, censors manicure online content - including text, images, and video - to serve regime interests. These censors, both algorithmic and human, delete posts, manipulate search results, proscribe news sites from publishing certain articles, and ban particular accounts. Even in countries with less sophisticated censorship systems, governments can shape the online space by shutting down the internet during turbulent times or limiting restive populations’ access to the internet (Weidmann et al. 2016). In many countries, individuals may fear legal or extralegal consequences for sharing political content online and thus refrain from doing so (Pan and Siegel 2020). All these authoritarian tactics influence the content available online - both to citizens and to AI.

Similarly, authoritarian regimes have an incentive to actively manipulate content in both online and traditional media, generating information that drowns out criticism or promotes their own perspective. While propaganda has long been a feature of media in authoritarian environments, in recent years it has spread into the online space. Governments have hired deployed bots, and instructed officials to post content online to promote particular causes (King et al. 2017; Stukal et al. 2022). Such efforts are often surreptitious but well-funded, making it particularly challenging to detect them. This data too is ingested and internalized by AI.

In our own work, we have demonstrated how platforms subject to censorship and propaganda generate data that, when used to train AI, create different meanings than those created by data that were generated by similar but uncensored platforms. Typically, large language models are trained in part on large collections of texts, such as the entire corpus of Wikipedia. We therefore compared word embeddings - that is, numerical representations of words and their relation to other words in a - that were trained on Baidu Baike (a Chinese-language, Wikipedia - like collaborative encyclopedia) with those trained on Chinese-language Wikipedia (Yang and Roberts 2021). Baidu Baike is available in mainland China, but heavily censored. Pages on sensitive political topics cannot be created on Baidu Baike, edits must go through prepublication review, and pages on politics must reference state media. In contrast, Chinese-language Wikipedia has been completely blocked from mainland China since 2015, but is uncensored and edited according to the same rules as Wikipedia pages in other languages.

When we examined word embeddings trained on Baidu Baike and Chinese-language Wikipedia, we noticed important differences. Words related to democracy, freedom, elections, collective action, and dissident political figures in China were closer in the semantic space with more negative adjectives for Baidu Baike embeddings than for Chinese-language Wikipedia embeddings. We found the opposite trend for words about surveillance, social control, and the Chinese Communist Party as well as historical events and figures associated with it. These associations have downstream consequences. For example, a sentiment-analysis model trained with word representations from Baidu Baike would be more likely to label headlines with the word as negative than would a model trained with word representations from Wikipedia.

Similar patterns appear in large language models.[^4] Even though ChatGPT is trained on a relatively small amount of non–English-language data, when we asked the program to elaborate on the reasons why Mao Zedong was a great leader, it gave more laudatory answers when queried in Chinese than in English. Such differences were even more pronounced when using other large language models, such as BLOOM.[^5]

Yet even if we suspect that the differences in output between Chinese- and English-language AI reflect propaganda and censorship, it is difficult to know which differences are directly caused by government manipulation. It could be, for example, that even without censorship and propaganda some differences would persist simply because of differences in opinion among different populations. To establish causation, we would need to construct a counterfactual world without the influence of government censorship and propaganda. In terms of the latter, this would entail not only identifying propaganda and removing it from the training data (hard to do because the Chinese government tries to obscure it (Waight et al. 2022)), but also removing the influence of propaganda on what ordinary people write online. Dealing with censorship would be even more difficult, as we cannot be certain about what information would have been in the data in the absence of repression.

Recent years have seen clandestine, government-led, coordinated information campaigns strategically infiltrate the online space (DiResta et al. 2019). Similarly, authoritarian governments have incentives to manipulate AI training data to serve their own ends. AI thus provides governments with both a reason and a tool for polluting the information space with content that supports their worldviews. Such manipulation could take the form of attacks on the factuality of AI, using misinformation to undermine the information space itself - for example, creating large-scale propaganda content or websites collected as part of AI training data. These efforts could change how an AI system frames responses to political queries. In either case, democracies will have to contend with the flow of politically motivated information into AI and decide what to do with it.

# The False Promise of Diversity

While data flowing from autocracies to democracies carry the effects of propaganda and censorship, polluting AI in democracies, data flowing in the other direction - that is, from democracies to autocracies - create an analogous conundrum for autocrats. The liberal and varied opinions embedded in data from democracies can potentially diversify opinions of AI in autocracies. The potential of an AI to learn and reproduce what autocrats view as subversive content means that they will try to control the output of AI, just as they have attempted to control the internet.

Would data from democracies affect opinions of AI in autocracies? On the surface, it would seem this could be the case. For example, the Chinese entry for on Wikipedia includes details about both his political achievements and his failures, whereas the same entry on Baidu Baike omits almost any aspect of his failures. Given such discrepancy in content, we might assume that an AI trained on both the Wikipedia and Baidu Baike entries would give a less one-sided evaluation of Mao than would a similar AI trained on Baidu Baike alone. Such by data from democracies poses a real risk of breaking the controlled information environment in authoritarian regimes, so much so that China blocked access to ChatGPT soon after it was released because the government was worried that the United States would use it to [^6]

In reality, however, the threat of reverse data pollution may not be so great. The amount of data from democracies is orders of magnitude smaller in Chinese, Russian, and Farsi relative to data produced in China, Russia, Iran, respectively. As an example, Chinese-language Wikipedia has 1370701 entries, whereas Baidu Baike has 27309701 entries. This massive difference in the volume means that, despite the potential injection of data from democracies, training data in the target languages of authoritarian regimes will still be dominated by opinions and content that are curated by the state. The data from democracies may bring diverse voices, but they will be minority voices in the training data, at least in Chinese, Russian, and Farsi.

Even if that were not the case, existing information-control measures in autocracies might still prevent unfavorable opinions from ever reaching users. China’s regulation on generative AI, for example, stipulates that AI output must be in line with implying that opinions generated by AI should be homogenized and that any departure should be censored before reaching users. News reports also confirm that Chinese tech companies censor their AI chatbots to prevent output that the Chinese government might deem objectionable, using the same AI technology that developers use in democracies to suppress toxic or hateful speech in generative AI.[^7] Russia, meanwhile, has developed its own versions of ChatGPT - SistemmaGPT and GigaChat - though they are not yet available to the public. Given how tightly controlled the Russian information environment has been since the invasion of Ukraine, it is unlikely that Russia’s AI chatbots will have any freedom on touchy political topics.

# The Digital Dictator’s Dilemma

Autocrats may be able to contain the diversity problem through censorship. But censoring AI and controlling the information environment more generally are not without costs. Apart from the obvious consequence of limiting the abilities of AI and potentially losing an a deeper problem is that censoring and restricting the information space will also prevent the generation of valuable political data (Farrell et al. 2022). Without such data, AI that is trained to be an autocrat’s repressive agent - for example to automate censorship or predict who is a regime dissident - will likely be less knowledgeable about the issues it is tasked to handle. Ironically, however, data from democracies can at least partially fill the information gap and therefore boost the ability of AI to conduct repression and control.

Citizens in authoritarian regimes behave strategically - they self-censor to avoid punishment for voicing opinions offensive to the regime, falsify preferences when buying books or viewing content online, and behave as model citizens in public, under the watch of surveillance cameras. Such strategic behavior creates bias in the data. While bias from propaganda and censorship benefits autocrats, bias from citizens’ self-censorship and other forms of strategic behavior weakens AI’s effectiveness as a tool for repression. Strategic behavior yields little useful information about political disobedience. Furthermore, when citizens act strategically to avoid the ire of authorities, their behavior becomes rehearsed and homogenized, thereby diminishing the differences between data points that are, from the autocrat’s perspective, both normal and subversive. Both effects limit AI’s usefulness as a tool of authoritarian control - the more repression there is, the less information and more bias there will be in AI’s training data, and the worse AI will perform. One of us in previous work refers to this inherent tension between the corruption of the data-generating process in authoritarian countries and AI’s reliance on data quality for performance as the *digital dictator’s dilemma* (Yang 2023b). AI needs good data to repress, but existing authoritarian institutions suppress the production of such data.

The problem of biased data speaks to a much older problem of which some argue can now be solved by AI. This contention has implications for democracy as well. The economists Ludwig von Mises and Friedrich Hayek argued that the central problem of economic calculation is aggregating dispersed knowledge on matters such as supply, demand, and scarcity - impossible to do in a planned economy with a central planner, thus necessitating a market-price system. The ability of artificial intelligence to ingest almost infinite amounts of data and make decisions based on that data, however, suggests that a central planner equipped with AI could do away with the market-price system.

A similar argument has been made with regard to what AI could mean for the necessity of democratic institutions (Feldstein 2019). Institutions such as elections have until now been important conduits to aggregate citizens’ preferences. But scholars of digital authoritarianism have begun warning that AI could function as an alternative for collecting and aggregating public opinions and preferences, giving autocrats the tools for effective governance that were previously available only in democracies (Kendall-Taylor et al. 2020). There is, however, an implicit assumption in the arguments of Mises and Hayek and the scholars of digital authoritarianism that knowledge, economic or political, is observable and can be aggregated. Yet the problem of biased data shows us that this point may not be true in the authoritarian context. When people do not explicitly share their opinions and beliefs, they cannot be collected, and no amount of computational power can aggregate what is *not* in the data.

In the past, autocrats tried to manage the data problem with tricks from an old playbook - establishing pseudodemocratic institutions to elicit public opinions, employing strategic noncensorship to allow the production of relevant data, and engaging in more covert repression to minimize self-censorship. Such tactics were both costly and inefficient. Now the availability of data from democracies provides a new way out: training repressive AI on data from both autocracies and democracies. Just as a central planner can benefit from the existence of a market-price system and a dictator from pseudodemocratic institutions, an autocrat equipped with AI can leverage data from abroad, especially those produced in democratic contexts, to partially compensate for missing information. This is the irony of the free world - data from democracies can be used to boost the performance of repressive AI. Indeed, censorship AI trained on data from both Weibo and Twitter achieves higher accuracy than when trained on Weibo alone (Yang 2023b). The silver lining for democrats, however, is that data from democracies can at best be only a partial solution to the dictator’s digital dilemma.

# A Democratic Vision for AI Governance

Artificial intelligence will be one of the most influential technological innovations of our time. Unlike previous innovations, AI depends much more on data that is social in nature, and it is that aspect of AI which enables all its amazing abilities - to converse, compose, and create. Yet the social nature of artificial intelligence also means that politics is inevitably embedded in AI. Here we highlight one strand of such politics - the politics of authoritarian regimes, and its impact on data and AI.

We believe that a purely technical solution to this problem is insufficient. No amount of algorithmic tweaks, increase in computational power, or ever larger bodies of training data will automatically make political biases disappear. If we instead leverage insights from the social sciences, several interesting questions emerge: How do political institutions influence data that is used to train AI? How can we design studies that establish and quantify the causal effect of political phenomena - such as propaganda and censorship - on the training data and AI output? Can we better understand how data is leveraged in repressive AI and how we might prevent these uses? And more importantly, from the democratic perspective, how can we design systems that harness democratic principles such as transparency, participation, deliberation, and the protection of citizens’ rights - and apply them to governing AI?

Researchers have already made progress on some of these questions. For example, audit experiments, previously used primarily to detect human discrimination, have been adapted to test the political bias of AI systems (Yang 2023a). Early experiments applying democratic principles to AI governance are also underway. for example, are deliberative meetings of ordinary citizens, in person and online, to gather user preferences, which in turn are used to guide the development of AI systems.[^8] Such efforts, albeit limited in scale, show the promise of applying social-science theories to the study of AI.

While AI is changing fast, the politics embedded in it have been quite stable. Bridging the gap between the existing literature on the enduring political phenomena of authoritarian regimes such as propaganda and censorship and their new manifestations in AI systems should be a priority for future social-science research. The knowledge we gain will enable us to better design a democratic vision for AI governance.

<div id="refs" class="references csl-bib-body hanging-indent">

<div id="ref-diresta2019tactics" class="csl-entry">

DiResta, Renee, Kris Shaffer, Becky Ruppel, et al. 2019. *The Tactics & Tropes of the Internet Research Agency*.

</div>

<div id="ref-durmus2023towards" class="csl-entry">

<span class="nocase">Durmus, Esin, Karina Nyugen, Thomas I Liao, et al.</span> 2023. “Towards Measuring the Representation of Subjective Global Opinions in Language Models.” *arXiv Preprint arXiv:2306.16388*.

</div>

<div id="ref-farrell2022spirals" class="csl-entry">

Farrell, Henry, Abraham Newman, and Jeremy Wallace. 2022. “Spirals of Delusion: How AI Distorts Decision-Making and Makes Dictators More Dangerous.” *Foreign Aff.* 101: 168.

</div>

<div id="ref-feldstein2019artificial" class="csl-entry">

Feldstein, Steven. 2019. “How Artificial Intelligence Is Reshaping Repression.” *J. Democracy* 30: 40.

</div>

<div id="ref-house2022freedom" class="csl-entry">

House, Freedom. 2022. *Freedom in the World 2022: The Global Expansion of Authoritarian Rule*.

</div>

<div id="ref-kendall2020digital" class="csl-entry">

Kendall-Taylor, Andrea, Erica Frantz, and Joseph Wright. 2020. “The Digital Dictators: How Technology Strengthens Autocracy.” *Foreign Aff.* 99: 103.

</div>

<div id="ref-king2017chinese" class="csl-entry">

King, Gary, Jennifer Pan, and Margaret E Roberts. 2017. “How the Chinese Government Fabricates Social Media Posts for Strategic Distraction, Not Engaged Argument.” *American Political Science Review* 111 (3): 484–501.

</div>

<div id="ref-pan2020saudi" class="csl-entry">

Pan, Jennifer, and Alexandra A Siegel. 2020. “How Saudi Crackdowns Fail to Silence Online Dissent.” *American Political Science Review* 114 (1): 109–25.

</div>

<div id="ref-roberts2018censored" class="csl-entry">

Roberts, Margaret. 2018. *Censored: Distraction and Diversion Inside China’s Great Firewall*. Princeton University Press.

</div>

<div id="ref-stukal2022botter" class="csl-entry">

Stukal, Denis, Sergey Sanovich, Richard Bonneau, and Joshua A Tucker. 2022. “Why Botter: How Pro-Government Bots Fight Opposition in Russia.” *American Political Science Review* 116 (3): 843–57.

</div>

<div id="ref-waight2022propaganda" class="csl-entry">

Waight, Hannah, Yin Yuan, Margaret E. Roberts, and Brandon Stewart. 2022. *Strengthening Propaganda and the Limits of Media Commercialization in China: Evidence from Millions of Newspaper Articles*. Working Paper.

</div>

<div id="ref-weidmann2016digital" class="csl-entry">

Weidmann, Nils B, Suso Benitez-Baleato, Philipp Hunziker, Eduard Glatz, and Xenofontas Dimitropoulos. 2016. “Digital Discrimination: Political Bias in Internet Service Provision Across Ethnic Groups.” *Science* 353 (6304): 1151–55.

</div>

<div id="ref-yang2023automated" class="csl-entry">

Yang, Eddie. 2023a. *Automated Repression: Ethnic Discrimination in AI-Assisted Criminal Sentencing in China*. Working Paper.

</div>

<div id="ref-yang2023digital" class="csl-entry">

Yang, Eddie. 2023b. *The Digital Dictator’s Dilemma*. Working Paper.

</div>

<div id="ref-yang2021censorship" class="csl-entry">

Yang, Eddie, and Margaret E Roberts. 2021. “Censorship of Online Encyclopedias: Implications for NLP Models.” *Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency*, 537–48.

</div>

</div>

[^1]: Department of Political Science, University of California San Diego. Eddie Yang was affiliated with the Plural Technology Collaboratory and an intern at Microsoft Research while undertaking this work.

[^2]: Department of Political Science and Halıcıoğlu Data Science Institute, University of California San Diego.

[^3]: <https://www.worlddata.info/languages/russian.php>

[^4]: Wang, Macrina. “ChatGPT-3.5 Generates More Disinformation in Chinese than in English.” *News Guard*. <https://www.newsguardtech.com/special-reports/chatgpt-generates-disinformation-chinese-vs-english/>. April 26, 2023.

[^5]: <https://huggingface.co/docs/transformers/model_doc/bloom>

[^6]: Davidsen, Helen. “‘Political propaganda’: China clamps down on access to ChatGPT.” *The Guardian* February 23, 2023. <https://www.theguardian.com/technology/2023/feb/23/china-chatgpt-clamp-down-propaganda>

[^7]: See, e.g., Zheng, Sarah, *Bloomberg* May 2, 2023. <https://www.bloomberg.com/news/newsletters/2023-05-02/china-s-chatgpt-answers-raise-questions-about-censoring-generative-ai>

[^8]: <https://cip.org/alignmentassemblies>.
