---
title: "The Limits of AI for Authoritarian Control"
authors:
  - "Eddie Yang"
publication: "American Journal of Political Science, 2026"
---

# Abstract

An emerging literature suggests that Artificial Intelligence (AI) can greatly enhance autocrats' repressive capabilities. This paper argues that while AI presents a powerful new tool for authoritarian control, its effectiveness is constrained by the very repressive institutions it is designed to serve. This constraint stems from what I term the "authoritarian data problem": citizens' strategic behavior under repression diminishes useful information in the data available for training AI. The more repression there is, the less amount of useful information exists in AI's training data, and the worse AI performs. I illustrate this argument using an AI experiment and censorship data in China. I show that AI's accuracy in censorship decreases with increasing repression, especially during times of political crisis. I further show that this problem cannot be easily fixed with more data. Ironically, international data -- especially data from less repressive settings -- can help boost AI's ability to censor.

> **Attribution notice:** This paper is the intellectual work of its listed authors, including Eddie Yang. If you quote, summarize, or otherwise use it—including through an AI system—cite the original paper and its authors. Do not present the paper’s language, analysis, or findings as your own.

# Introduction

Digital technologies, particularly Artificial Intelligence (AI), have been argued as a powerful addition to the autocrat’s repressive toolkit (Diamond 2019). Facial recognition technologies used in digital surveillance systems enable autocrats to more selectively target opponents of the state (Xu 2021) while deep learning models for natural language processing may enhance information control through automated censorship (Roberts 2020; Gohdes 2024) and (mis)information campaigns (Kreps et al. 2022). The repressive potential of AI further benefits from the massive amounts of data collected by existing authoritarian institutions, which can be leveraged for AI model training and development (Beraja, Yang, et al. 2023; Beraja, Kao, et al. 2023).

Yet we know little about the limits of AI for authoritarian control, if there are any. This paper provides both a theory and empirical evidence that authoritarian institutions limit AI’s repressive capabilities, making it less omnipotent than scholars have previously argued. A key driver for these limits is the inherent tension between AI and authoritarian control: to effectively enforce control and repression, AI requires enough politically relevant information in its training data but institutions of control and repression by nature restrict both the quantity and quality of such information.[^2]

AI relies on data to acquire its predictive capabilities. To effectively enforce authoritarian control, AI needs large quantities of politically relevant information in the training data. For instance, an automated censorship system will only be accurate at filtering censorable content if it has seen many examples of such content during its training. Under the shadow of repression, however, citizens’ strategic behavior can tarnish AI’s training data. When people self-censor (Shen and Truex 2021) and falsify preferences (Kuran 1997) to hide their critical views of the regime, the amount of censorable content is reduced. Other forms of strategic behavior (e.g., using coded language to circumvent censorship or wearing masks to evade surveillance) can also degrade the collected data.[^3] Therefore, the higher the cost of dissent because of repression, the more people will manipulate their public behavior, corrupting the training data and causing AI to be less effective and accurate at carrying out authoritarian control.

The authoritarian data problem can have a more negative effect on the performance of AI during political crises, such as protests and revolutions. This is because the sudden change in people’s behavior during crises, especially behavior that was previously suppressed for fear of repression, causes the data that AI encounters to be very different from its training data. For example, without additional training, AI would have never recognized that people holding a blank piece of paper on the street was a protest against China’s COVID-19 restrictions. The shift in data distribution in crises is particularly bad for autocrats: they need AI to perform the best during times of political turmoil but it is exactly in such times that AI fumbles in performance.

While autocrats can collect more data to further train their repressive AI systems, simply relying on increasing the *quantity* of data is unlikely to solve the AI performance problem – as long as people maintain strategic behavior, the additional data will suffer from the same *quality* issues that cause AI’s underperformance. However, data from democracies – generated largely without the same political constraints as in the authoritarian context – can potentially lessen the authoritarian data problem.

To empirically test the theory, I focus on the use of AI for censorship as a case study. Specifically, I use a novel experiment to recreate commercial censorship AI systems and test their performance given different political conditions. I use millions of social media posts from the Chinese social media platform, Weibo, and Chinese tweets from Twitter as the training data.[^4] Notably, I exploit a rare opportunity to use an automated censorship service from a technology company in China to obtain the political sensitivity of each post. The political sensitivity scores allow me to model self-censorship and preference falsification by creating missingness in the training data. For example, to model a highly repressive environment where there is a high degree of preference falsification and self-censorship, the training data will have few social media posts with high political sensitivity scores. The set-up also allows me to test AI’s performance during crises as well as the effect of data from international sources (Twitter) on AI’s censorship accuracy.

The experiment establishes three sets of results. First, as the information environment becomes more repressive and there is more preference falsification and self-censorship, the resulting information loss in the data causes a drop in the accuracy of AI in classifying which social media posts should be censored. Second, the drop in AI’s accuracy is substantially larger during crises than normal times, with the majority of the misclassifications being false negatives (censorable posts misclassified as ). Third, doubling the amount of data from Weibo (domestic data source) has a marginal effect on the accuracy of censorship AI, while data from Twitter (international data source) helps increase accuracy, despite being a fraction of the Weibo data in size. The improvement from the Twitter data, however, does not fully close the accuracy gap caused by the authoritarian data problem. Through text analysis of the Weibo and Twitter data, I give suggestive evidence that the insufficiency of the Twitter data is likely due to its differences in discourse from the domestic Weibo data. This is consistent with a recent study that shows Venezuelan activists changed their discourse once they went into exile (Esberg and Siegel 2023).

Taken together, the theory and empirical results highlight that the classic strategic behavior by citizens in authoritarian regimes (Kuran 1997; Wintrobe 2000; Jiang and Yang 2016; Roberts 2018) now manifest in new forms to hamper digital dictators who wish to use AI for authoritarian control. While the literature on technology and autocracy has shown that AI can be useful for autocrats, the paper highlights the (understudied) limits of AI for authoritarian control. It does so by challenging an implicit assumption of the existing literature – that more data means more accurate predictions from AI (Feldstein 2019). In contrast, the paper shows that (distributionally) biased data can lead to less accurate predictions and simply adding more biased data will not solve the problem, even for state-of-the-art language models like those developed by DeepSeek. Ironically, however, the paper points out that biases created by preference falsification and self-censorship may be partially alleviated by the availability of data from contexts outside of the authoritarian regime.

More broadly, the paper contributes to our understanding of autocrats’ strategy for information control and regime survival. As modern autocrats move away from mass repression and rely more on the manipulation of the information environment (Guriev and Treisman 2019, 2020), censorship and information gathering have become essential for authoritarian control. Traditionally, autocrats face a trade-off: in order to gather necessary information for regime survival, autocrats need to relax restrictions on the freedom of expression and the press, but doing so risks generating dissent and allowing citizens to learn about the regime’s corruption or incompetence (Egorov et al. 2009; Egorov and Sonin 2020). While existing studies have focused on how *domestic* mechanisms, such as elections (Cox 2009; Rozenas 2010; Miller 2015) and strategic (non-)censorship (King et al. 2013; Lorentzen 2014; Chen and Xu 2017), allow autocrats to strike a delicate balance in the trade-off, the role of *international* sources of information (e.g., diaspora and independent media) has been less explored.

The paper demonstrates one way autocrats can integrate international sources of information (e.g., Twitter) to boost authoritarian control – by using such information to train more accurate censorship AI, while excluding citizens from accessing the same information. Qualitative evidence suggests that the use of international sources of information is already happening systematically, on a large scale, and across authoritarian regimes. On the other hand, the paper also points out the potential limits of such an approach for autocrats, as the paper joins an emerging literature (Esberg and Siegel 2023) in highlighting the difference in content between domestic and international sources of information.

# Background: AI and Autocracy

In this paper, AI refers to computer programs that are capable of performing tasks that typically require human intelligence. Examples of AI performing tasks include playing chess, recognizing faces in surveillance videos, and classifying whether a social media post should be censored. Essentially, AI can be seen as a technology of prediction (Agrawal et al. 2019): predicting the best next move in chess, the identity of a face, and the political nature of some content.

A key underlying technology that powers AI is deep learning – algorithms that are capable of extracting complex relationships from data. Typically, training deep learning models follows a two-stage process: 1) a pre-training stage where models are trained on diverse datasets to obtain general capabilities and 2) a fine-tuning stage where the pre-trained model is adapted to a specific task using customized datasets. For example, to train a censorship AI, one can first obtain a pre-trained model, which is usually trained on large corpora of general text. The pre-trained model is then fine-tuned on a censorship-specific dataset to improve its ability to carry out censorship. In the paper, I focus on data and training in the second stage as fine-tuning has a large impact on AI’s performance on specific tasks.

In the fine-tuning stage, just like OLS regression, deep learning models take as input some features $`X`$ with their corresponding outcomes or labels $`Y`$ and fit a function $`Y = f(X)`$. Unlike OLS regression, deep learning models generally do not pre-specify the relationship between $`X`$ and $`Y`$ but rather use a data-driven approach to learn the functional form of $`f(\cdot)`$. Additionally, deep learning models are usually much more complex, involving upward of billions of parameters.

A key contributing factor to AI’s recent success is the availability of large amounts of high-quality data. Such data enables deep learning models to extract complex relationships that are essential for complicated tasks such as playing chess and carrying out conversations. For example, chess-playing AI AlphaGo Zero was trained on $`4.9`$ million chess games (<span class="nocase">Silver et al.</span> 2017) and AI chatbots like ChatGPT are trained on trillions of words scraped from books and the internet. These models are transforming the modern way of life. Students now rely on AI chatbots for answers to questions and assignments and people increasingly use self-driving technologies to assist with their driving.

Like other areas of society, political institutions have also incorporated the use of AI in their decision-making process. For example, 11 U.S. states and 178 additional counties in other states are using algorithmic risk assessment tools to assist judges in making bail decisions.[^5] India has used facial recognition systems to verify voter identity in elections. Perhaps even more so than democracies, authoritarian regimes have embraced AI to automate tasks like surveillance (Xu 2023), censorship, meting out criminal sentences in place of judges (Yang 2023), and other repressive tasks (Kendall-Taylor et al. 2020). Yet despite AI’s growing importance in politics, evaluations of its impact are rare in political science, with a few notable exceptions (Xu 2021; Allie 2023; Imai et al. 2023).

# Theory: Why Authoritarian Politics Constrain AI

AI derives its capabilities from the data it is trained on. The ideal situation, therefore, is to have a large quantity of high-quality data, as both scale and quality are critical to achieving robust and accurate AI performance. When the training data is scarce or problematic (biased, low-quality etc.), the output of AI often becomes subpar (Sajith and Kathala 2024). A well-known example of bad data causing issues in AI is the case of racial bias in facial recognition models. Facial recognition systems tend to mis-recognize faces with darker skin tones at a much higher rate (Cook et al. 2019; <span class="nocase">Vangara et al.</span> 2019; Robinson et al. 2020). The disparity in error rate is attributed to racial imbalance in the training samples, with a lack of racial minority, especially African-American, faces in the data (Buolamwini and Gebru 2018; Leslie 2020). Subsequently, much effort has focused on increasing the training data quantity and quality through more racially balanced samples (Kärkkäinen and Joo 2019; Wang et al. 2021).

## Data Deficiencies in Authoritarian Regimes

Given the importance of data in dictating the performance of AI, both practitioners and scholars have argued that authoritarian regimes may have an advantage in developing AI systems, due to their ability to collect large amounts of data on their citizens (Lee 2018; Feldstein 2019). This is evident in areas such as healthcare, where countries like China are leading the global market in medical AI, thanks to its readily available data from public hospitals.[^6]

In politics, however, the same argument can fall apart. In the authoritarian setting, in particular, several mechanisms can significantly diminish the quantity and quality of political data. For example, the absence of free and fair elections, an important information-gathering channel, can deprive the state of systematic data on citizens’ genuine preferences, grievances, and levels of dissatisfaction (Miller 2015). Similarly, widespread political apathy, a subdued or co-opted opposition, or even a genuine lack of discontent can lead to a dearth of critical voices in the political training data. In all cases, citizens may behave truthfully and voice their genuine opinions. Yet either no systematic channel exists to collect such data or the data that is available is largely non-political and non-critical in nature.

Additionally, under the shadow of repression, citizens have an incentive to act strategically: individuals may falsify their true preferences, avoid surveillance cameras, or use coded language online. Bureaucrats, fearing repercussions for delivering bad news or failing to meet targets, may misreport local statistics, painting an overly optimistic picture.[^7] Such strategic responses deliberately contaminate the data, introducing noise, falsehoods, and omissions, thereby undermining data for repressive tasks like surveillance and censorship.

Mechanisms such as the absence of elections and political apathy primarily reduce the quantity of relevant data, while citizens’ strategic behavior under duress can degrade both the quantity and the quality of the information the state can collect. Consequently, both sets of mechanisms can undermine AI performance. Imagine, in healthcare, if patients’ medical histories were poorly recorded or if patients systematically misreported their symptoms during diagnosis, such data would likely cause even the most advanced medical AI screening systems to miss critical diagnostic cues, thereby compromising their accuracy and overall efficacy. Indeed, a slew of problems in authoritarian regimes has been linked to bad or missing data, including inefficient governance (Wallace 2022; Trinh 2023) and the sudden collapse of regimes (Kuran 1991; Lohmann 1994). Yet, less studied is the fact that the use of AI in politics suffers just as much, if not more, from these data deficiencies. Just as facial recognition systems trained predominantly on faces of light-skinned faces struggle to recognize darker-skinned faces, AI trained to automate repression and censorship can be crippled by bad data stemming from both systemic information deficits and citizens’ strategic behavior.

In the context of censorship (the setting for the empirical section), data deficiencies can arise from several of the above-mentioned mechanisms. From the perspective of an AI developer creating censorship tools, the ideal training data would contain a diversity of public opinions. Crucially, this includes opinions deemed censorable by the state (e.g., opinions critical of the government). However, a co-opted opposition, political apathy, or a lack of genuine discontent can lead to muted public discourse on policies and government criticism, thereby reducing the quantity of valuable data.

Additionally, citizens’ strategic behavior when voicing public opinions can further degrade the available data. This strategic behavior can manifest in several forms. First, citizens may falsify their public preferences under the perceived threat of punishment (Kuran 1997). While citizens may harbor grievances against the autocrat, the regime, or specific policies in private, the fear of censorship and repression can lead them to suppress these private preferences and instead publicly (Shih 2008; Wedeen 2015). This suppression reduces the amount of data on the more extreme or critical opinions and creates a mismatch between the distribution of citizens’ private preferences and the public data that the autocrat collects. An implication of this mismatch, as shown by several studies, is that publicly expressed popular support for authoritarian regimes is often higher than the actual level of support (Jiang and Yang 2016; Robinson and Tannenberg 2019; Hale 2022; Nicholson and Huang 2023).

Relatedly, individuals can also compromise data quality through self-censorship (Berinsky 1999; Shen and Truex 2021). Instead of falsifying their preferences, people engaging in self-censorship simply refrain from voicing censorable public opinion at all. By self-censoring, citizens can avoid the psychological cost of preference falsification (Crabtree et al. 2020) as well as potential punishment from the regime. Furthermore, other forms of strategic behavior, such as using coded language in online discussions, can further contribute to the degradation of data quality. In this strategic setting, the severity of the data problem depends on the cost of voicing dissent: the higher the cost, the less such information appears in the data, and the more severe the problem becomes (Tannenberg 2022).

## Automating Autocracy with Bad Data

Data deficiencies in the political realm pose significant challenges to the effective automation of authoritarian control. When AI systems are developed using data where useful signals are reduced, obscured, made more subtle, or even falsified, their performance can be significantly impaired. Specifically, AI suffers two interrelated problems stemming from data deficiencies in the authoritarian context: 1) a reduction in sensitive yet valuable signals within the training data, and 2) an increased difficulty in the prediction task itself. While mechanisms such as the absence of elections primarily contribute to the first problem, strategic behaviors like preference falsification and self-censorship exacerbate both problems. Given that strategic behavior can affect both the quantity and quality aspects of the authoritarian data problem, I focus on strategic behavior in the following discussion. To illustrate these two problems, consider the stylized example of censorship AI in Figure <a href="#fig:toy" data-reference-type="ref" data-reference="fig:toy">1</a>.

In this example, the AI’s training data consists of content that should and should not be censored. Both kinds of content are generated from the unobserved distribution of political sensitivity.[^8] As an example, content with sensitivity scores above $`0.5`$ are treated as censorable and those with scores below $`0.5`$ are treated as . The goal of a censorship AI system is thus to identify censorable content while minimizing false positives (i.e., safe content misclassified as censorable). However, with strategic behaviors like preference falsification and self-censorship, much of content with high political sensitivity (e.g., anti-regime content) is not publicly expressed. As a result, the distribution of observed content becomes right-censored, in that content with high sensitivity (represented by the shaded region in Figure <a href="#fig:toy" data-reference-type="ref" data-reference="fig:toy">1</a>) is missing from the available data. This reduces the amount of data with political sensitivity above $`0.5`$, resulting in a smaller amount of censorable content in the training data (Problem 1).

Furthermore, the absence of this high-sensitivity content (the shaded region) effectively reduces the discernible differences between observed censorable and safe content. In other words, the absence of more overtly critical or extreme expressions means that the remaining censorable and safe content in the observed data become more similar to each other. As the overall content becomes more homogeneous, it becomes more difficult for the AI to accurately distinguish between content that warrants censorship and those that do not (Problem 2). Even for the observed censorable content, strategic adaptations like the use of coded language or adversarial content manipulations (e.g., using homophone substitution to swap out sensitive words[^9]) to circumvent censorship obscure the censorable nature of the content, making the prediction problem even more difficult and further exacerbating Problem 2.

<figure id="fig:toy" data-latex-placement="ht">
<div class="center">
<embed src="figs/toy_paper_4.pdf" style="width:95.0%" />
</div>
<p><span><em>Notes:</em> The shaded regions in the data generating processes (DGP) of the training data and the normal time test data indicate right-censoring. This causes content with high sensitivity to be missing in the observed data. Given the observed training data, censorship AI solves a binary classification problem of predicting whether content should be censored or not. Test data is used to evaluate the performance of censorship AI.</span></p>
<figcaption>Stylized Example of Censorship AI Training and Testing</figcaption>
</figure>

The impact of strategic behavior on problems 1 and 2 depends on the extent of such strategic behavior: as the level of repression and censorship increases and people respond with more preference falsification and self-censorship (causing the shaded region in Figure <a href="#fig:toy" data-reference-type="ref" data-reference="fig:toy">1</a> to grow larger), the data problems become more severe, potentially further reducing the effectiveness of the AI system.

## AI during Crises

Arguably, while autocrats care about AI’s effectiveness in authoritarian control during normal times, it is during periods of sudden political crisis when its performance becomes paramount. Contrary to the autocrats’ wishes, however, I argue that the authoritarian data problem is likely to cause a greater decline in the performance of AI during sudden political crises than during normal times. Here, normal times refer to periods of political and social stability – in authoritarian regimes – when the extent of strategic behavior is relatively stable. In the context of AI, this means the test data the AI encounters is similar to its training data (as represented by the equivalence of the training and normal time test data generating processes in Figure <a href="#fig:toy" data-reference-type="ref" data-reference="fig:toy">1</a>). On the other hand, crises refer to times of abrupt political turmoil, such as protests and coups. As I argue below, crises often entail significant changes in the data distribution, which can hamper AI performance.

There are several forces at play that cause further underperformance of AI during crises. First, during these moments, phenomena like information cascades and revelation of regime weakness allow citizens to learn about shared grievances and embolden them to reduce or even abandon preference falsification and self-censorship (Lohmann 1994; Kuran 1997). This means the test data during a crisis is drawn from a wider, less censored distribution (i.e., the shaded region in Figure <a href="#fig:toy" data-reference-type="ref" data-reference="fig:toy">1</a> shrinks or disappears). This mismatch between the AI’s training data and the test data encountered during the crisis can significantly degrade AI performance, as much content typical of crises was suppressed during normal times and thus unseen during training. Furthermore, the shape of the data generating distribution may be entirely different during crises. Instead of a dominance of less sensitive content during normal times, the modal content may be sensitive and censorable during crises. This further widens the between AI’s training data and the crisis test data, causing a larger performance drop.

Second, new types of content and new forms of behavior may emerge during crises that are absent in the AI’s training data: protest movements might rapidly coin new slogans, symbols, or hashtags specific to the unfolding events. Actions may be given new meanings and discussions of entirely new topics may spring up. For instance, during protests against China’s COVID-19 restrictions, people holding a blank piece of paper became a symbol of resistance. During the Arab Spring in Egypt, an eyepatch became a powerful symbol of solidarity with protesters who had been shot in the eye by security forces. Similarly, the three-finger salute, first popularized by *The Hunger Games* books and films, later became a symbol of rebellion for the pro-democracy movements in Thailand and Myanmar. These emergent forms of expression represent patterns the AI system either never encountered or associated with different meanings during its training on data from normal times. Unlike content and behavior suppressed during normal times, these new types of content and emergent behaviors may be even more difficult for AI to anticipate and generalize to. Consequently, AI systems may fail to identify or correctly interpret these novel signals, leading to significant blind spots in its coverage precisely when the regime perceives the greatest threat.

Fundamentally, the driving force for AI’s underperformance in crises is that crises often entail sudden changes in people’s behavior. If such changes in behavior are not well represented (meaning there are few such examples) in the training data, then the AI is ill-equipped to deal with the crisis situation. Conversely, if the crisis unfolds slowly or if similar crises are repeated over time, the gradual and repeated nature would allow these emergent behaviors to be incorporated into the training data. This would, in turn, reduce the distribution shift and enable the AI to adapt, thereby lessening the performance decline when similar situations arise.

## The Role of International Data

What can autocrats do in light of the data constraints on AI? One approach is to simply collect more data. However, the additional data will likely suffer from the same quality issues if it is collected from the same data generating process that is tainted by strategic behavior. In other words, sampling more from the biased distribution does not correct for the (distributional) bias. Once there is enough data for AI to learn about the biased distribution, simply collecting more data without changing the constraints under which citizens generate data should have a marginal impact on the performance of AI.

On the other hand, however, if autocrats can somehow collect data from the right-censored parts of the data generating process (i.e., content that is self-censored or that reflects citizens’ withheld private preferences), then the performance of AI can be improved. One way autocrats can do so is by collecting data that is not generated *domestically* but *internationally*, especially from democracies where citizens do not face the same political constraints. For instance, content from diaspora communities on international social media platforms such as Twitter and Facebook may contain valuable information that is suppressed domestically. A censorship AI that is trained on domestic data augmented by international data may thus be more accurate in censorship than the AI trained on domestic data alone. For autocrats, collecting data from international sources not only boosts the performance of AI but also has the advantage of keeping the level of repression and censorship unchanged at home. Qualitative evidence suggests that such practice is already used systematically on a large scale and across authoritarian regimes.[^10]

How well data augmentation from international sources works depends on how similar such data is to the right-censored parts of the data generating process. In particular, there needs to be sufficient overlap in the topics and semantics between international and domestic sources. Furthermore, data from international sources needs to be diverse enough in terms of political sensitivity to cover the entire span of right-censoring in domestic sources. Existing evidence suggests that international sources of information are qualitatively different from domestic sources, both in terms of topical distribution as well as political sensitivity (Esberg and Siegel 2023). Such differences will limit the effect of data augmentation on AI performance.

## Summary

In summary, I leverage theories of citizens’ strategic behavior in authoritarian regimes to explain the (under-)performance of AI for authoritarian control. Specifically, the theory implies the following two sets of hypotheses.

**Repression-performance trade-off:**

1.  As repression and censorship increase and people engage in more strategic behavior, the performance of AI on authoritarian control will become worse.

2.  The drop in AI’s performance is larger during times of political crisis than during normal times.

**Data augmentation:**

1.  More data collection under the same data generating process has a marginal impact on performance.

2.  Data from international (especially democratic) sources can improve AI’s performance, but may not solve the authoritarian data problem entirely.

# Data and Research Design

I chose AI used to automate censorship as the empirical setting. The choice is motivated by the fact that recent advances in AI have focused on textual data and have lead to widespread adoption of automated censorship technology in many authoritarian countries (Shahbaz et al. 2023). In the context of censorship, AI solves a binary classification problem: given a social media post, predict whether its label should be $`0`$ (not censor) or $`1`$ (censor). In practice, a censorship AI is trained by fine-tuning a pre-trained model with labeled censorship data. The pre-trained model is usually a general open-source deep learning model and the labeled censorship data consists of social media posts with their corresponding censorship labels.[^11]

## Repression-performance Trade-off

To test the theory’s hypotheses on the repression-performance trade-off, the ideal empirical set-up would involve multiple parallel worlds where the AI technology is fixed but the data generating process is subject to varying degrees of strategic behavior. The performance of censorship AI models from these worlds could then be compared.

To approximate the ideal set-up, I conduct an AI experiment that recreates as closely as possible the actual training of censorship AI models in practice, while varying the training data to reflect different degrees of strategic behavior. Specifically, the experiment uses 1) the same AI algorithm that many technology companies use, 2) training data consisting of millions of real-world user-generated content, and 3) training procedures leveraging state-of-the-art computing hardware. Using a unique dataset of social media posts for which the political sensitivity is known, the experiment compares the accuracy of censorship AI models trained on the different training data. The experiment first uses domestic data to test hypotheses 1a, 1b, and 2a, and then incorporates international data to test hypothesis 2b.

To construct the domestic training data, I first combine two datasets of Chinese social media posts from previous studies (Fu and Zhu 2020; Hu et al. 2020). The social media posts, totaling more than $`10`$ million in size, are on the topics of COVID-19 and were posted on Weibo, a Chinese social media platform, during the early period of the COVID-19 pandemic (Dec. 2019 – Feb. 2020). I focus on the early period of the pandemic because this was when censorship of COVID-19 topics had not caught up[^12] and therefore the social media posts have a relatively wide distribution of political sensitivity. The combined dataset serves as the basis from which I construct different versions of training data and use them to train censorship AI models specifically for COVID-19.

To get the political sensitivity of the social media posts, I use an automated censorship service from a Chinese technology company.[^13] The service is sold to smaller social media companies to help conduct censorship. It takes the text of social media posts as input and outputs a political sensitivity score that ranges from $`0`$ to $`1`$ for each post, with $`1`$ being the most sensitive. In the experiment, the political sensitivity scores serve as the latent variable. I use the service’s default sensitivity score of $`0.5`$ as the threshold to generate the binary censorship labels – social media posts with scores above $`0.5`$ have a label of $`1`$ (censor) and posts with scores below $`0.5`$ have a label of $`0`$ (not censor). The social media posts and their censorship labels can then be used as training data for censorship AI. Following standard industry practice, I down-sample social media posts with labels of $`0`$ to partially account for the imbalance in the proportion of the two classes ($`0`$ and $`1`$) of labels. Therefore, the main sample has a size of $`1`$ million social media posts.

To model different degrees of data missingness due to strategic behavior like preference falsification and self-censorship, I use the main sample to construct training datasets that differ in their distribution of political sensitivity. As the top row of Figure <a href="#fig:design" data-reference-type="ref" data-reference="fig:design">2</a> shows, I construct five versions of training dataset with varying degrees of missingness. To model the case where there is no missingness due to strategic behavior, I use the entire sample as the training dataset. The other four versions use different thresholds ($`0.9`$, $`0.8`$, $`0.7`$, $`0.6`$) above which the corresponding social media posts are missing from the training dataset. A threshold of $`0.6`$ means that only social media posts with sensitivity scores between $`0`$ and $`0.6`$ are in the training dataset. This models the most extreme case in which the regime is highly repressive and there is a high degree of preference falsification and self-censorship.

The design assumes that citizens have perfect information about what content is censorable and that they respond accordingly given the level of repression. In the appendix, I account for imperfect information and coordination by allowing five percent of the data from the missing part of the distribution to leak into the training datasets. The substantive conclusions remain unchanged.

<figure id="fig:design" data-latex-placement="ht">
<div class="center">
<embed src="figs/design_paper_v3.pdf" />
</div>
<p><span><em>Notes:</em> Graphical representation of the research design. There are five versions of training dataset corresponding to different degrees of strategic behavior. The test datasets are drawn from the same distributions of their corresponding training datasets. All datasets are drawn from the full distribution. Note that this is a stylized representation. The shapes of the actual distributions are different from the graph.</span></p>
<figcaption>Experimental Design</figcaption>
</figure>

For each version of the training dataset, I train a separate censorship AI model on it. Specifically, I use the Chinese version of BERT (Bidirectional Encoder Representations from Transformers; Devlin et al. (2018)) as the pre-trained model and fine-tune it on the training datasets for censorship. BERT is a deep learning model with $`110`$ million parameters and was developed by Google. Since its introduction, BERT has been one of the most popular deep learning models for prediction and is widely used in commercial applications. In the Appendix, I provide details about how the model is used in practice for censorship based on fieldwork in technology companies. To account for the uncertainty from data sampling and the stochastic nature of the fine-tuning process, each version of the training dataset is used to train $`25`$ models with the training data shuffled each time. This allows me to obtain uncertainty estimates for model performance. More details about the training procedure are included in Appendix A.2.

To evaluate the performance of the different censorship AI models, I follow the theory and construct two kinds of test data: normal time and crisis (second and third rows of Figure <a href="#fig:design" data-reference-type="ref" data-reference="fig:design">2</a>). Social media posts in the normal time test data are drawn with the same level of missingness as the corresponding version of the training dataset, whereas the crisis test data is always drawn from the full distribution. Mirroring the theory, this is to model the situation in which citizens maintain their level of preference falsification and self-censorship during normal times but start revealing their true preferences during crises. Both sets of test data are sampled from social media posts that are not in the training data. To be able to compare performance evaluated on different test data, each test dataset is a balanced sample of 1000 positive labels (i.e. censor) and 1000 negative labels (i.e. not censor). To measure the performance of censorship AI, I use accuracy, defined as the fraction of correct predictions over the total number of predictions, in the main text and report other measures of performance in the Appendix. Here, a correct prediction means that the model predicts a censorship label that is the same as the ground truth label.

Admittedly, the design only captures one kind of content change during crises: the revelation of previously suppressed information. It does not, however, fully model the emergence of entirely new topics or behaviors during crises. This is a limitation of the design because, as argued previously, these emergent phenomena can be even more difficult for AI systems to anticipate and generalize to. Consequently, the performance degradation observed in the experiment may represent a conservative estimate of the challenges AI would face during complex, real-world crises.

## Data Augmentation

To test the effect of more data on AI’s performance (hypothesis 2a), I follow the same experimental set-up as above but double the size of the initial sample from 1 million to 2 million while keeping the distribution of political sensitivity unchanged. A new set of censorship AI models are trained using the larger training datasets and their accuracy is compared with the original models.

To test the effect of data from international sources (hypothesis 2b), I scraped all 558322 Chinese tweets from Twitter that are on the same COVID-19 topics and were posted during the same period as the Weibo data. The political sensitivity of the Twitter data is also obtained through the automated censorship service. I then construct the Twitter dataset of 219111 tweets with political sensitivity scores above $`0.5`$ and use it to augment the Weibo training datasets.[^14] Notably, the same Twitter dataset is used to augment all training datasets, assuming that there is no data missingness from international sources as a result of changing domestic repression levels. A new set of censorship AI models are trained using the augmented training datasets and their accuracy is compared with the original models.

# Results

I first present evidence of the repression-performance trade-off. I then show evidence that adding more domestic data during training has a marginal impact on AI’s performance but data augmentation from international sources results in a larger improvement in AI’s censorship accuracy.

## More Repression, Worse AI Performance

Figure <a href="#fig:main" data-reference-type="ref" data-reference="fig:main">3</a> presents evidence of the repression-performance trade-off. It shows the accuracy of censorship AI models trained with datasets of varying degrees of data missingness. The threshold (x-axis) indicates the political sensitivity score above which data is missing from the training dataset. The threshold of $`1.0`$ means the training dataset has no missing data and the threshold of $`0.6`$ has the most missing data. Model accuracy is evaluated on both the normal time and crisis test data.

<figure id="fig:main" data-latex-placement="ht">
<div class="center">
<img src="figs/main.png" />
</div>
<p><span><em>Notes:</em> Each threshold value represents a version of the training dataset. Uncertainty estimates are obtained based on the predictions of <span class="math inline">25</span> models for each threshold.</span></p>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Evaluations on the normal time data (blue) show that as data missingness increases as a result of strategic behavior, the accuracy of the censorship AI model decreases, with the worst-performing model being trained on the dataset with the most missingness. A similar downward trend is also observed for the crisis test data (red). In line with the theory, the drop in model accuracy is significantly larger in crises, when people reveal their true preferences, than in normal times. In the appendix, I show that the drop in model accuracy is largely due to an increase in false negatives (censorable content predicted to be ). This is a particularly bad situation for autocrats as false negatives allow transmission of politically sensitive information among citizens and thus may be more costly for autocrats than false positives (censoring more than they should).

Table <a href="#tab:example" data-reference-type="ref" data-reference="tab:example">1</a> provides two examples of social media posts and the corresponding predictions from different models. The first example expresses disappointment at the government’s handling of the pandemic and the second example mocks a research lab for promoting traditional Chinese medicine. In both cases, the censorship model trained on data with no missing data (threshold = 1.0) correctly predicts their censorship labels while the model trained with missing data (threshold = 0.6) gets both wrong.

<div id="tab:example">

|  |  |  |  |
|:---|:--:|:--:|:--:|
| Social media post (translated) | Ground | Prediction | Prediction |
|  | truth | (thres.=1.0) | (thres.=0.6) |
| Originally, I had a lot of confidence in how the pandemic was being handled since the central government took over, but now all these developments are really disappointing. \[Sad emoji\] | censor | censor | not censor |
| The Wuhan Institute of Virology, a world-class P4 biosafety lab, believes that Shuanghuanglian oral liquid can inhibit the virus. Now is truly the pinnacle of traditional Chinese medicine history. \[Sarcastic emoji\] | censor | censor | not censor |

Example Social Media Posts and Censorship Predictions

</div>

Critically, the decrease in model accuracy and the performance gap between normal time and crisis test data, as observed in Figure <a href="#fig:main" data-reference-type="ref" data-reference="fig:main">3</a>, depend on the strategic nature of citizen behavior. The same patterns would not be observed if people withhold content in a way that is unrelated to political sensitivity. Figure <a href="#fig:random" data-reference-type="ref" data-reference="fig:random">4</a> shows one such scenario in which data is missing at random. The x-axis indicates the percentage of missing data as a fraction of the original 1 million sample. Despite having missing data at the same scale as the set-up in Figure <a href="#fig:main" data-reference-type="ref" data-reference="fig:main">3</a>, Figure <a href="#fig:random" data-reference-type="ref" data-reference="fig:random">4</a> shows no significant performance difference across models and between normal times and crises.

<figure id="fig:random" data-latex-placement="ht">
<div class="center">
<img src="figs/random.png" />
</div>
<p><span><em>Notes:</em> Each threshold value represents a version of the training dataset. Uncertainty estimates are obtained based on the predictions of <span class="math inline">25</span> models for each threshold.</span></p>
<figcaption>Model Performance with Non-strategic Missing Data</figcaption>
</figure>

In the Appendix, I provide additional evidence that the substantive conclusions are robust to various changes to the experimental set-up, such as using a larger BERT model, changing the censorship decision rule (e.g., from $`0.5`$ to $`0.4`$), allowing some leakage of the missing data into the training data, adding various performance enhancing techniques, and using state-of-the-art large language models like those developed by DeepSeek and Alibaba.

## Marginal Impact of More Domestic Training Data

Given the previous results, one of the ways autocrats may choose to respond to the data problem is to collect more data and train the model on a larger dataset. Figure <a href="#fig:double" data-reference-type="ref" data-reference="fig:double">5</a> presents the result of doubling the size of the training dataset on model accuracy. Specifically, Figure <a href="#fig:double" data-reference-type="ref" data-reference="fig:double">5</a> shows the difference in accuracy, on both test data, between models trained with double the amount of training data and those trained with the original training datasets. Across settings, the improvement in model accuracy from doubling the size of the training data is marginal – the largest accuracy improvement is smaller than two percentage points.

Figure <a href="#fig:double" data-reference-type="ref" data-reference="fig:double">5</a> thus provides evidence that additional data that is collected under the same informational environment where there is preference falsification and self-censorship has a marginal impact on the performance of censorship AI.

<figure id="fig:double" data-latex-placement="ht">
<div class="center">
<img src="figs/double.png" />
</div>
<p><span><em>Notes:</em> Y-axis shows the average difference in accuracy between models trained on the original Weibo data and models trained on the larger (doubled in size) Weibo data. Each threshold value represents a version of the training dataset. Uncertainty estimates are obtained based on the predictions of <span class="math inline">25</span> models for each threshold.</span></p>
<figcaption>Effect of Doubling Domestic Training Data</figcaption>
</figure>

In the Appendix, I show that the breakdown of the errors by models trained with the larger training datasets follows a similar trend to the original models – the false positive rate is low and stays relatively stable across different thresholds but the false negative rate increases drastically as data missingness increases.

## Accuracy Improvement from International Data

While additional data collected domestically provides little improvement to model accuracy, data from international sources, generated largely without the same political constraints, should boost model performance. Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> provides evidence that augmenting the original Weibo training data with data from Twitter improves model accuracy, especially for performance during crises.

<figure id="fig:twitter" data-latex-placement="ht">
<div class="center">
<img src="figs/twitter.png" />
</div>
<p><span><em>Notes:</em> Y-axis shows the average difference in accuracy between models trained on the original Weibo data and models trained on the Twitter-augmented data. Each threshold value represents a version of the original Weibo training dataset. Uncertainty estimates are obtained based on the predictions of <span class="math inline">25</span> models for each threshold.</span></p>
<figcaption>Effect of Twitter Data Augmentation</figcaption>
</figure>

Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> compares the accuracy of models trained on datasets augmented by the Twitter data with models trained on the original Weibo training datasets. Similar to domestic data augmentation, the Twitter data augmentation provides a marginal improvement on the normal time test data. This is because the normal time test data is sampled from the same distribution as the original training data. In this case, the decrease in performance for both sets of models (original and Twitter-augmented) is due to the increase in similarity between censorable and content rather than a mismatch in distribution between the training and test data. As augmentation does not change the fact that the prediction problem for the normal time data becomes more difficult as the threshold decreases, the Twitter data thus provides litter accuracy improvement.

On the other hand, Twitter data augmentation improves the accuracy of censorship AI models during crises. Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> shows that, when there is missing data (thresholds $`0.6 - 0.9`$), the accuracy of the models trained on the augmented datasets is substantially higher than the models trained on the original Weibo data, with the largest improvement being more than six percentage points. This is despite the fact that the Twitter data is only about one-fifth of the Weibo data in size. Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> thus shows that Twitter data can partially compensate for the missing data from Weibo and reduce the mismatch in distribution between the training data and crisis test data.

It is important to note, however, that the accuracy improvement from Twitter data is limited, in that the models’ accuracy is still significantly lower than that of the models trained on the full Weibo data (threshold=1.0). One potential explanation for this is that the content on Twitter is different from the content on Weibo so relying on Twitter data augmentation cannot fully compensate for the missing Weibo data.

Figure <a href="#fig:stm" data-reference-type="ref" data-reference="fig:stm">7</a> provides suggestive evidence that the content in the Weibo data is indeed different from the content in the Twitter data. It shows the topic prevalence across the two data sources. Topic prevalence is estimated with a structural topic model (Roberts et al. 2014) using the combined Weibo and Twitter data[^15], where the number of topics is set to 15. Topic labels are written manually based on the top keywords and the most representative posts for each topic.[^16]

Figure <a href="#fig:stm" data-reference-type="ref" data-reference="fig:stm">7</a> shows that politically sensitive topics (e.g., discussions of virus origins, political system, and criticism of lockdown policies) are substantially more prevalent in the Twitter data than in the Weibo data. For the topic on virus origins and international research, for example, it is more than three times as prevalent in the Twitter data than in the Weibo data. Additionally, international issues (topic on global pandemic and international impact) are also more prevalent in the Twitter data. In contrast, content on Weibo consists of more discussions of domestic issues such as daily updates on COVID-19 cases and local community prevention and control. Furthermore, content on Weibo includes more support (instead of criticism) for local and national COVID measures, with the topic on support for Wuhan and national solidarity being 8.8 percentage points more prevalent in the Weibo data than in the Twitter data. In Appendix E.2, I provide evidence that the difference in content propagates to censorship AI models, where the models’ internal representation of the Twitter and Weibo data shows spatial differences between the two data sources.

<figure id="fig:stm" data-latex-placement="ht">
<div class="center">
<img src="figs/tp1.png" />
</div>
<figcaption>Topic Prevalence across Weibo and Twitter</figcaption>
</figure>

Together, Figure <a href="#fig:double" data-reference-type="ref" data-reference="fig:double">5</a> and Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> provide evidence for the theory’s data augmentation hypotheses: more data from domestic sources has a marginal impact on model accuracy but data from international sources helps improve model performance. Figure <a href="#fig:twitter" data-reference-type="ref" data-reference="fig:twitter">6</a> also shows the limit of data augmentation from international sources in boosting censorship AI’s performance and Figure <a href="#fig:stm" data-reference-type="ref" data-reference="fig:stm">7</a> suggests that the difference in content between domestic and international sources likely contributes to the limit.

# Discussion

Artificial Intelligence has become a key technology in the autocrats’ toolkit and will be increasingly so in the foreseeable future. Its ability to ingest vast amounts of data and make predictions based on the data – a capability powerfully demonstrated by the recent success of large language models like DeepSeek – undoubtedly enables contemporary autocrats to sift through information at a scale their historical counterparts could not have imagined. Despite AI’s powerful potential for authoritarian control, I argue that there are inherent limits to AI’s ability and that these limits are the result of existing authoritarian institutions. Just like their traditional counterparts, digital autocrats face a dilemma between repression and information: the more repression there is, the less political information there will be in the data, and the worse AI will perform. Regardless of its capabilities, AI cannot process or aggregate information that is not observed.

The theory and empirical findings of this paper provide some nuance to the ongoing debate on the effect of AI on authoritarian control. By problematizing the argument that more data means better prediction and better control and by bringing to the forefront the issue of data quality, this paper argues that the general equilibrium effect of AI may not be as favorable toward autocrats as the existing literature has argued.

An important limitation of the paper is that, in order to identify the effect of data on AI, it considers AI in isolation. In reality, AI is often used as a complement to, rather than a replacement for, humans in repressive tasks. For censorship, the explosion of online content often requires that ther is human-AI collaboration, which can take the form of AI serving as the initial content screener and human censors then reviewing content flagged by AI.[^17] It could be that human censors would then be able to correct mistakes made by AI. However, as shown in the experiment, if the majority of misclassifications are false negatives, it may be more difficult for humans to mitigate the problem. Additionally, human censors may also update the training data to lessen the impact of new content during crises. Future work could explore these dynamics further.

Furthermore, the theory of the paper relies on the assumption that in the face of increasing repression and censorship, people will falsify their preferences and self-censor more, causing greater data missingness. This is not a completely innocuous assumption. Although there is substantial empirical evidence supporting this assumption (Fu et al. 2013; Huang 2015; Tanash et al. 2017) and it is in fact the premise of the dictator’s dilemma in Wintrobe (2000), studies have shown that repression can generate both chilling and backlash effects (Huang 2018; Pan and Siegel 2020).[^18] The scope conditions for the backlash effect identified in the literature are that repression and censorship are overt and visible to the public and that they are not strong enough to stifle most citizens’ reactions (Pan and Siegel 2020; Roberts 2020). In the context of digital repression and censorship, which are more covert and all-encompassing by nature (Xu 2021), these scope conditions may be too stringent, potentially limiting the backlash effect. On the other hand, if there is indeed a substantial backlash effect, by the logic of the theory, this can have an unintended consequence of providing valuable information to the training data and boosting repressive AI’s performance.

The theory also points to similar unintended consequences of political phenomena that work in the digital autocrats’ favor. For example, polarization in authoritarian regimes can make the prediction problem easier. This is because, as the online discussion polarizes, the censorable content will be easier to identify by AI as their similarity with non-censorable content decreases. This serves as an additional channel, on top of the ones the existing literature has identified (Svolik 2018; Svolik 2019), through which (would-be) autocrats can use polarization to strengthen their rule.

Similarly, the theory suggests that if there are alternative, non-domestic platforms on which citizens can express dissent, then the repression-performance trade-off may be partially mitigated when autocrats also collect data from these platforms. Several recent studies have documented the migration of dissent from domestic to international platforms (Hobbs and Roberts 2018; Esberg 2022; Esberg and Siegel 2023). In the context of AI, this can work in the autocrats’ favor, as this allows them to collect uncensored information without changing the repressive environment domestically. However, as the paper demonstrates, the effect of international data may be limited in its impact on AI performance, especially when discussions from international sources diverge from domestic sources.

While not explicitly spelled out, the paper points to the possibility that data from democracies boosting authoritarian AI is only half the story. By the same logic, data from authoritarian regimes can serve to contaminate AI from democracies. Given that major AI companies in the U.S. and Europe are relying on ever larger datasets to train their AI models, it is likely that data tainted by censorship and propaganda can influence the output of these models (Yang and Roberts 2021, 2023). This can be especially concerning considering that such AI models are being deployed in important areas such as education and criminal justice. Documenting data leakages from authoritarian regimes and quantifying their effect on AI are worth exploring in future research.

**Appendix *for***

<div class="center">

**The Limits of AI for Authoritarian Control**\
Eddie Yang\
(Purdue)

</div>

<div class="enumerate">

Censorship AI Model Details

Model Training Details

Details on the Weibo and Twitter Data

Political Sensitivity Service

Details on the Training Datasets

International Social Media Data Collection by Authoritarian Regimes

Alternative Measure of Performance

Error Rate Results

AI Performance across Topics

Robustness to Weaker Assumptions and Performance Enhancing Techniques

Keywords for Each Topic from Structural Topic Model

AI Model’s Internal Representation of Weibo and Twitter Data

</div>

# Censorship AI Model Details

Except for the results on larger model and alternative model architecture in Section <a href="#robustness" data-reference-type="ref" data-reference="robustness">16</a>, the pre-trained BERT model used in the paper is the Chinese BERT (bert-base-chinese).[^19] The model has 110 million parameters and has been shown to perform well on a variety of Chinese prediction tasks.

Based on information gathered in fieldwork, the model and its variants have been used extensively for commercial applications by technology companies. Among other applications, variants of the model have been used for censorship and more generally for content moderation (e.g., detecting pornography and spam). In contrast to more recent generative AI models such as ChatGPT, BERT is better suited for prediction tasks and is in general much cheaper and faster for inference/prediction.

# Model Training Details

The BERT models are fine-tuned using the **transformers** library provided by Hugging Face.[^20] To fine-tune the models, the social media posts in the training data need to be converted into strings of tokens (tokenization) that correspond to the internal dictionary of the BERT model. Tokenization is also provided as part of the **transformers** library.

During training, the F1 score is used as the evaluation metric to track model performance. The F1 score is defined as $`F1 = 2 * (\text{precision} * \text{recall})/(\text{precision} + \text{recall})`$, where precision is given by $`(\text{no. of true positives})/(\text{no. of true positives} + \text{no. of false positives})`$ and recall is given by $`(\text{no. of true positives})/(\text{no. of true positives} + \text{no. of false negatives})`$.

Early stopping was used to prevent over-fitting. Specifically, training was stopped if the F1 score on the validation set did not improve for three epochs. To speed up training, I used mixed precision training (fp16) for all models in the experiment. The full list of hyperparameter values is provided in Table <a href="#tab:hyper" data-reference-type="ref" data-reference="tab:hyper">2</a>.

<div id="tab:hyper">

| Hyperparameter          |     Value     |
|:------------------------|:-------------:|
| maximum token length    |      128      |
| fp16                    |     True      |
| batch size              |      512      |
| learning rate           |    0.0001     |
| learning rate scheduler |    cosine     |
| early stopping patience |       3       |
| warmup steps            |      500      |
| maximum training steps  |     10000     |
| optimizer               | AdamW (fused) |

Hyperparameters

</div>

Notes: The table shows the hyperparameter values used in the training experiment.

The fine-tuning experiment was conducted using Nvidia A100 and H100 GPU with 80GB memory. The total GPU hours are around 1700.

# Details on the Weibo and Twitter Data

Weibo data from Fu and Zhu (2020) were collected by the authors based on a list of $`40`$ COVID-19-related keywords. The data contains 1230353 posts that were posted between December 1, 2019 and February 27, 2020 on Weibo.

Weibo data from Hu et al. (2020) were collected for a longer time span (December 1, 2019 - December 31, 2020) and were based on a more extensive list of keywords. To ensure compatibility, I use a subset of the data that includes posts that were posted between December 1, 2019 and February 27, 2020 and contain at least one of the $`40`$ keywords in Fu and Zhu (2020). The subset contains 8518113 Weibo posts.

All Weibo posts were anonymized to remove tags and other user information.

Twitter data was collected using the Twitter research API with the following restrictions: 1) the tweets were posted between December 1, 2019 and February 27, 2020; 2) the tweets contain at least one of the $`40`$ keywords in Fu and Zhu (2020); and 3) the language of the tweets is identified as Chinese by Twitter. The restrictions were used to ensure compatibility with the Weibo data. Similar to the Weibo data, the Twitter data was anonymized to remove tags and other user information.

# Political Sensitivity Service

The political sensitivity service is provided by Baidu, a major Chinese technology company, and is publicly available as of the writing of this paper.[^21] The service uses a combination of banned keywords collected by Baidu and deep learning models to assign political sensitivity to text. The sensitivity score ranges from $`0`$ to $`1`$, with $`1`$ being the most sensitive. Table <a href="#tab:keyword" data-reference-type="ref" data-reference="tab:keyword">3</a> reports several keywords (and keyword combinations) that were flagged by the service. All of them seem to be sensible keywords that could be considered sensitive, especially during the early COVID-19 pandemic.

<div class="CJK*">

UTF8gbsn

<div id="tab:keyword">

| Keyword(s)     | Translation          |
|:---------------|:---------------------|
| 政府, 蛀虫     | government,          |
|                | parasite             |
| 武汉, 问责     | Wuhan,               |
|                | accountability       |
| 湖北, 瞒报     | Hubei,               |
|                | withhold information |
| 颜色革命       | color revolution     |
| 中国经济, 衰退 | Chinese economy,     |
|                | slowdown             |

Flagged Keywords

</div>

Notes: The table shows example keywords (and their translations) that have been flagged by the censorship service.

</div>

# Details on the Training Datasets

From the combined Weibo data, a $`2.3`$ million sample was labeled by the political sensitivity service. The size of the sample was primarily determined by resource constraints. From the $`2.3`$ million labeled data, a 2 million sample was constructed where censorable posts were upsampled to reduce label imbalance (there are many more content than censorable content). The 2 million sample is then randomly split into two datasets of 1 million social media posts, one of which serves as the baseline Weibo data for fine-tuning models. Table <a href="#tab:data" data-reference-type="ref" data-reference="tab:data">4</a> shows the summary statistics of different versions of the baseline training datasets.

The entirety of the Twitter data was labeled by the political sensitivity service. Twitter posts for which the sensitivity scores are above $`0.5`$ are then used as the Twitter augmentation dataset.

<div id="tab:data">

| Training dataset | Thres-hold | positive labels | negative labels |
|:-----------------|:-----------|:----------------|:----------------|
| Version \#1      | 1.0        | 183511          | 816489          |
| Version \#2      | 0.9        | 119139          | 816489          |
| Version \#3      | 0.8        | 92872           | 816489          |
| Version \#4      | 0.7        | 64142           | 816489          |
| Version \#5      | 0.6        | 36738           | 816489          |

Summary Statistics of Training Datasets

</div>

Notes: The table shows the the number of posts by censorship status across the five versions of the training dataset.

Figure <a href="#fig:dists" data-reference-type="ref" data-reference="fig:dists">8</a> shows the distribution of the Weibo and Twitter training data. Both distributions show a bimodal shape, where there is a large quantity of both non-sensitive and very sensitive content. In comparison, the Twitter data comprises a larger proportion of sensitive data whereas the majority of the Weibo data is non-sensitive.

<figure id="fig:dists" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/dists.png" />
</div>
<figcaption>Distributions of Training Data</figcaption>
</figure>

# International Social Media Data Collection by Authoritarian Regimes

Here I present some qualitativa evidence that social media data from international platforms such as Twitter, Facebook, YouTube and TikTok is being collected en mass by authoritarian regimes.

Publicly available information suggests that large sclae Chinese Twitter data has been collected and used for AI training in China. The Natural Language Processing and Information Retrieval sharing platform hosted by the Beijing Institute of Technology shows that at least a hundred million Chinese tweets have been collected and from which five million is made publicly available.[^22] The Peacock Chinese Twitter Corpus (PCTC) is another dataset of 4.9 millionn Chinese tweets.[^23] Information gathered in fieldwork also confirms that international social media data is being used to augment AI training data by Chinese technology companies.

Similarly, leaked documents from Russia suggest that Russia is monitoring and collecting massive amount of social media data from platforms like Twitter, Facebook, YouTube, and Tiktok and is in the process of using such data to develop automated censorship systems.[^24] In particular, documents show that one Russian company has been collecting data on the scale of 140 million messages in Russian and other languages spoken in the former Soviet Union and 40 million images per day from Facebook, Instagram, TikTok, Twitter, and other social media platforms since 2014.[^25]

# Alternative Measure of Performance

In addition to accuracy, another commonly used metric to evaluate the performance of deep learning models is the F1 score. The F1 score is defined as
``` math
F1 = 2 * \frac{\text{precision} * \text{recall}}{\text{precision} + \text{recall}}
```
where precision is given by $`(\text{no. of true positives})/(\text{no. of true positives } + \text{ no. of false positives})`$ and recall is given by $`(\text{no. of true positives})/(\text{no. of true positives } + \text{ no. of false negatives})`$.

In simple terms, precision is the ability of a model to correctly identify positive instances (true positives) out of the total instances it predicts as positive. It focuses on minimizing false positives, meaning the instances that are wrongly classified as positive (censor). Recall is the ability of a model to correctly identify all the positive instances (true positives) out of the total actual positive instances. It focuses on minimizing false negatives, meaning the instances that are wrongly classified as negative (not censor).

<figure id="fig:f1" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/main_f1.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

The F1 score combines precision and recall into a single metric by taking their harmonic mean. The harmonic mean gives more weight to lower values, so the F1 score will be high only if both precision and recall are high. It ranges between 0 and 1, with 1 indicating perfect performance and 0 indicating poor performance.

Figure <a href="#fig:f1" data-reference-type="ref" data-reference="fig:f1">9</a> reports the F1 scores for the censorship AI models trained on different training datasets. Similar to the main results, the result based on the F1 score shows that as data missingness increases, the performance of the censorship AI models becomes worse and the drop in performance is significantly larger for the crisis test data than for the normal time test data.

# Error Rate Results

While accuracy/F1 serves as an indicator of the overall performance of censorship AI models, it does not reveal the type of error that the models make. Specifically, the models’ errors can be false positives (where prediction is censorship but the actual label is non-censorship) or false negatives (where prediction is non-censorship but the actual label is censorship). The types of error have important implications for authoritarian rule, as false negatives (failing to censor) allow transmission of politically sensitive information among citizens and thus may be more costly for autocrats than false positives (censoring more than they should).

<figure id="fig:double_rate" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/err_rate.png" style="width:70.0%" />
</div>
<figcaption>Error Rate by Error Type - Original Models</figcaption>
</figure>

Figure <a href="#fig:double_rate" data-reference-type="ref" data-reference="fig:double_rate">10</a> breaks down the models’ errors by type. It reports the false positive rate, defined as $`\displaystyle\frac{\text{No. of false negatives}}{\text{Total no. of positives}}`$, and false negative rate, defined as $`\displaystyle\frac{\text{No. of false positives}}{\text{Total no. of negatives}}`$, for different censorship AI models. As Figure <a href="#fig:double_rate" data-reference-type="ref" data-reference="fig:double_rate">10</a> shows, the false positive rate is low and stays relatively stable across different thresholds. However, as data missingness increases, the false negative rate increases substantially, with the largest false negative rate more than three times that of the smallest. This is true for both the normal time test data and the crisis test data, with a larger increase in false negative rate during crises. Therefore, Figure <a href="#fig:double_rate" data-reference-type="ref" data-reference="fig:double_rate">10</a> points to a particularly bad situation for autocrats as censorship AI models are more likely to not censor truly censorable content when data missingness increases. Figure <a href="#fig:double_rate_2" data-reference-type="ref" data-reference="fig:double_rate_2">11</a> reports the error rates for models trained on double the amount of domestic data and the patterns are similar.

<figure id="fig:double_rate_2" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/err_rate_double.png" style="width:70.0%" />
</div>
<figcaption>Error Rate by Error Type - Doubled Training Data</figcaption>
</figure>

Intuitively, when censorable information is missing in the training data, signals about what should be censored (e.g., keywords that should trigger censorship) become more sparse in the data. As a result, more censorable posts will be able to pass the censorship models undetected because the models could not identify censorable markers from these posts.

# AI Performance across Topics

Given that the theory of the paper distinguishes between sudden political crises and crises that are either gradual in nature or have repeated over time, we might expect that the effectiveness of AI censorship varies by topic. Certain topics, such as critiques of government incompetence during a pandemic, may look similar to government critiques at other moments in time. In contrast, newer topics or topics specific to the pandemic may be more challenging for AI to accurately classify.

To test this hypothesis, I use data from Weiboscope (Fu et al. 2013) to train a new set of censorship AI models and then use the new models to test their performance on the COVID-19 Weibo data. The Weiboscope data contains posts from more than 350000 Chinese weibo users who have more than 1000 followers as well as from an additional random sample of Weibo users. The data was collected in 2012 and contain 226841249 posts. Importantly, the Weiboscope data records whether a post is censored. This is achieved by revisiting a post multiple times and observing whether it returns the error message (Fu et al. 2013).

To construct the training data, I take all 62017 censored posts and randomly sample the same number of uncensored posts from the Weiboscope data. In total, the training data has 124034 Weibo posts. I then fine-tune the BERT model on the training data and test the models’ performance on the COVID-19 Weibo data across different topics.

Figure <a href="#fig:weiboscope" data-reference-type="ref" data-reference="fig:weiboscope">12</a> shows the models’ accuracy across different topics in the COVID-19 Weibo data. The overall accuracy of the models is $`0.578`$, only slightly better than random guessing. This is unsurprising given the training data is eight years older than the test data.

<figure id="fig:weiboscope" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/weiboscope.png" style="width:70.0%" />
</div>
<figcaption>Performance of Weiboscope Models across Topics</figcaption>
</figure>

On the other hand, the accuracy shows large variations across topics. In particular, the models are much more accurate on topics related to 1) political system and social impact, 2) passing of Dr. Li Wenliang, and 3) Protective supplies than other topics. Close reading of these topics reveals that 1) contains many general criticism of the government, 2) frequently mentions the central government, and 3) often involves discussions of Taiwan and Hong Kong. It is possible that these recurring themes that have historically been subject to censorship in China caused the models to preform better on these topics. It also shows that, even for new topics (e.g., passing of Dr. Li Wenliang), if its discussion overlaps with past topics that have been subject to censorship, the censorship model’s performance decline may not be as severe as that on entirely new topics. Overall, Figure <a href="#fig:weiboscope" data-reference-type="ref" data-reference="fig:weiboscope">12</a> gives suggestive evidence that the effectiveness of AI censorship may vary by topic, with better effectiveness on topics that overlap with past, recurring themes.

# Robustness to Weaker Assumptions and Performance Enhancing Techniques

This section reports additional results from weaker assumptions and performance enhancing techniques commonly used in deep learning. The substantive conclusions from the main results hold for the additional tests reported in this section.

Figure <a href="#fig:leak" data-reference-type="ref" data-reference="fig:leak">13</a> shows the accuracy results when $`5\%`$ of the data above the threshold is allowed in the training data. This relaxes the assumption that citizens have perfect information about what content is censorable and that they are using a pure strategy of self-censorship and preference falsification according to the threshold. Noticeably, allowing data leakages improves the accuracy of the models trained on data with missing data (threshold \< 1.0). The improvement is larger for the crisis test data.[^26] However, the accuracy gaps between models of different thresholds still persist. Except for threshold = 0.6, the drop in accuracy is also larger for crises than normal times. Essentially, the model’s accuracy in crises is determined by a horse race between accuracy loss due to distribution shift (between the training and test data) and accuracy increase due to the prediction problem becoming easier (as the most politically sensitive hence easier to classify posts are now in the test data).

<figure id="fig:leak" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/leak.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Figure <a href="#fig:weight" data-reference-type="ref" data-reference="fig:weight">14</a> shows the results of using sample weights in the training. As Table <a href="#tab:data" data-reference-type="ref" data-reference="tab:data">4</a> shows, there is an imbalance in the number of censorable and posts in the training data. To correct for this imbalance, I use $`\displaystyle \frac{1}{prop_{c}} + 1`$ as sample weights in the training of censorship AI models, where $`prop_{c}`$, $`c \in \{\text{censorable, safe\}}`$ indicates the proportions of censorable and content in the training data. 1 is added to the inverse weighting for training stability. Sample weighting improves the performance of the models trained on data with missing data but does not eliminate the performance gaps between 1) models trained on data with different degrees of missingness and 2) the normal time and crisis test data.

<figure id="fig:weight" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/weight.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Figure <a href="#fig:weight_leak" data-reference-type="ref" data-reference="fig:weight_leak">15</a> combines data leakage and sample weighting. Here we see the benefits from both measures - the accuracy of the model goes up across the board, although the general patterns from the main results still hold.

<figure id="fig:weight_leak" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/weight_leak.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Another measure that can potentially change the accuracy of the censorship model is changing the decision rule - the value of political sensitivity above which posts are labeled as censorable. The main results in the paper uses $`0.5`$ as the decision rule. However, if the sensitivity of many censorable posts is predicted just below $`0.5`$, a lower decision rule can potentially improve the performance of the model. On the other hand, lowering the decision rule can also introduce more false positives, thus negatively affecting model performance.

Figure <a href="#fig:decision" data-reference-type="ref" data-reference="fig:decision">16</a> reports the accuracy results from using $`0.4`$ as the decision rule. The results are little changed as comapared to the main results and the substantial conclusions remain unchanged.

<figure id="fig:decision" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/rule.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Figure <a href="#fig:tradeoff" data-reference-type="ref" data-reference="fig:tradeoff">17</a> presents the trade-offs between lowering decision rule (hence fewer false negatives) and the false positive rate. To achieve an accuracy of 85% on censored content, the models with the lowest threshold would have on average a 9% false positive rate during normal times and 26% false positive rate during crises. Furthermore, to achieve an accuracy of 90% on censored content (which typically is the lowest acceptable value), models with the lowest threshold would have on average a 11% false positive rate during normal times and 59% false positive rate during crises. These high false positive rates may potentially be high costs for a censor even when they are less sensitive to false positives.

Additionally, to make models with lower thresholds achieve higher accuracy on censored content, the decision rule has to be set extremely low (often smaller than 0.01). An implication of such a small decision rule is that the model may become highly sensitive to minor variations in content and is therefore unlikely to generalize well beyond the specific test setting.

<figure id="fig:tradeoff" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/tradeoff.png" />
</div>
<figcaption>Trade-off between false positives and false negatives</figcaption>
</figure>

Figure <a href="#fig:large" data-reference-type="ref" data-reference="fig:large">18</a> shows the accuracy results for a set of larger models. Instead of the $`110`$ million parameter BERT model used for the main results, Figure <a href="#fig:large" data-reference-type="ref" data-reference="fig:large">18</a> uses the larger $`340`$ million parameter BERT model. Despite being three times larger, the larger models do not show a significant accuracy improvement or any deviation from the patterns of the main results.

<figure id="fig:large" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/large.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Figure <a href="#fig:electra" data-reference-type="ref" data-reference="fig:electra">19</a> shows the accuracy results for a set of models with a different model architecture. Instead of the BERT model, Figure <a href="#fig:electra" data-reference-type="ref" data-reference="fig:electra">19</a> uses the Chinese version of the ELECTRA model (Clark et al. 2020) as the deep learning model for training. The model has 102 million parameters and is trained using an adversarial framework. Figure <a href="#fig:electra" data-reference-type="ref" data-reference="fig:electra">19</a> shows that that the alternative model architecture yields similar results and the substantial conclusions are unchanged.

<figure id="fig:electra" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/electra.png" />
</div>
<figcaption>Model Performance across Training Datasets</figcaption>
</figure>

Given the continued development of AI, I also test the robustness of the results on newer, larger language models. Figure <a href="#fig:llm" data-reference-type="ref" data-reference="fig:llm">20</a> shows the accuracy for two such models: DeepSeek LLM 7B, a 7-billion parameter large language model (LLM) from DeepSeek (a Chinese artificial intelligence company), and Qwen3 4B, a 4-billion parameter LLM from Alibaba. Due to their size, these models were trained using 10% of the training data, with five models per threshold. As Figure <a href="#fig:llm" data-reference-type="ref" data-reference="fig:llm">20</a> indicates, the substantial conclusions remain largely the same. The exception is that, at the lowest threshold (0.6), the models’ accuracy is higher during crises than during normal times. This is possibly a consequence of the smaller training data, which makes it more challenging for models to learn to distinguish borderline sensitive from non-sensitive data.

<figure id="fig:llm" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/llm.png" />
</div>
<figcaption>LLM Performance across Training Datasets</figcaption>
</figure>

# Keywords for Each Topic from Structural Topic Model

**Topic 1:**\
Highest Probability Words: Virus, USA, China, COVID-19, Wuhan, Research, Expert\
FREX Words: Influenza, Vaccine, Seafood, Huanan, Inhibition, Paper, Coptis\
Score Words: USA, Virus, Influenza, Vaccine, Iran, WHO, SARS\
**Topic 2:**\
Highest Probability Words: Epidemic, China, Country, Government, Society, People, Control\
FREX Words: Stock Market, Democracy, Highly Pathogenic, Subtype, Culling, Socialism, Denmark\
Score Words: Humanity, China, Chinese People, World, Epidemic, Economy, Disaster\
**Topic 3:**\
Highest Probability Words: Wuhan, Lockdown, Quarantine, Really, Reason, Sadness, Many\
FREX Words: Everywhere, Brain, That Kind, Let Go, Owner, Fool, In the City\
Score Words: Lockdown, Wuhan, Animal, Really, Wild, Sadness, Quarantine\
**Topic 4:**\
Highest Probability Words: Confirmed, Cases, Pneumonia, New, Cumulative, Deaths, Novel\
FREX Words: Confirmed, Cases, New, Cumulative, Report, Contacts, -Time\
Score Words: Cases, Confirmed, New, Cumulative, Discharged, Deaths, Recovered\
**Topic 5:**\
Highest Probability Words: Pneumonia, Novel, Virus, Corona, Infection, Epidemic, News\
FREX Words: Aerosol, Paperclip, Manuscript, Sp, Fecal-Oral, Department Store\
Score Words: Corona, Novel, Virus, Pneumonia, Infection, Health, Transmission\
**Topic 6:**\
Highest Probability Words: Wuhan, Supplies, Hospital, Hubei, Donation, Personnel, Support\
FREX Words: Huoshenshan, Red Cross, Dali, Requisition, Leishenshan, Charity, Guidance Team\
Score Words: Supplies, Hospital, Donation, Medical Team, Support, Huoshenshan, Dali\
**Topic 7:**\
Highest Probability Words: Epidemic, End, Hope, Reason, Stay at Home, Video, War Against the Virus\
FREX Words: Check-in, Promise, Valentine’s Day, Exercise, Stay Home, Hotpot, Quick Click\
Score Words: Check-in, End, Go Out, Lottery, Quick Click, Promise, Happiness\
**Topic 8:**\
Highest Probability Words: Japan, China, South Korea, Hong Kong, COVID-19, Government, Epidemic\
FREX Words: Japan, South Korea, Italy, Princess, Diamond, Cruise Ship, Tokyo\
Score Words: Japan, South Korea, Diamond, Italy, Princess, Cruise Ship, Hong Kong\
**Topic 9:**\
Highest Probability Words: Patient, Hospital, Wuhan, Treatment, Sick, Zhong Nanshan, Quarantine\
FREX Words: Fever, Nucleic Acid, Fangcang, Traditional Chinese Medicine, Clinic, Recovery, Plasma\
Score Words: Patient, Hospital, Treatment, Symptoms, Fever, Sick, Zhong Nanshan\
**Topic 10:**\
Highest Probability Words: Doctor, Li Wenliang, Reason, Frontline, Candle, Nurse, Passed Away\
FREX Words: Li Wenliang, Candle, Passed Away, Elderly, Unfortunately, Daughter, Rumor\
Score Words: Doctor, Li Wenliang, Candle, Passed Away, Nurse, Elderly, Rumor\
**Topic 11:**\
Highest Probability Words: Epidemic, Map, Show, Henan, Reason, Zhejiang, Time\
FREX Words: Map, School Opening, School, Student, Hardcore, Ahh, University\
Score Words: Map, School Opening, Henan, Show, School, Jiangxi, Prison\
**Topic 12:**\
Highest Probability Words: Epidemic, Prevention and Control, Work, Personnel, Community, Residential Area, Epidemic Prevention\
FREX Words: Highway, Registration, Traffic Police, Passenger Transport, Sub-bureau, Highway, Scenic Spot\
Score Words: Prevention and Control, Police, Residential Area, Command Center, Public Security, Notice, Highway\
**Topic 13:**\
Highest Probability Words: Mask, Protection, Disinfection, Go Out, Wash Hands, Medical, Contact\
FREX Words: Alcohol, Pharmacy, Ventilation, Taiwan, Elevator, Wuhan, Cleaning\
Score Words: Mask, Disinfection, Wash Hands, Medical, Taiwan, Go Out, Alcohol\
**Topic 14:**\
Highest Probability Words: Stay Strong, Wuhan, Epidemic, China, Fight, Reason, Tribute\
FREX Words: Fist, Believe, Relay, Heroes, Song, Hello, Creation\
Score Words: Stay Strong, Tribute, Wuhan, Frontline, Medical Staff, China, Frontline\
**Topic 15:**\
Highest Probability Words: Epidemic, Resumption of Work, Enterprise, Company, Production, Prevention and Control, Impact\
FREX Words: Resumption of Work, Enterprise, Employee, Resumption of Production, Salary, Bank, Return to Post\
Score Words: Enterprise, Resumption of Work, Resumption of Production, Production, Company, Market, Employee

# AI Model’s Internal Representation of Weibo and Twitter Data

Figure <a href="#fig:pca" data-reference-type="ref" data-reference="fig:pca">21</a> provides evidence that the difference in content between Weibo and Twitter propagates to censorship AI models. It shows how a censorship AI model trained on the combined Weibo-Twitter dataset internally represents the Weibo and Twitter data. Specifically, I choose the censorship AI model that is trained on the entire Weibo and Twitter data ($`\text{threshold} = 1.0`$), so that I can obtain the internal representation of all training data. For presentational purposes, I randomly sample 2000 social media posts from the Weibo and Twitter data respectively. I then use the model to obtain the embeddings (internal representation) of the combined 4000 social media posts. As the embeddings are high-dimensional, I use t-distributed stochastic neighbor embedding (t-SNE) to reduce the dimensionality of the embeddings and plot the distributions in a two-dimensional space in Figure <a href="#fig:pca" data-reference-type="ref" data-reference="fig:pca">21</a>. Essentially, t-SNE is a statistical technique that models high-dimensional data in a two-dimensional space such that similar data points are closer to each other and dissimilar data points are further apart with high probability.

As Figure <a href="#fig:pca" data-reference-type="ref" data-reference="fig:pca">21</a> shows, the overall shapes of the two data sources are quite similar: both plots display a curved, hook-like pattern. This is to be expected as both sources are collected based on the same COVID topics using a common list of keywords. However, the distribution of the data points is quite different between the two sources. Weibo’s plot has a denser concentration of points on the upper right curve whereas Twitter’s plot has a sparser concentration around the same area but a denser concentration on the separate cluster on the left loop region.

<figure id="fig:pca" data-latex-placement="hbt!">
<div class="center">
<img src="figs_ajps/tsne.png" />
</div>
<figcaption>Model’s Internal Representation of Weibo and Twitter Data</figcaption>
</figure>

<div id="refs" class="references csl-bib-body hanging-indent">

<div id="ref-agrawal2019artificial" class="csl-entry">

Agrawal, Ajay, Joshua S Gans, and Avi Goldfarb. 2019. “Artificial Intelligence: The Ambiguous Labor Market Impact of Automating Prediction.” *Journal of Economic Perspectives* 33 (2): 31–50.

</div>

<div id="ref-allie2023facial" class="csl-entry">

Allie, Feyaad. 2023. “Facial Recognition Technology and Voter Turnout.” *The Journal of Politics* 85 (1): 328–33.

</div>

<div id="ref-beraja2023ai" class="csl-entry">

Beraja, Martin, Andrew Kao, David Y Yang, and Noam Yuchtman. 2023. “AI-Tocracy.” *The Quarterly Journal of Economics* 138 (3): 1349–402.

</div>

<div id="ref-beraja2023data" class="csl-entry">

Beraja, Martin, David Y Yang, and Noam Yuchtman. 2023. “Data-Intensive Innovation and the State: Evidence from AI Firms in China.” *The Review of Economic Studies* 90 (4): 1701–23.

</div>

<div id="ref-berinsky1999two" class="csl-entry">

Berinsky, Adam J. 1999. “The Two Faces of Public Opinion.” *American Journal of Political Science*, 1209–30.

</div>

<div id="ref-buolamwini2018gender" class="csl-entry">

Buolamwini, Joy, and Timnit Gebru. 2018. “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification.” *Conference on Fairness, Accountability and Transparency*, 77–91.

</div>

<div id="ref-chen2017authoritarian" class="csl-entry">

Chen, Jidong, and Yiqing Xu. 2017. “Why Do Authoritarian Regimes Allow Citizens to Voice Opinions Publicly?” *The Journal of Politics* 79 (3): 792–803.

</div>

<div id="ref-clark2020electra" class="csl-entry">

Clark, Kevin, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020. “Electra: Pre-Training Text Encoders as Discriminators Rather Than Generators.” *arXiv Preprint arXiv:2003.10555*.

</div>

<div id="ref-cook2019demographic" class="csl-entry">

Cook, Cynthia M, John J Howard, Yevgeniy B Sirotin, Jerry L Tipton, and Arun R Vemury. 2019. “Demographic Effects in Facial Recognition and Their Dependence on Image Acquisition: An Evaluation of Eleven Commercial Systems.” *IEEE Transactions on Biometrics, Behavior, and Identity Science* 1 (1): 32–41.

</div>

<div id="ref-cox2009authoritarian" class="csl-entry">

Cox, Gary W. 2009. “Authoritarian Elections and Leadership Succession, 1975-2004.” *APSA 2009 Toronto Meeting Paper*.

</div>

<div id="ref-crabtree2020cults" class="csl-entry">

Crabtree, Charles, Holger L Kern, and David A Siegel. 2020. “Cults of Personality, Preference Falsification, and the Dictator’s Dilemma.” *Journal of Theoretical Politics* 32 (3): 409–34.

</div>

<div id="ref-devlin2018bert" class="csl-entry">

Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. “Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding.” *arXiv Preprint arXiv:1810.04805*.

</div>

<div id="ref-diamond2019road" class="csl-entry">

Diamond, Larry. 2019. “The Road to Digital Unfreedom: The Threat of Postmodern Totalitarianism.” *Journal of Democracy* 30 (1): 20–24.

</div>

<div id="ref-egorov2009resource" class="csl-entry">

Egorov, Georgy, Sergei Guriev, and Konstantin Sonin. 2009. “Why Resource-Poor Dictators Allow Freer Media: A Theory and Evidence from Panel Data.” *American Political Science Review* 103 (4): 645–68.

</div>

<div id="ref-egorov2020political" class="csl-entry">

Egorov, Georgy, and Konstantin Sonin. 2020. *The Political Economics of Non-Democracy*. National Bureau of Economic Research.

</div>

<div id="ref-esberg2023employment" class="csl-entry">

Esberg, Jane. 2022. “Employment Restriction as Repression: Evidence from Argentina’s Film Industry.” *Working Paper*.

</div>

<div id="ref-esberg2023exile" class="csl-entry">

Esberg, Jane, and Alexandra A Siegel. 2023. “How Exile Shapes Online Opposition: Evidence from Venezuela.” *American Political Science Review* 117 (4): 1361–78.

</div>

<div id="ref-farrell2022spirals" class="csl-entry">

Farrell, Henry, Abraham Newman, and Jeremy Wallace. 2022. “Spirals of Delusion: How AI Distorts Decision-Making and Makes Dictators More Dangerous.” *Foreign Aff.* 101: 168.

</div>

<div id="ref-feldstein2019road" class="csl-entry">

Feldstein, Steven. 2019. “The Road to Digital Unfreedom: How Artificial Intelligence Is Reshaping Repression.” *Journal of Democracy* 30 (1): 40–52.

</div>

<div id="ref-fu2013assessing" class="csl-entry">

Fu, King-wa, Chung-hong Chan, and Michael Chau. 2013. “Assessing Censorship on Microblogs in China: Discriminatory Keyword Analysis and the Real-Name Registration Policy.” *IEEE Internet Computing* 17 (3): 42–50.

</div>

<div id="ref-fu2020did" class="csl-entry">

Fu, King-wa, and Yuner Zhu. 2020. “Did the World Overlook the Media’s Early Warning of COVID-19?” *Journal of Risk Research* 23 (7-8): 1047–51.

</div>

<div id="ref-gohdes2024repression" class="csl-entry">

Gohdes, Anita R. 2024. *<span class="nocase">Repression in the Digital Age: Surveillance, Censorship, and the Dynamics of State Violence</span>*. Oxford University Press. <https://doi.org/10.1093/oso/9780197743577.001.0001>.

</div>

<div id="ref-guriev2019informational" class="csl-entry">

Guriev, Sergei, and Daniel Treisman. 2019. “Informational Autocrats.” *Journal of Economic Perspectives* 33 (4): 100–127.

</div>

<div id="ref-guriev2020theory" class="csl-entry">

Guriev, Sergei, and Daniel Treisman. 2020. “A Theory of Informational Autocracy.” *Journal of Public Economics* 186: 104158.

</div>

<div id="ref-hale2022authoritarian" class="csl-entry">

Hale, Henry E. 2022. “Authoritarian Rallying as Reputational Cascade? Evidence from Putin’s Popularity Surge After Crimea.” *American Political Science Review* 116 (2): 580–94.

</div>

<div id="ref-hiruncharoenvate2015algorithmically" class="csl-entry">

Hiruncharoenvate, Chaya, Zhiyuan Lin, and Eric Gilbert. 2015. “Algorithmically Bypassing Censorship on Sina Weibo with Nondeterministic Homophone Substitutions.” *Proceedings of the International AAAI Conference on Web and Social Media* 9: 150–58.

</div>

<div id="ref-hobbs2018sudden" class="csl-entry">

Hobbs, William R, and Margaret E Roberts. 2018. “How Sudden Censorship Can Increase Access to Information.” *American Political Science Review* 112 (3): 621–36.

</div>

<div id="ref-hu2020weibo" class="csl-entry">

Hu, Yong, Heyan Huang, Anfan Chen, and Xian-Ling Mao. 2020. “Weibo-COV: A Large-Scale COVID-19 Social Media Dataset from Weibo.” *arXiv Preprint arXiv:2005.09174*.

</div>

<div id="ref-huang2015propaganda" class="csl-entry">

Huang, Haifeng. 2015. “Propaganda as Signaling.” *Comparative Politics* 47 (4): 419–44.

</div>

<div id="ref-huang2018pathology" class="csl-entry">

Huang, Haifeng. 2018. “The Pathology of Hard Propaganda.” *The Journal of Politics* 80 (3): 1034–38.

</div>

<div id="ref-imai2023experimental" class="csl-entry">

Imai, Kosuke, Zhichao Jiang, D James Greiner, Ryan Halen, and Sooahn Shin. 2023. “Experimental Evaluation of Algorithm-Assisted Human Decision-Making: Application to Pretrial Public Safety Assessment.” *Journal of the Royal Statistical Society Series A: Statistics in Society* 186 (2): 167–89.

</div>

<div id="ref-jiang2016lying" class="csl-entry">

Jiang, Junyan, and Dali L Yang. 2016. “Lying or Believing? Measuring Preference Falsification from a Political Purge in China.” *Comparative Political Studies* 49 (5): 600–634.

</div>

<div id="ref-karkkainen2019fairface" class="csl-entry">

Kärkkäinen, Kimmo, and Jungseock Joo. 2019. “Fairface: Face Attribute Dataset for Balanced Race, Gender, and Age.” *arXiv Preprint arXiv:1908.04913*.

</div>

<div id="ref-kendall2020digital" class="csl-entry">

Kendall-Taylor, Andrea, Erica Frantz, and Joseph Wright. 2020. “The Digital Dictators: How Technology Strengthens Autocracy.” *Foreign Aff.* 99: 103.

</div>

<div id="ref-king2013censorship" class="csl-entry">

King, Gary, Jennifer Pan, and Margaret E Roberts. 2013. “How Censorship in China Allows Government Criticism but Silences Collective Expression.” *American Political Science Review* 107 (2): 326–43.

</div>

<div id="ref-kreps2022all" class="csl-entry">

Kreps, Sarah, R Miles McCain, and Miles Brundage. 2022. “All the News That’s Fit to Fabricate: AI-Generated Text as a Tool of Media Misinformation.” *Journal of Experimental Political Science* 9 (1): 104–17.

</div>

<div id="ref-kuran1991now" class="csl-entry">

Kuran, Timur. 1991. “Now Out of Never: The Element of Surprise in the East European Revolution of 1989.” *World Politics* 44 (1): 7–48.

</div>

<div id="ref-kuran1997private" class="csl-entry">

Kuran, Timur. 1997. *Private Truths, Public Lies: The Social Consequences of Preference Falsification*. Harvard University Press.

</div>

<div id="ref-lee2018ai" class="csl-entry">

Lee, Kai-Fu. 2018. *AI Superpowers: China, Silicon Valley, and the New World Order*. Houghton Mifflin.

</div>

<div id="ref-leslie2020understanding" class="csl-entry">

Leslie, David. 2020. “Understanding Bias in Facial Recognition Technologies.” *arXiv Preprint arXiv:2010.07023*.

</div>

<div id="ref-lohmann1994dynamics" class="csl-entry">

Lohmann, Susanne. 1994. “The Dynamics of Informational Cascades: The Monday Demonstrations in Leipzig, East Germany, 1989–91.” *World Politics* 47 (1): 42–101.

</div>

<div id="ref-lorentzen2014china" class="csl-entry">

Lorentzen, Peter. 2014. “China’s Strategic Censorship.” *American Journal of Political Science* 58 (2): 402–14.

</div>

<div id="ref-miller2015elections" class="csl-entry">

Miller, Michael K. 2015. “Elections, Information, and Policy Responsiveness in Autocratic Regimes.” *Comparative Political Studies* 48 (6): 691–727.

</div>

<div id="ref-nicholson2023making" class="csl-entry">

Nicholson, Stephen P, and Haifeng Huang. 2023. “Making the List: Reevaluating Political Trust and Social Desirability in China.” *American Political Science Review* 117 (3): 1158–65.

</div>

<div id="ref-pan2020saudi" class="csl-entry">

Pan, Jennifer, and Alexandra A Siegel. 2020. “How Saudi Crackdowns Fail to Silence Online Dissent.” *American Political Science Review* 114 (1): 109–25.

</div>

<div id="ref-roberts2018censored" class="csl-entry">

Roberts, Margaret E. 2018. “Censored.” In *Censored*. Princeton University Press.

</div>

<div id="ref-roberts2020resilience" class="csl-entry">

Roberts, Margaret E. 2020. “Resilience to Online Censorship.” *Annual Review of Political Science* 23: 401–19.

</div>

<div id="ref-roberts2014structural" class="csl-entry">

Roberts, Margaret E, Brandon M Stewart, Dustin Tingley, et al. 2014. “Structural Topic Models for Open-Ended Survey Responses.” *American Journal of Political Science* 58 (4): 1064–82.

</div>

<div id="ref-robinson2019self" class="csl-entry">

Robinson, Darrel, and Marcus Tannenberg. 2019. “Self-Censorship of Regime Support in Authoritarian States: Evidence from List Experiments in China.” *Research & Politics* 6 (3): 2053168019856449.

</div>

<div id="ref-robinson2020face" class="csl-entry">

Robinson, Joseph P, Gennady Livitz, Yann Henon, Can Qin, Yun Fu, and Samson Timoner. 2020. “Face Recognition: Too Bias, or Not Too Bias?” *Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops*, 0–1.

</div>

<div id="ref-rozenas2010forced" class="csl-entry">

Rozenas, Arturas. 2010. “Forced Consent: Information and Power in Non-Democratic Elections.” *APSA 2010 Annual Meeting Paper*.

</div>

<div id="ref-sajith2024training" class="csl-entry">

Sajith, Aryan, and Krishna Chaitanya Rao Kathala. 2024. “Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?” *arXiv Preprint arXiv:2411.15821*.

</div>

<div id="ref-shahbaz2023repressive" class="csl-entry">

Shahbaz, Adrian, Allie Funk, Kian Vesteinsson, et al. 2023. “The Repressive Power of Artificial Intelligence.” *Freedom House*.

</div>

<div id="ref-shen2021search" class="csl-entry">

Shen, Xiaoxiao, and Rory Truex. 2021. “In Search of Self-Censorship.” *British Journal of Political Science* 51 (4): 1672–84.

</div>

<div id="ref-shih2008nauseating" class="csl-entry">

Shih, Victor Chung-Hon. 2008. “‘Nauseating’ Displays of Loyalty: Monitoring the Factional Bargain Through Ideological Campaigns in China.” *The Journal of Politics* 70 (4): 1177–92.

</div>

<div id="ref-silver2017mastering" class="csl-entry">

<span class="nocase">Silver, David, Julian Schrittwieser, Karen Simonyan, et al.</span> 2017. “Mastering the Game of Go Without Human Knowledge.” *Nature* 550 (7676): 354–59.

</div>

<div id="ref-svolik2018polarization" class="csl-entry">

Svolik, Milan. 2018. “When Polarization Trumps Civic Virtue: Partisan Conflict and the Subversion of Democracy by Incumbents.” *Available at SSRN 3243470*.

</div>

<div id="ref-svolik2019polarization" class="csl-entry">

Svolik, Milan W. 2019. “Polarization Versus Democracy.” *Journal of Democracy* 30 (3): 20–32.

</div>

<div id="ref-tanash2017decline" class="csl-entry">

Tanash, Rima S, Zhouhan Chen, Dan S Wallach, and Melissa Marschall. 2017. “The Decline of Social Media Censorship and the Rise of Self-Censorship After the 2016 Failed Turkish Coup.” *FOCI@ USENIX Security Symposium*.

</div>

<div id="ref-tannenberg2022autocratic" class="csl-entry">

Tannenberg, Marcus. 2022. “The Autocratic Bias: Self-Censorship of Regime Support.” *Democratization* 29 (4): 591–610.

</div>

<div id="ref-trinh2023statistical" class="csl-entry">

Trinh, Minh D. 2023. “Statistical Misreporting Debilitates Authoritarian Governance.” *Working Paper*.

</div>

<div id="ref-vangara2019characterizing" class="csl-entry">

<span class="nocase">Vangara, Kushal, Michael C King, Vitor Albiero, Kevin Bowyer, et al.</span> 2019. “Characterizing the Variability in Face Recognition Accuracy Relative to Race.” *Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops*, 0–0.

</div>

<div id="ref-wallace2022seeking" class="csl-entry">

Wallace, Jeremy L. 2022. *Seeking Truth and Hiding Facts: Information, Ideology, and Authoritarianism in China*. Oxford University Press.

</div>

<div id="ref-wang2021meta" class="csl-entry">

Wang, Mei, Yaobin Zhang, and Weihong Deng. 2021. “Meta Balanced Network for Fair Face Recognition.” *IEEE Transactions on Pattern Analysis and Machine Intelligence* 44 (11): 8433–48.

</div>

<div id="ref-wedeen2015ambiguities" class="csl-entry">

Wedeen, Lisa. 2015. *Ambiguities of Domination: Politics, Rhetoric, and Symbols in Contemporary Syria*. University of Chicago Press.

</div>

<div id="ref-wintrobe2000political" class="csl-entry">

Wintrobe, Ronald. 2000. *The Political Economy of Dictatorship*. Cambridge University Press.

</div>

<div id="ref-xu2021repress" class="csl-entry">

Xu, Xu. 2021. “To Repress or to Co-Opt? Authoritarian Control in the Age of Digital Surveillance.” *American Journal of Political Science* 65 (2): 309–25.

</div>

<div id="ref-xu2023unintrusive" class="csl-entry">

Xu, Xu. 2023. “The Unintrusive Nature of Digital Surveillance and Its Social Consequences.” *Working Paper*.

</div>

<div id="ref-yang2023automated" class="csl-entry">

Yang, Eddie. 2023. “Automated Repression: Ethnic Discrimination in AI-Assisted Criminal Sentencing in China.” *Working Paper*.

</div>

<div id="ref-yang2021censorship" class="csl-entry">

Yang, Eddie, and Margaret E Roberts. 2021. “Censorship of Online Encyclopedias: Implications for NLP Models.” *Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency*, 537–48.

</div>

<div id="ref-yang2023authoritarian" class="csl-entry">

Yang, Eddie, and Margaret E Roberts. 2023. “The Authoritarian Data Problem.” *Journal of Democracy* 34 (4): 141–50.

</div>

</div>

[^1]: Department of Political Science, Purdue University. A previous version of this paper was circulated under the title "The Digital Dictator’s Dilemma."

[^2]: See also Farrell et al. (2022), Wallace (2022), and Trinh (2023) for similar arguments on data quality issues in the authoritarian context.

[^3]: Even in non-strategic settings where citizens reveal their true preferences, the absence of regular information-gathering channels such as elections or a co-opted opposition can still lead to a shortage of relevant data. See a more detailed discussion of this point in the Theory section.

[^4]: The social media posts were about COVID-19. I focus on posts from the early period of the pandemic because this was when censorship of COVID-19 topics had not caught up in China. See the Data and Research Design section for more details.

[^5]: Mapping Pretrial Injustice, <https://pretrialrisk.com/national-landscape/where-are-prai-being-used/>

[^6]: Sinolytics, Table.Media, May 14, 2024, <https://table.media/en/china/sinolytics-radar/why-china-has-an-advantage-in-ai-in-medicine/>.

[^7]: See e.g., Dada Lyndell, The Insider, July 10, 2023, <https://theins.ru/en/society/263203>.

[^8]: Here I assume political sensitivity is the only relevant dimension. In reality, other dimensions such as topic, user, and platform may also influence the content generating process.

[^9]: Hiruncharoenvate et al. (2015)

[^10]: See Appendix B.5 for a more detailed discussion.

[^11]: Similar set-up has been widely used in training AI for content moderation. See e.g., Google’s Perspective API: <https://developers.perspectiveapi.com/s/about-the-api-model-cards>.

[^12]: According to one Chinese technology company, the technology to automatically censor COVID-19 topics was not put to use until around Feb. 27, 2020. See <https://bit.ly/4j1l1pr>.

[^13]: See Appendix B.3 for more details on the service and its usage by social media companies.

[^14]: I exclude tweets with labels of $`0`$ from the Twitter dataset as missingness in the Weibo datasets only comes from social media posts with positive labels.

[^15]: For this analysis, both censorable and safe posts from Twitter and Weibo are included.

[^16]: See Appendix E.1 for the list of topic keywords.

[^17]: See e.g., Jessica Batke and Laura Edelson, The China File, June 30, 2025, <https://locknet.chinafile.com/the-locknet/part-2/#humans-and-machines-censoring-together>.

[^18]: See Roberts (2020) for a survey of the debate.

[^19]: <https://huggingface.co/google-bert/bert-base-chinese>

[^20]: <https://huggingface.co/docs/transformers/index>

[^21]: <https://ai.baidu.com/tech/textcensoring>

[^22]: <http://www.nlpir.org/wordpress/2018/02/01/nlpir-500%E4%B8%87%E6%9D%A1twitter%E5%86%85%E5%AE%B9%E8%AF%AD%E6%96%99%E5%BA%93/>

[^23]: <https://figshare.com/articles/dataset/Peacock_Chinese_Twitter_Corpus_PCTC_/13489239/1>

[^24]: See e.g., <https://istories.media/stories/2023/02/08/vnutri-mashini-tsenzuri/>

[^25]: <https://static.istories.media/uploaded/documents/0b809ea16feb42c7b8c91b022e45bd6b.pdf>

[^26]: There is no change for the model trained on the full distribution of data (threhold = 1.0) as there is no missing data to begin with.
