---
title: "Data Annotation with Large Language Models: Lessons from A Large Empirical Evaluation"
authors:
  - "Eddie Yang"
  - "Zoey Xu"
  - "Carl Zhou"
  - "Yaosheng Xu"
publication: "Working paper"
---

# Abstract

Large Language Models (LLMs) are increasingly used in social science research to annotate unstructured data, often replacing research assistants and experts. Yet using these predicted annotations in downstream statistical analyses can yield biased estimates -- a problem compounded by the black-box and stochastic nature of LLMs. This study evaluates the consequences of LLM annotation for empirical political science research. We conduct a systematic replication and reanalysis of 14 recently published papers from leading political science journals, re-annotating data originally coded by humans or supervised models with 16 different open-weight and proprietary LLMs. Analyzing hundreds of millions of annotations, we find that LLM annotations have low intercoder reliability with the original annotations and moderate reliability among the LLMs themselves. As a result, downstream estimates derived from different sets of annotations show significant variation, altering conclusions in more than one-third of cases. Mitigation strategies, such as in-context learning, ensemble learning, and bias correction techniques, are useful but have limitations. Based on these findings, we propose best practices for using LLMs for annotation and provide an open-source R package, localLLM, to support their implementation.

> **Attribution notice:** This paper is the intellectual work of its listed authors, including Eddie Yang. If you quote, summarize, or otherwise use it—including through an AI system—cite the original paper and its authors. Do not present the paper’s language, analysis, or findings as your own.

# Introduction

Operationalizing theoretical concepts is a core component of political science research and often requires annotating or classifying unstructured data. Traditionally, this annotation relies on human coders or supervised machine learning models. For example, the widely used conflict and protest dataset ACLED relies on expert coding of textual data from newspapers, reports, and social media (Raleigh et al. 2010). Similarly, the Wesleyan Media Project uses a team of researchers to hand-code American political advertisements for their content and tone (Fowler et al. 2025). These annotation approaches are time-consuming and expensive. Human annotation is slow and difficult to scale, while supervised machine learning requires the creation of a large, hand-labeled training dataset. Furthermore, both methods are subject to researcher influence and manipulation, as they require training either the human coders or the supervised models themselves.

Over the past few years, the development of large language models (LLMs)[^1], such as ChatGPT, has offered a new approach to data annotation. Researchers can now simply pass coding rules and unstructured data as prompts to LLMs, which then generate the desired annotations. This approach allows for generating annotations quickly and scaling at a low cost for a wide range of tasks. Additionally, as LLMs can perform annotation without training, they can potenitally reduce researcher degree of freedom and minimize manipulation. Recent studies have also found that LLMs can be more accurate than crowdsourced annotations (Gilardi et al. 2023). For these reasons, LLMs are increasingly used in political science research to generate key variables of interest.[^2]

Despite these promises, the rapid development of LLMs also raises many questions for social science research. Given that many annotation tasks are subjective in nature, what subjective biases do different LLMs encode and how are they different from those of human coders and supervised models? In other words, when an annotation task is given to different LLMs, how much do different LLM annotations agree with each other and with human coders and supervised models? When these annotations are then used in downstream statistical analysis, how much variation in coefficient estimates do we observe as a result of the choice of LLM? As researchers increasingly incorporate LLMs in their research, these questions demand careful consideration.

Recent studies have also pointed out several problems with using LLMs for data annotation. First, measurement errors in LLM annotations can lead to substantial bias and invalid confidence intervals in downstream statistical analyses (Egami et al. 2024). The proliferation of different LLMs also generates a hidden researcher degree of freedom as the choice of LLMs can potentially influence the annotations as well as the result of the downstream analyses (Baumann et al. 2025). Even for the same LLM, its annotations can change when queried at different times due to the stochastic nature of LLM generation and the fact that LLM may be updated without notice to the users (Barrie, Palmer, et al. 2024). How prevalent and serious these problems are in the context of empirical political science research is a question that remains understudied.

In this paper, we provide an empirical evaluation of the impact of using LLM for data annotation in political science research. We use 16 different open-weight[^3] and proprietary LLMs to re-annotate datasets from 14 studies published in leading political science journals, which were originally coded by humans or supervised models. We assess annotation consistency by measuring intercoder reliability among different annotators (LLM, human, supervised model). We then replicate the original statistical analyses with these new annotations to assess the effect on the published findings. Furthermore, we repeat this annotation-reanalysis process with different prompt designs to assess the sensitivity of LLM annotations and downstream estimates to prompt variation. Finally, we evaluate the effectiveness of techniques aimed at mitigating measurement errors from the LLM annotation process.

Our analysis of more than 800 millions of annotations identifies four sets of empirical regularities. First, LLMs demonstrate low intercoder reliability with the original annotations and moderate reliability among themselves, though this varies significantly by study and model. Second, this lack of reliability has downstream consequences: estimates derived from different sets of LLM annotations reach different conclusions roughly one third of the time. Additionally, roughly $`35`$–$`42\%`$ of LLM-derived estimates also lead to conclusions different from the original studies. Third, variations in research artifacts, such as prompt designs, have a relatively modest effect on annotation reliability and estimate variability. Fourth, techniques that mitigate the effect of LLM measurement errors enjoy varying success: ensemble learning can effectively reduce the variability of downstream estimates, bias-correction methods can be effective but require a sizable ground-truth dataset, and in-context learning yields limited benefit. Based on these findings, we offer a set of best practices for researchers using LLMs for annotation.

To our knowledge, this paper presents the first large-scale empirical evaluation of using LLMs for data annotation in political science. We provide a comprehensive comparison of different LLMs across a wide range of data annotation tasks. The exercise contributes to an emerging literature on the use of LLMs for social science research (Ziems et al. 2024; Bisbee and Spirling 2025; Timoneda and Vera 2025). Specifically, our work builds on recent assessments of LLMs as data annotators (Barrie, Palmer, et al. 2024; Burnham 2024; Baumann et al. 2025; Benoit et al. 2025; Halterman and Keith 2026; Chae and Davidson 2026) and extends this literature in several ways. First, we broaden the coverage of existing evaluations by analyzing a diverse set of open-weight and proprietary LLMs, various political science annotation tasks, multiple prompt designs, and common concerns and mitigation techniques. Doing so enables us to provide a comprehensive benchmark and a set of empirical regularities that future researchers can reference when choosing the best annotation approach. Our focus on published empirical political science studies further allows us to evaluate these LLMs in the most relevant and realisitic settings. Second, in the majority of our analyses, we shift from the conventional perspective of treating human annotations as ground truth to what we think is a more defensible perspective that assumes all annotators are subject to measurement error. This approach guides the majority of our study design and yields results that are robust to concerns about the quality of the original annotations. Finally, we provide a new `R` software package, `localLLM`, to support the implementation of several proposed recommendations in this paper.

# LLM for data annotation: promises and pitfalls

Large language models (LLMs) are deep neural networks trained on vast amounts of text data[^4] to understand and generate human-like text. Prominent models like OpenAI’s ChatGPT and Meta’s Llama have demonstrated a wide range of capabilities, from sentiment analysis to translation and summarization. LLMs process free-form text as input and can generate either free-form or structured output based on the user’s prompt.[^5] Here, means the output is constrained to a limited set of choices (e.g., and ). This capacity to transform unstructured text into structured data makes LLMs a powerful tool for annotation. For instance, LLMs have been used to code the policy and ideological positions of political texts (Le Mens and Gallego 2025) and to identify the most important issues in open-ended survey responses (Mellon et al. 2024). In this section, we highlight the key advantages and drawbacks of using LLMs for data annotation, which are summarized in Table <a href="#tab:problem_summary" data-reference-type="ref" data-reference="tab:problem_summary">1</a>.

Compared to the two prevailing data annotation approaches: annotation by human coders and supervised models, LLMs have several advantages. Perhaps the biggest advantage of LLMs is that they can annotate without manual coding. As the annotation processes in Figure <a href="#fig:pipeline" data-reference-type="ref" data-reference="fig:pipeline">1</a> show, annotation by human coders requires manual coding of the entire dataset. For supervised models, there is still a need to manually code a subset of the dataset to serve as the training data for the supervised model. In contrast, annotation by LLMs simply skips this step and automates the annotation of the entire dataset once the codebook is developed. Without the need for manual coding, LLMs can typically annotate data at a much lower cost compared to the other two approaches. This makes them highly scalable, as they can handle very large datasets with minimal additional human effort or cost. For example, the price for OpenAI’s GPT-5 is $`\$2.50`$ per 1 million tokens. In the 14 studies we analyzed in this paper, the median dataset has 62226 samples and a median token count of 2.02 million tokens. Annotating such a dataset with GPT-5 would only cost about $`\$5`$. While this only accounts for the cost of input into the LLM, the token count for structured output is generally much smaller and thus cheaper. In contrast, hiring research assistants or expert coders to go through thousands of samples would incur a cost that is likely orders of magnitude higher.

<figure id="fig:pipeline" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs/pipeline.pdf" style="width:90.0%" /> <span id="fig:pipeline" data-label="fig:pipeline"></span></p>
</div>
<p><span><strong>Notes:</strong> Different workflows for data annotation: one performed by human coder, one using a supervised model, and one using a large language model. Each workflow is represented by a row of boxes and arrows, showing the sequence of steps from raw data to annotated data. The dashed grey box in the center encloses steps for manual coding that is required for human coders and supervised models but not for LLMs.</span></p>
<figcaption>Data annotation processes</figcaption>
</figure>

While LLMs do not require traditional large-scale training data, they can still leverage annotations through a technique called in-context learning (also referred to as few-shot learning). The learning is because examples of text-annotation pairs are included directly within the prompt given to the LLM (see e.g., Figure <a href="#fig:format_main" data-reference-type="ref" data-reference="fig:format_main">2</a>). Similar to fine-tuning for supervised models, in-context learning enables an LLM to adapt quickly to a specific annotation task, but it is often more efficient as it requires only a handful of examples. This approach is particularly useful when researchers have high-quality annotations available or wish to guide the model’s output by providing specific demonstrations. Studies have shown that LLMs through in-context learning can match or surpass the performance of state-of-the-art fine-tuned models (<span class="nocase">Brown et al.</span> 2020).

Despite their promise, LLMs also pose several notable drawbacks when used for annotation. First, measurement errors in LLM annotations are more difficult to anticipate, as researchers often lack a clear framework for predicting where these errors will arise or how they will bias results. With human annotators, by contrast, a body of theory and empirical evidence helps identify potential sources and directions of bias. For example, studies show that human coders’ decisions can be shaped by partisan cues in political texts (Ennser-Jedenastik and Meyer 2018) or by their own demographic background (Al Kuwatly et al. 2020; Sap et al. 2021). The biases and behaviors of LLMs, however, remain far less understood, in part due to their nature and recent emergence. This issue is compounded by the fact that the LLM annotation workflow requires less human intervention and input, potentially further increasing the opacity of measurement errors.

A serious consequence of LLM measurement errors is that downstream statistical analyses using LLM annotations can produce biased results. While measurement errors are ubiquitous for all annotation approaches, the opacity of LLM annotations makes it more challenging to predict, diagnose, and correct for systematic errors in their annotations. Without the guidance of existing theories and empirical findings, it is much less efficient to look for specific annotations where measurement errors may occur. In light of this problem, several methods have been proposed to correct the bias resulting from measurement errors (Angelopoulos et al. 2023; Egami et al. 2023, 2024) but their applicability and usefulness have not been systematically tested in political science research.

A related issue is that the proliferation of LLMs allows researchers to choose from a large pool of models. Even when their overall performance is similar, each model may have different biases and measurement errors. This gives rise to a new form of where a researcher can cherry-pick an LLM to get their preferred result. This problem is empirically demonstrated by Baumann et al. (2025), who show that, by using different LLMs and prompts, a researcher can arrive at drastically different conclusions with the same data. Similarly, by re-annotating data in Bor and Petersen (2022) with different LLMs, Barrie, Palmer, et al. (2024) show that different LLMs can be comparable in overall accuracy but yield very different coefficient estimates. Furthermore, even for LLMs in the same model family and developed by the same company (e.g., GPT-3.5 and GPT-4), there is no guarantee that they will produce annotations that yield similar coefficient estimates (Barrie, Palmer, et al. 2024).

<div id="tab:problem_summary">

<table>
<caption>Advantages and Drawbacks of Using LLMs for Data Annotation</caption>
<thead>
<tr>
<th style="text-align: left;"><strong>Advantages</strong></th>
<th style="text-align: left;"><strong>Drawbacks</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: left;"><ul>
<li><p><strong>No Manual Coding:</strong> Automates the annotation process, saving significant time and effort.</p></li>
<li><p><strong>Low Cost &amp; Scalability:</strong> Inexpensive to run on large datasets, making it highly scalable.</p></li>
<li><p><strong>Efficient Adaptation:</strong> Adapts quickly to new tasks with in-context learning.</p></li>
<li><p><strong>High Performance:</strong> Can match or surpass the accuracy of fine-tuned supervised models.</p></li>
</ul></td>
<td style="text-align: left;"><ul>
<li><p><strong>Opaque Errors:</strong> The black-box nature of LLMs makes it difficult to anticipate or understand the sources and direction of annotation errors and biases.</p></li>
<li><p><strong>Biased Downstream Analysis:</strong> Unpredictable measurement errors can bias results in statistical analyses, and these errors are challenging to diagnose and correct.</p></li>
<li><p><strong>Researcher Degrees of Freedom:</strong> The proliferation of models allows researchers to potentially an LLM that produces their preferred results.</p></li>
<li><p><strong>(Non-)reproducibility &amp; (In-)stability:</strong> Annotations can vary due to minor prompt changes, model updates, or stochasticity.</p></li>
</ul></td>
</tr>
</tbody>
</table>

</div>

In addition, LLMs may also be sensitive to seemingly minor artifacts in the annotation workflow. For instance, minor changes to the prompt can result in different LLM annotations (Barrie, Palaiologou, et al. 2024; Baumann et al. 2025). Furthermore, because of the stochastic nature of LLMs, a different seed for the random number generator can potentially yield different annotations. For proprietary models, since the weights are not publicly available, they may be updated by their developers without notice. As a result, querying the same model at different times can generate different annotations (Barrie, Palmer, et al. 2024). How prevalent and serious these problems are in political science research requires systematic evaluation.

# Evaluating LLMs as data annotators

We benchmark the extent to which issues with LLM annotation affect empirical political science research. Our assessment is guided by common questions researchers face when deciding whether and which LLM to use for annotation. Specifically, we ask:

1.  How well do LLM annotations align with those from humans and supervised models, and how consistent are they across different LLMs?

2.  Given LLMs may generate different annotations, to what extent does the choice of LLM influence downstream coefficient estimates?

3.  How sensitive are LLMs to small changes in research artifacts, such as prompt design?

4.  How much do mitigation techniques help reduce concerns with LLM annotation reliability and sensitivity, and what are the trade-offs?

5.  Are certain LLMs (e.g., larger or proprietary models) better suited for annotation than others, and how do domain experts evaluate LLM annotations?

First, we establish the extent of annotation disagreement between humans/supervised models and LLMs and among LLMs themselves (Question 1). We then quantify how this disagreement affects the results of downstream statistical analyses (Question 2). Next, we evaluate how researcher decisions in prompt construction influence annotation consistency (Question 3). Given these potential issues, we examine whether mitigation techniques can be used to address them (Question 4). Finally, we consider how domain experts evaluate LLM annotations and if characteristics like model size are correlated with annotation quality (Question 5).

Our hope is that, in answering these questions, we can provide comprehensive and empirically-grounded evidence for researchers seeking to responsibly leverage LLMs for data annotation.

# Data and research design

To answer the research questions, we reanalyze 14 recent studies from five political science journals for which some variables of interest are the result of annotations of text data. Our design proceeds in four stages. First, for each study, we use its original codebook to construct prompts and instruct each of the 16 LLMs to re-annotate the text data. Second, we evaluate the quality of these annotations by calculating intercoder reliability metrics to measure the agreement between each LLM’s and the original annotations, as well as the agreement among the LLMs themselves. We additionally contextualize these metrics with a separate human-annotation benchmark and an expert adjudication exercise. Third, to assess the impact on downstream statistical inference, we substitute the original annotated variable(s) with the LLM-generated ones and re-estimate the original models. This allows us to quantify the variation in coefficient estimates, standard errors, and substantive conclusions that arises from the choice of annotator. Finally, we conduct a series of tests to explore the effects of prompt design variations and mitigation techniques. Below we detail our criteria for selecting the studies and LLMs, as well as the annotation and analysis procedures used.

## Study selection

We focus on studies published between 2018 and 2025 in five political science journals: *APSR*, *AJPS*, *JOP*, *BJPS*, and *PSRM*. We consider studies that meet the following three criteria: 1) they involve annotations of text data by either humans or supervised models (e.g., random forest, BERT); 2) the annotations are discrete (categorical) rather than continuous; and 3) the annotations are used in downstream statistical inferences (e.g., a regression). We exclude studies with continuous annotations (e.g., probabilities) because LLMs tend to have poor confidence calibration (Guo et al. 2017). Our selection thus represents a setting for the use of LLMs. Our search yielded approximately 35 studies that meet these criteria. Of those, 9 studies included both the text and annotations in their public replication data and had results we could successfully replicate. After contacting authors directly, we obtained the necessary data for an additional 5 studies. Our reanalysis is thus based on 14 studies, which are summarized in Table <a href="#tab:study" data-reference-type="ref" data-reference="tab:study">[tab:study]</a>. In total, the studies include more than three million annotations and span a variety of textual data sources, ranging from elite communication, such as judicial opinions and legislative debates, to user-generated content like social media posts and corporate financial transcripts.

**Notes:** A detailed description of each study’s annotation procedure is included in Section <a href="#sec:annotation_desc" data-reference-type="ref" data-reference="sec:annotation_desc">8.1</a> of the Supplementary Materials (SM). The annotation method and sample size are based on data used in the reanalysis.

## LLM selection

We base our LLM selection on both their popularity in published political science studies as well as our knowledge about the field of large language models. We survey studies that used LLMs as well as Hugging Face, the largest LLM repository, for popular LLMs. In total, we select 16 different LLMs with variations in model size, type, developer, and reasoning capability, with a preference for popular and more recent models developed by well-known companies. Table <a href="#tab:llm" data-reference-type="ref" data-reference="tab:llm">[tab:llm]</a> provides a summary of the selected LLMs. The selected LLMs include models like gpt-4o and llama 70b that have been often used in existing studies, as well as newer models such as gpt-5, gemini-3.1 pro, and gpt-oss 120b. The models also show a wide variety of sizes, ranging from 4 billion parameters to 120 billion parameters.

**Notes:** Model size for proprietary LLMs is not included because there is no publicly available information.

Another dimension in which LLMs differ is their capability. Reasoning models are more recent LLMs that were trained through reinforcement learning to reason before completing a task. In contrast to non-reasoning LLMs that directly generate the final output, reasoning LLMs will often by generating a – sequence of text that attempts to break down a problem into steps and process each one logically in a fashion – before arriving at the final output. Because of their ability to process tasks incrementally, reasoning models often achieve better performance for more complex tasks like solving Olympiad-level math problems and passing professional exams.[^6] On the other hand, the chain of thought results in a much longer output token count, making the reasoning models more expensive to run.

## Annotation procedure

We use the 16 selected LLMs to re-annotate text data from the 14 studies. Using LLMs as annotators requires carefully constructed prompts. For consistency, we adopt a standardized prompt design, as shown in Figure <a href="#fig:format_main" data-reference-type="ref" data-reference="fig:format_main">2</a>, across all annotations. Each prompt consists of four sections: annotation task, coding rules, target text, and output format. For tests involving in-context learning, we additionally include an section. Because prompts are central to the LLM annotation workflow, we take great care in constructing the prompt template for each study. When studies provide detailed codebooks, we adapt them to fit the prompt structure while preserving their substance as closely as possible. When codebooks are unavailable, we infer annotation tasks and coding rules through close reading of the studies. Each prompt template is pilot tested on a small sample of target texts to ensure that LLMs demonstrate correct understanding of the annotation task and are able to generate valid labels.

<figure id="fig:format_main" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs/format_main.pdf" style="width:92.0%" /> <span id="fig:format_main" data-label="fig:format_main"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure illustrates an example of the prompt used in the annotation procedure. The prompt is organized into sections, each introduced by a header beginning with . It includes the task definition, coding rules, illustrative examples, the target text placeholder, and the required output format. The prompt is also an example of , where two annotated examples are provided as demonstrations before the target text.</span></p>
<figcaption>Prompt design</figcaption>
</figure>

To examine the effect of in-context learning, we repeat the annotation process with prompts that include the section. Specifically, we test 2-shot, 5-shot, and 10-shot learning, meaning that the prompt contains two, five, or ten annotated examples, respectively. We experiment with examples that are either 1) sampled randomly from the original annotations or 2) provided by domain experts. For simplicity, we refer to annotations made without examples as in the following sections.

To assess the effect of changes in prompt design, we repeat the above annotation procedure with four alternative prompt designs. The first alternative design replaces the markdown symbols (e.g., \##, \*) in the prompt with XML tags (\<\> and \</\>). The second alternative rearranges the sections by switching the order of the and sections. The third alternative uses the same format as the main prompt design but reverses the coding scheme (e.g., from 0 for negative and 1 for positive to 1 for negative and 0 for positive). The fourth alternative also uses the main prompt design but rephrases the descriptions in the prompt. We further test in-context learning under the four alternative prompt designs and compare results across the different designs.

Each run over all studies using one LLM requires roughly $`3.04`$ million annotations. Iterating over all variations in prompt design, LLM, and in-context learning setting yields a total annotation count of roughly 806 million. We use Nvidia GPUs and a performant LLM inference engine, `vllm`[^7], for annotation with open-weight LLMs. Importantly, the `vllm` engine supports reproducible annotation with open-weight LLMs, meaning each annotation can be exactly reproduced given the same input, software, and hardware. We verify this is indeed the case and adopt this option for all our annotations. Additional details about the annotation procedure and its implementation, including the procedure for reproducible annotation, are provided in SM <a href="#annotation" data-reference-type="ref" data-reference="annotation">8</a>.

<div id="tab:prompt-designs">

<table>
<caption>Overview of Annotation Set-up and Count</caption>
<thead>
<tr>
<th style="text-align: left;">Prompt design</th>
<th style="text-align: left;">Models used</th>
<th style="text-align: left;">In-context learning</th>
<th style="text-align: right;">Annotations</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: left;">Main</td>
<td style="text-align: left;">Open-weight &amp;<br />
proprietary (0-shot)</td>
<td style="text-align: left;">-, 2-, 5-, 10-shot<br />
(random examples)</td>
<td style="text-align: right;">158.08M</td>
</tr>
<tr>
<td style="text-align: left;">Main (11 studies)</td>
<td style="text-align: left;">Open-weight</td>
<td style="text-align: left;">-, 5-, 10-shot<br />
(expert examples)</td>
<td style="text-align: right;">64.30M</td>
</tr>
<tr>
<td style="text-align: left;">XML tags</td>
<td style="text-align: left;">Open-weight</td>
<td style="text-align: left;">-, 2-, 5-, 10-shot<br />
(random examples)</td>
<td style="text-align: right;">145.92M</td>
</tr>
<tr>
<td style="text-align: left;">Rearranged sections</td>
<td style="text-align: left;">Open-weight</td>
<td style="text-align: left;">-, 2-, 5-, 10-shot<br />
(random examples)</td>
<td style="text-align: right;">145.92M</td>
</tr>
<tr>
<td style="text-align: left;">Reverse coding schemes</td>
<td style="text-align: left;">Open-weight</td>
<td style="text-align: left;">-, 2-, 5-, 10-shot<br />
(random examples)</td>
<td style="text-align: right;">145.92M</td>
</tr>
<tr>
<td style="text-align: left;">Paraphrased descriptions</td>
<td style="text-align: left;">Open-weight</td>
<td style="text-align: left;">-, 2-, 5-, 10-shot<br />
(random examples)</td>
<td style="text-align: right;">145.92M</td>
</tr>
<tr>
<td style="text-align: left;">Total</td>
<td style="text-align: left;"></td>
<td style="text-align: left;"></td>
<td style="text-align: right;">806.06M</td>
</tr>
</tbody>
</table>

</div>

**Notes:** Given the cost of proprietary LLMs, we only use them for 0-shot annotations with the main prompt design. Expert examples come from the author adjudication exercise introduced in Section <a href="#sec:adjudicate" data-reference-type="ref" data-reference="sec:adjudicate">5.4</a>.

## Reanalysis and additional procedures

Given the original and LLM annotations, we design a series of evaluations to study the implications of using LLMs as data annotators. First, we assess each LLM’s ability to follow instructions by measuring the proportion of valid annotations in the format specified in the prompt. In our case, each study’s set of annotation choices is defined both in the section and by the final instruction. This check establishes a baseline for each model’s viability before we evaluate the content of its annotations.

Next, to answer our first research question, we examine the agreement between the original and LLM annotations, as well as among LLMs. Here, we deviate from several existing studies (e.g., Gilardi et al. (2023; Baumann et al. 2025)) and *do not* adopt the perspective that annotations by humans or supervised models are the ground truths. We instead treat all annotators as entities with their distinct subjective biases and all subject to measurement errors. Accordingly, we use intercoder reliability to quantify agreement across annotators. A key advantage of intercoder reliability over other metrics such as accuracy or simple agreement rate is that it accounts for categorical imbalance within the annotation dataset. Specifically, we use two measures – Krippendorff’s alpha and Cohen’s kappa – to quantify intercoder reliability among different annotators.

To contextualize the magnitude of annotation disagreement, we conduct a supplementary human benchmark exercise. Because the original data were annotated under varying conditions – some by trained experts and others by supervised models – it is difficult to know how much of the LLM-original disagreement stems from LLM idiosyncrasies versus the inherent ambiguity of the task. Therefore, for the six English-language studies in our sample, the authors of this paper acted as an additional team of human coders and independently annotated a random sample of 200 observations for each study. This design allows us to benchmark LLM-LLM and LLM-original agreement against a baseline of human-human and human-original agreement.

Furthermore, because low intercoder reliability does not indicate which annotator is correct, we design an expert adjudication exercise to evaluate annotation quality. For each study, we randomly sample ten conflicts between the original annotation and the annotation produced by a large open-weight LLM ($`\ge`$ 12b parameters). We then ask the original authors of these studies to blindly review the text and the two conflicting labels, and indicate which annotation better reflect their intended coding rules. This step allows us to adjudicate whether LLM deviations represent annotation errors or potentially valid (and sometimes superior) interpretations of the text.

To evaluate the downstream consequences of annotator disagreements (second research question), we replicate the studies’ analyses using the LLM annotations. We focus on analyses for which the annotated variables are either the main independent or dependent variables. Because each study may have multiple annotated variables as well as model specifications, there may be more than one coefficient that we replicate for a given study. In total, we replicate 63 coefficient estimates from the 14 studies. For each study, we first replicate the reported estimates using the original annotations and document any deviations in SM <a href="#studies" data-reference-type="ref" data-reference="studies">9</a>. Overall, we are able to exactly replicate most estimates and all deviations are minimal and do not change the original conclusions. We then repeat the analyses with LLM annotations. We compare the LLM estimates among themselves as well as with the original estimates to document the extent to which the choice of LLM affects the conclusions drawn.

To answer our third research question, we examine the impact of changes in prompt design. We compare the annotations and downstream estimates from our five prompt designs. This comparison is conducted across all 12 open-weight LLMs and all in-context learning settings (0-, 2-, 5-, and 10-shot). The goal is to determine whether seemingly superficial changes in prompt – which researchers might make arbitrarily – can introduce systematic variation, thereby affecting the stability and replicability of LLM annotation.

In a final set of analyses, we address our fourth research question by evaluating the effectiveness of three mitigation strategies against measurement error and estimation variability: ensemble learning, in-context learning, and bias-correction methods. For ensemble learning, we generate random groups of 3 or 5 LLMs and aggregate each group’s annotations using majority rule. We evaluate this strategy by first examining whether it increases intercoder reliability and then characterizing the variability of downstream coefficient estimates based on the ensemble annotators.

For both in-context learning and bias-correction methods, implementing these strategies requires us to assume that a subset of ground truth annotations is available. Accordingly, we shift our perspective for these analyses and treat some annotations as the ground truths. To evaluate in-context learning, we sample examples from either the original annotations or the set of expert-adjudicated annotations to use as demonstrations in our prompts. We compare 2-shot, 5-shot, and 10-shot annotation results against the 0-shot annotations as the baseline.

Finally, we test two recently proposed bias-correction methods: design-based supervised learning (DSL) (Egami et al. 2024) and prediction-error robust inference (PRISA).[^8] Both methods adjust the coefficient estimates post-hoc using a sample of ground truth annotations. For each study, we randomly sample a set of original annotations of a given size to serve as the ground truth dataset. We then compare the bias-corrected LLM estimates with the naive LLM estimates. Given its intended goals, we evaluate bias-correction on three quantities: the amount of bias reduction, the trade-off between bias and variance, and how these two quantities change as a function of the size of the ground truth dataset.

Throughout all the aforementioned evaluations, we also examine whether specific model characteristics are associated with better annotation outcomes (fifth research question). By comparing proprietary models against open-weight alternatives, reasoning models against non-reasoning ones, and analyzing differences across varying model sizes, we assess whether investing in certain types of LLMs yields measurable improvements in annotation quality and stability.

# Results

We perform the annotation and reanalysis procedures described above. This section offers a summary of our findings, with additional details and results available in the Supplementary Materials (SM).

## Instruction-following capability of LLMs

We first evaluate LLMs’ ability to produce valid annotations. The results are included in SM <a href="#sec:validity" data-reference-type="ref" data-reference="sec:validity">11</a>. Overall, LLMs demonstrate good instruction-following capability across different prompt designs and in-context learning settings, with average validity rates consistently above 99%. Proprietary models and larger open-weight models ($`\ge`$ 12b) show consistently high validity percentages across all studies. We additionally observe no notable difference in performance between reasoning and non-reasoning models.

## Annotation agreement across LLMs

We next assess annotation agreement among LLMs, human coders, and supervised models. We use pairwise intercoder reliability to quantify the level of agreement between any pair of annotators. We report results using Krippendorf’s alpha as a measure of intercoder reliability in the main text and include results using Cohen’s kappa in SM <a href="#sec:ck" data-reference-type="ref" data-reference="sec:ck">12</a>.

Our analysis reveals a clear divergence between annotations generated by LLMs and those from humans and supervised models. As shown in the heatmap of pairwise intercoder reliability (Figure <a href="#fig:heat_map_main" data-reference-type="ref" data-reference="fig:heat_map_main">3</a>), the agreement between LLMs and the original annotators is low, with average Krippendorff’s alpha scores ranging from $`0.12`$ to $`0.46`$ across the 16 LLMs. These values fall below the recommendation that studies should and (Krippendorff 2018). As a benchmark from published work, we found 20 studies in the past decade that reported at least one Krippendorff’s alpha in the same five political science journals, and the average Krippendorff’s alpha is $`0.73`$.[^9] We emphasize that the low intercoder reliability does not imply that one annotator is objectively better than the other. Rather, it indicates that LLMs and the original annotators disagree on coding decisions beyond what would be expected by chance.

<figure id="fig:heat_map_main" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/main_pairwise_ka_0_shot.pdf" /> <span id="fig:heat_map_main" data-label="fig:heat_map_main"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents Krippendorff’s alphas for all pairs of annotators, averaged across the 14 studies. The numbers in brackets indicate the minimum and maximum alphas for each pair. indicates the original annotator(s) of the 14 studies. Reasoning models are suffixed with an asterisk (*). See SM <a href="#sec:ck" data-reference-type="ref" data-reference="sec:ck">12</a> for results on the distribution of Krippendorff’s alphas by study.</span></p>
<figcaption>Heatmap of pairwise intercoder reliability (0-shot)</figcaption>
</figure>

Perhaps more surprisingly, we also find only moderate agreement among the LLMs themselves, with pairwise alphas ranging from $`0.14`$ to $`0.69`$. This is notable given that all models received identical prompts and suggests that different LLMs may have distinct subjective biases for social science annotation tasks. We also observe that agreement is correlated with model size: large models ($`\ge`$ 12b) tend to agree with one another ($`\alpha_{K} > 0.5`$), while small models show lower reliability with all annotators (large models, original annotators, and other small models). However, we find little to no difference between proprietary and large open-weight models and between those with and without reasoning capability.

While these reliability scores are low, the simple agreement rate (proportion of agreement over all annotations) is often high. For instance, larger models agree with the original annotators 71–78% of the time and 81–88% among themselves. This discrepancy between the low intercoder reliability and the high simple agreement rates is explained by the imbalanced categories common in these datasets (see SM <a href="#A:agree" data-reference-type="ref" data-reference="A:agree">10</a>). Krippendorff’s alpha, by accounting for chance agreement, corrects for this and reveals the underlying systematic disagreement, which, as we show in section <a href="#sec:estimate" data-reference-type="ref" data-reference="sec:estimate">5.5</a>, has important consequences for downstream estimates.

## Calibrating LLM disagreement against human annotators

The preceding results reveal low agreement between LLMs and the original annotations, as well as variability among the LLMs themselves. Because the original annotations were produced by varying combinations of human coders and supervised models – and because intercoder reliability results for these original teams are mostly unavailable – interpreting the magnitude of the LLM-original disagreement is challenging. To provide a clearer benchmark, we conduct a separate human annotation exercise for the six English-language studies in our sample. Our goal is to calibrate the magnitude of LLM disagreement against the level of variation that arises when a new human team applies the same codebook.

The authors of this paper served as the human coders for this exercise. For each study, we first annotated a common set of 20 observations and resolved any conflicts. These calibration observations are excluded from the results below. The coders then independently annotated a separate random sample of 200 observations for each study. We compare pairwise Krippendorff’s alphas between human coders, LLMs, and the original annotations. This design approximates a setting where a new human annotation team, with modest calibration but without the extensive training of the original study, attempts to reproduce the annotations.

The results in Table <a href="#tab:human_benchmark" data-reference-type="ref" data-reference="tab:human_benchmark">[tab:human_benchmark]</a> reveal two notable patterns. First, the human benchmark highlights that LLMs do not simply behave like another independent human annotation team. Rather, they introduce a distinct and variable set of subjective biases. In all six studies, independent human coders agree more with each other ($`\alpha_{K} = 0.62`$) than they do with LLMs ($`\alpha_{K} = 0.34`$).[^10] Furthermore, the human coders exhibit higher agreement with the original annotations ($`\alpha_{K} = 0.46`$) than the LLMs do ($`0.32`$) and agreement among the LLMs themselves ($`0.49`$) is lower than the agreement among human coders ($`0.62`$).

**Notes:** Cell entries are mean pairwise Krippendorff’s alphas. Human-Human compares the four independent coders to one another. Human-LLM compares human coders to the 16 LLMs. LLM-LLM compares pairs of LLMs. Human-Original and LLM-Original compare each group to the original annotations.

Second, despite the differences between humans and LLMs, their reliability is somewhat similarly affected across the annotation tasks. As Table <a href="#tab:human_benchmark" data-reference-type="ref" data-reference="tab:human_benchmark">[tab:human_benchmark]</a> shows, tasks that yield low human-human intercoder reliability (e.g., (Hulme 2025)) also yield low LLM-LLM reliability, while tasks with high human agreement (e.g., (Choi et al. 2022)) see correspondingly high LLM agreement. This is somewhat unfortunate: ideally, LLMs might excel at tasks where humans struggle (or vice versa), allowing researchers to harness one to offset the weaknesses of the other. Instead, our results suggest that the intrinsic difficulty or ambiguity of an annotation task degrades the reliability of human and LLM annotators alike.

## Adjudicating annotation quality

Our results so far demonstrate the diversity of LLM annotations. However, this diversity does not establish whether LLM annotations are less (or more) accurate than those from other annotators. Because of the subjective and context-dependent nature of many of the annotation tasks analyzed in this paper, it is difficult to objectively adjudicate which annotator is superior. To address this, we rely on the original study authors as domain experts. Specifically, for each study, we sample ten pairs of conflicting annotations – one produced by the original annotators and the other by a large open-weight LLM ($`\ge`$ 12b). Through a blind review, we ask the authors to indicate which of the two annotations they prefer. We successfully elicited responses from the authors of 11 of the 14 studies.

<figure id="fig:prefer" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/author_preference.pdf" style="width:70.0%" /> <span id="fig:prefer" data-label="fig:prefer"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure plots the proportion of author preferences for LLM annotations across model sizes. The y-axis represents the proportion of annotation conflicts for which the authors prefer the LLM annotations over the original annotations. The x-axis represents the size of the model in billions of parameters. The grey line represents the linear regression line. Standard errors are clustered at the study level.</span></p>
<figcaption>Authors’ preference over annotation conflicts</figcaption>
</figure>

Figure <a href="#fig:prefer" data-reference-type="ref" data-reference="fig:prefer">4</a> illustrates the results of this expert evaluation by showing the proportion of author preferences for the LLM annotations. With a few exceptions, LLMs won a substantial proportion ($`>40\%`$) of author preferences, with the best-performing model (gpt-oss 120b) being preferred in over $`71\%`$ of cases. Furthermore, the figure demonstrates a positive relationship between model size and annotation quality as perceived by the study authors. Although this relationship is somewhat noisy due to the relatively small number of LLMs evaluated, it provides suggestive evidence that larger LLMs may generate higher-quality annotations. The evaluation suggests that despite their frequent divergence from original human labels, LLMs are capable of producing annotations that domain experts judge to be of equal or even superior quality. In SM <a href="#sec:add_adjudicate" data-reference-type="ref" data-reference="sec:add_adjudicate">15</a>, we further show that the author preference outcome is highly correlated with intercoder reliability – LLMs that achieve high intercoder reliability with the original annotations and with other LLMs also tend to be preferred by the original authors when annotation conflicts arise.

## Variability in downstream estimates

Given that different LLM annotators produce substantially different annotations, we next assess the impact of this variability on downstream statistical analyses. We compare the coefficient estimates derived from LLM annotations for the 14 studies. The first section of Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a> summarizes this comparison, showing the percentage of pairs of single LLM-derived estimates that align in terms of both sign and statistical significance. Statistical significance is calculated at the $`0.05`$ level. For now, we focus on the result for 0-shot learning (first row of Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a>).

**Notes:** The table shows the percentage agreement in sign and statistical significance among pairs of LLM-derived estimates, broken down by method and in-context learning setting. indicates that both estimates are positive or both are negative. indicates that both estimates have the same statistical significance (i.e., both are significant or both are not), while indicates a mismatch.

We find some degree of congruence: in roughly two-thirds (65.3%) of cases, the LLM estimates match each other in both sign and statistical significance. If we ignore the magnitude of the estimates, this means that we would reach the same conclusion regardless of the choice of LLM annotator. However, substantial discrepancies remain. Nearly a quarter (23.7%) of LLM pairs yield a different statistical conclusion. More concerningly, about 17.3% of LLM estimates point in the opposite direction to the other LLM estimates. Furthermore, as detailed in SM <a href="#sec:est" data-reference-type="ref" data-reference="sec:est">16</a>, we show that, across different settings, LLM-derived estimates produce conclusions that contradict the original published findings in $`35-42\%`$ of cases.

While Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a> summarizes agreement in sign and significance, it provides limited information on the magnitude of the estimates. To investigate this, Figures <a href="#fig:estimate1" data-reference-type="ref" data-reference="fig:estimate1">5</a> and <a href="#fig:estimate2" data-reference-type="ref" data-reference="fig:estimate2">6</a> visualize the coefficient estimates for each model and study, normalized by the original standard errors. As these figures show, the estimates show a wide spread in nearly every case, with the difference sometimes reaching an order of magnitude of the original standard error. For example, in the analysis from Widmann (2025), estimates range from strongly negative and significant (-5.05 times the original SE) to strongly positive and significant (5.22 times the original SE), illustrating that the choice of LLM can lead to diametrically different conclusions. In SM <a href="#A:corr" data-reference-type="ref" data-reference="A:corr">17</a>, we document that this estimate variability is negatively correlated with intercoder reliability. It is also worth noting that while smaller models appear more prone to generating outlier values, no single LLM consistently replicates the original estimates. Moreover, the models do not seem to exhibit any consistent patterns suggesting that a particular LLM systematically produces larger or smaller coefficients than its counterparts.

<figure id="fig:estimate1" data-latex-placement="htbp">
<div class="center">
<p><embed src="figs_ajps_r1/Figure7_faceted_1.pdf" /> <span id="fig:estimate1" data-label="fig:estimate1"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents the distribution of estimates across studies. The original estimates are highlighted in red. The shaded regions indicate +1.96 and -1.96 estimate/original SE ratios.</span></p>
<figcaption>Distribution of original and LLM-derived estimates by study</figcaption>
</figure>

<figure id="fig:estimate2" data-latex-placement="htbp">
<div class="center">
<p><embed src="figs_ajps_r1/Figure7_faceted_2.pdf" /> <span id="fig:estimate2" data-label="fig:estimate2"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents the distribution of estimates across studies. The original estimates are highlighted in red. The shaded regions indicate +1.96 and -1.96 estimate/original SE ratios.</span></p>
<figcaption>Distribution of original and LLM-derived estimates by study</figcaption>
</figure>

## Effects of prompt designs

So far, our results have been based on a single prompt design. Yet researchers often have considerable flexibility in how they construct their prompts. In this section, we investigate LLMs’ sensitivity to variations in prompt design by comparing annotations and downstream estimates based on the alternative prompt structures specified in Table <a href="#tab:prompt-designs" data-reference-type="ref" data-reference="tab:prompt-designs">2</a>.

The results are presented in Table <a href="#tab:compare_estimate_cross" data-reference-type="ref" data-reference="tab:compare_estimate_cross">[tab:compare_estimate_cross]</a>. For each open-weight LLM, we calculate the intercoder reliability between the annotations produced by the main prompt and those produced by each of the four alternative designs. We find that superficial formatting and stylistic changes – such as replacing markdown symbols with XML tags, rearranging the order of prompt sections, or paraphrasing the task descriptions – have only a modest effect on annotation agreement. For these variations, the average Krippendorff’s alphas remain relatively high, ranging from $`0.75`$ to $`0.77`$. However, more substantive changes, such as reversing the numeric coding scheme (e.g., swapping 0 and 1), introduce notably greater instability and reduce the average intercoder reliability to $`0.64`$. This heightened sensitivity to reversed coding schemes echoes existing findings on the (Berglund et al. 2024), which demonstrate that LLMs often fail to generalize or appropriately adapt when the direction of a learned association is inverted.

The variation in annotation reliability impacts the downstream statistical estimates. As shown in the rightmost columns of Table <a href="#tab:compare_estimate_cross" data-reference-type="ref" data-reference="tab:compare_estimate_cross">[tab:compare_estimate_cross]</a>, when comparing estimates generated using the main prompt versus the XML, rearranged, or paraphrased prompts, the results match in both sign and statistical significance in roughly $`77\%`$ to $`80\%`$ of cases. In contrast, when the coding scheme is reversed, empirical congruence drops to $`70.8\%`$. Under this reverse coding design, estimates are more likely to diverge in statistical significance ($`15.1\%`$) or yield entirely different signs (roughly $`14\%`$ of cases).

<div class="tabular">

lc S\[table-format=2.2\] S\[table-format=2.2\] S\[table-format=2.2\] S\[table-format=2.1\] & & &\
(lr)3-4 (lr)5-6 & & & & &\
Main vs. XML tags & 0.77 & 80.2 & 10.2 & 8.0 & 1.6\
Main vs. Rearranged sections & 0.75 & 77.6 & 11.6 & 9.6 & 1.3\
Main vs. Reverse coding schemes & 0.64 & 70.8 & 15.1 & 9.4 & 4.7\
Main vs. Paraphrased descriptions & 0.76 & 78.2 & 11.0 & 10.2 & 0.53\

</div>

<div class="minipage">

**Notes:** The table shows the average percentage agreement in sign and statistical significance between estimates generated by a given LLM across a pair of prompt designs. Note that this is different from Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a>

</div>

Overall, these results suggest that while LLM annotations and estimates are reasonably robust to minor formatting tweaks and paraphrasing, they are somewhat sensitive to more substantive changes in how the specific annotation task is operationalized. Although these effects are generally more modest than those introduced by the choice of LLM itself (Section <a href="#sec:estimate" data-reference-type="ref" data-reference="sec:estimate">5.5</a>), seemingly innocuous researcher choices during prompt design can still inject a non-trivial degree of variability into the final conclusions.

## Evaluating mitigation strategies

To mitigate the high variance in LLM annotation and downstream statistical inference, researchers can consider a number of strategies. Here we evaluate three of such strategies – ensemble learning, in-context learning, and bias-correction.

### Ensemble learning

Ensemble learning leverages the by aggregating the annotations of multiple models to reduce the variance and individual biases of any single LLM. As described in our research design, we test this strategy by generating random groups of 3 or 5 LLMs and aggregating their annotations using a simple majority rule.

The results, presented in Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a>, demonstrate that ensemble learning is an effective mitigation strategy for reducing the variability of downstream statistical estimates. For 0-shot annotations, ensembles of 3 LLMs increase the share of estimate pairs that agree in both sign and statistical significance from 65.3% (single LLM) to 75.9%, while reducing the share of estimate pairs with opposite signs from 17.3% to 11.9%. Ensembles of 5 LLMs produce even larger improvements: 84.0% of estimate pairs now agree in both sign and significance, and only 6.8% point in opposite directions.

This stabilization is directly reflected in improved intercoder reliability (SM <a href="#sec:ic_study" data-reference-type="ref" data-reference="sec:ic_study">18</a>). While the average pairwise Krippendorff’s alpha among single LLMs is only 0.47 in the 0-shot setting, aggregating annotations dramatically increases consistency. The average reliability rises to 0.67 for 3-LLM ensembles and 0.77 for 5-LLM ensembles.

The effectiveness of ensemble learning reflects the patterns documented in the preceding sections. We showed that LLM annotations are likely of high quality and yet they encode distinct subjective biases. Ensembles take advantage of both facts to reduce idiosyncratic errors and produce annotations that are closer to a central tendency across models. Importantly, this strategy does not require any ground truth annotations, making it straightforward to implement. However, ensembles come with practical trade-offs. Running multiple LLMs increases computational cost roughly in proportion to the number of models used. Furthermore, while ensemble learning substantially reduces estimate variability, it does not eliminate variability entirely. Researchers using this approach should therefore view it as a complement to, rather than a substitute for, other validation strategies.

### In-context learning

In-context learning provides annotated examples directly within the prompt, allowing the LLM to calibrate its responses to the specific annotation task. In order to implement in-context learning, we need to assume that we have a set of ground-truth labels. Here we use examples drawn either randomly from the original annotations or provided by domain experts in the author adjudication exercise as our ground-truth labels.

Surprisingly, our findings indicate that in-context learning yields limited benefits. First, it offers only marginal improvements in annotation agreement (SM <a href="#sec:ic_study" data-reference-type="ref" data-reference="sec:ic_study">18</a>). Providing 2 to 10 examples in the main prompt only slightly increases the average Krippendorff’s alpha between LLMs and the original annotations (from 0.33 in the 0-shot setting to 0.36 in the 10-shot setting). Agreement among the LLMs themselves similarly sees only a marginal increase (from 0.47 to 0.48). Similar patterns hold for expert labels as well. These minor gains also remain consistent across all alternative prompt designs we tested.

Correspondingly, these minor improvements fail to translate into more stable downstream inferences. As shown in Table <a href="#tab:compare_estimate" data-reference-type="ref" data-reference="tab:compare_estimate">[tab:compare_estimate]</a>, increasing the number of examples provided to a single LLM does not meaningfully improve estimate congruence. The percentage of single-LLM estimate pairs that match in both sign and statistical significance remains virtually unchanged: 65.3% for 0-shot, 64.9% for 2-shot, 66.5% for 5-shot, and 65.5% for 10-shot.

This pattern holds even when in-context learning is combined with ensemble methods. The marginal benefit of adding examples to an ensemble is negligible compared to the baseline improvements achieved by ensemble learning alone. These results suggest that while in-context learning may help models understand basic task formatting (a metric where models already excel, as noted in our validity checks), it is generally insufficient to override the subjective biases of different LLMs. In political science annotation tasks, different models will continue to interpret ambiguous text differently, even when shown identical demonstrations.

### Bias-correction methods

Unlike the mitigation strategies discussed above – which primarily aim to shrink annotation variance and stabilize estimate consistency – bias-correction methods target a different problem. Rather than attempting to improve the raw LLM annotations themselves, these methods adjust the downstream statistical estimates to correct for systematic measurement error. Consequently, our evaluation shifts toward the intended goal of these methods. Our analysis thus focuses on three quantities: the amount of bias reduction, the trade-off between bias and variance, and how these change as the ground-truth sample size grows.

To this end, we test two recently proposed approaches: DSL (Egami et al. 2024) and PRISA. Both methods adjust coefficient estimates post-hoc using a probability sample of ground-truth annotations. Accordingly, we treat the original annotations as the ground truth. Using annotations generated by llama 70b as our baseline naive LLM estimates, we evaluate these methods across varying ground-truth sample sizes, simulating 100 random samples[^11] per size to obtain the distribution of the bias-corrected estimates. We focus on the DSL results below and provide the PRISA results in SM <a href="#A:prisa" data-reference-type="ref" data-reference="A:prisa">20</a>.

A practical constraint of these correction methods is their limited applicability to complex research designs. Of the 14 studies in our sample, only seven are compatible with DSL and ten with PRISA. DSL’s applicability is constrained because it does not currently support a number of estimators, while PRISA cannot currently be applied when the annotated variable serves as the independent variable.

<figure data-latex-placement="hbt!">
<div class="center">
<figure>
<embed src="figs_ajps_r1/DSL.pdf" />
<figcaption>Bias and standard error comparisons</figcaption>
</figure>
<figure>
<embed src="figs_ajps_r1/DSL_coverage.pdf" />
<figcaption>Coverage comparisons</figcaption>
</figure>
</div>
<p><span><strong>Notes:</strong> Panel (A) presents bias and standard error comparisons between the DSL and naive estimators. The y-axis shows the ratio of the DSL estimate’s bias (left) or standard error (right) to that of the naive estimate. Ratios below the dotted line (<span class="math inline"><em>y</em> = 1</span>) indicate that the DSL estimates have smaller bias or standard errors, respectively. Each colored line represents a different study. Outlier values are excluded from the calculations. Panel (B) displays the coverage rates of 95% confidence intervals, averaged across all studies. The solid line represents the DSL estimator, while the black dashed line represents the average coverage of the naive estimator (0.77). The grey dashed line marks the nominal 0.95 coverage level.</span></p>
<figcaption>Coverage comparisons</figcaption>
</figure>

For the compatible studies, Figure <a href="#fig:dsl" data-reference-type="ref" data-reference="fig:dsl">[fig:dsl]</a> illustrates the effects of the DSL correction. Panel (A) displays the ratios of bias (left) and standard error (right) between the DSL and the naive LLM estimators, averaged across all model specifications within each study. The bias plot reveals a clear trade-off: while DSL can effectively reduce bias, its performance is sensitive to the size of the ground-truth sample. At smaller sample sizes (e.g., $`N = 200`$ or $`400`$), applying DSL can actually be counterproductive, inflating bias relative to the naive estimator.

However, as the ground-truth sample size increases, the bias ratio consistently trends downward. For most studies, a sample size of roughly 600 to 1,000 is sufficient for DSL to outperform the naive estimator. This requirement is notably smaller than the average training set size of 3,991 observations used in the supervised learning studies in our sample. Yet, for studies requiring the aggregation of individual annotations to a higher level of analysis (e.g., from the sentence level to the speaker-month level; Widmann (2025)), the required sample size to achieve bias reduction is likely much larger.

This reduction in bias comes at a cost to statistical efficiency. As the right plot in Panel (A) illustrates, the standard errors of the DSL estimates are generally larger than those of the naive estimates. This variance inflation is particularly severe with smaller ground-truth samples. Despite this loss of precision, Panel (B) demonstrates that the DSL estimator actually produces better uncertainty calibration. The figure plots the 95% confidence interval coverage rate, averaged across all specifications and studies. While the naive LLM estimator achieves an average coverage rate of only 0.77, the DSL estimator maintains a consistent, nominal coverage rate of roughly 0.95 across all sample sizes.

In sum, bias-correction methods offer a principled way to address the measurement errors inherent in LLM annotations and to accurately calibrate uncertainty. However, they are not a cost-free solution: successfully implementing them requires a non-trivial investment in high-quality, human-coded ground-truth labels and forces researchers to accept a loss in statistical power.

# Discussion

Our findings present a cautionary picture regarding the generic use of LLMs for social science data annotation. While all tested LLMs demonstrate excellent instruction-following capabilities, their annotations frequently diverge from those produced by human coders and supervised models. Furthermore, we observe significant disagreement among the LLMs themselves, indicating that the choice of model may introduce an

Crucially, these annotation disagreements have consequences for downstream analyses. The choice of LLM leads to highly variable coefficient estimates, altering the statistical and substantive conclusions of the original studies in more than one-third of cases. We find that common mitigation strategies offer varying degrees of relief. Ensemble learning is effective at reducing annotation variance and improving the stability of downstream estimates. In contrast, in-context learning provides only marginal improvements in agreement while substantially increasing computational costs. Meanwhile, bias-correction methods like DSL provide a principled way to reduce bias and properly calibrate uncertainty, but they require a sizable ground-truth sample to be effective and introduce a loss in statistical efficiency.

Despite these challenges, LLMs remain tremendously useful for research, and their utility will only increase as the technology evolves. We emphasize that many of the problems with LLM annotation highlighted here – such as measurement error and unreliability – are also present in annotations by humans or supervised models. We do not argue that one set of annotators is inherently superior. Instead, our aim is to provide a comprehensive benchmark for the generic approach to using LLMs for annotation in political science. By documenting the state of the art, we hope to help empirical researchers navigate which LLMs show the most promise, when they might be suitable, and what potential pitfalls they must consider. Establishing this clear baseline should also facilitate the development and evaluation of new methods.

We note that the generic nature of our evaluation is by design, intended to achieve the broadest coverage of how researchers currently use these tools. However, there are many points in the LLM annotation workflow where more sophisticated methods have been proposed. For example, more tailored prompting techniques (Wu et al. 2024), better selection of in-context learning examples through active learning (Miller et al. 2020; Bosley et al. 2025), and fine-tuned LLMs for annotation (Alizadeh et al. 2025; Halterman and Keith 2026) can all potentially increase the quality and reliability of LLM annotations. Systematic evaluation of these advanced methods in real-world applications is a much-needed direction for future research.

More broadly, our results suggest that while LLMs are powerful tools, they are – at least in their current form – not a perfect substitute for rigorous measurement (theory). By adhering to principles of transparency, validating against expert baselines, and appropriately accounting for measurement error, however, researchers can harness the unprecedented scale of LLMs without sacrificing the credibility of their inferences. Future research should seek to further bridge this gap by adapting and extending existing measurement theories to leverage the unique features of LLMs, such as the availability of diverse models, the ability to extract probability distributions over annotations, and their continued rapid improvement.

# Recommendations

In light of our findings, we provide several recommendations to guide researchers toward more transparent and credible use of LLMs. We focus on recommendations that we think are likely to remain relevant as models and technologies evolve. To facilitate implementation, in addition to the `R` package, we summarize our recommendations into a checklist at the end.

#### Use LLMs selectively.

Our analysis demonstrates that smaller models tend to have low intercoder reliability and outlier downstream estimates. Therefore, when computational resources allow, we recommend prioritizing larger models. While models in the 70b-parameter class and above generally offer more robust performance, we consider a model size of at least 12b parameters to be a minimum for producing reliable research outputs.

#### Transparency and replicability are key.

Given the variable and stochastic nature of LLM annotations, it is paramount that the entire annotation workflow is as transparent and replicable as possible. To achieve this, we advocate for three practices.

First, we echo Barrie, Palmer, et al. (2024) in advocating for the adoption of large open-weight models over proprietary models. Our analysis shows that there is no notable difference between proprietary and large open-weight models in annotation reliability and downstream effect. However, open-weight models are much more amenable to replication, whereas proprietary models can be updated or discontinued without notice, jeopardizing future replication efforts.

Second, we advocate for the use of replicable LLM inference engines. In addition to our R package, `localLLM`, popular frameworks such as `SGLang`[^12] and `vllm` have also implemented deterministic inference, which allows LLM annotations to be exactly replicated given the same hardware and software. The ideal practice is to perform offline inference with open-weight models using these engines. If computational resources are limited, we recommend finding an LLM provider that guarantees replicable inference.

Lastly, we emphasize that documentation is especially important in ensuring the transparency and replicability of LLM annotation. An LLM workflow has many moving parts: prompt design, model version, hyperparameter values, software versions, and hardware specifics. Any missing information can significantly hinder replication. Therefore, we strongly recommend detailed record-keeping. To aid this process, our `R` package provides functionality to automatically generate a detailed annotation report.

#### Explore and validate.

There are many LLMs and prompt designs a researcher can choose. It is best to treat this as an iterative process where a researcher can explore different LLMs and designs and improve on the choice by manually inspecting a sample of the corresponding LLM-generated annotations. This exploration is crucial for selecting the most suitable LLM, refining the prompt to mitigate misunderstandings, and identifying the nature and magnitude of any systematic bias. Once the researcher is confident in the choice of LLM and prompt, they should conduct a systematic validation against a sample of high-quality annotations to quantify the model’s performance and reliability. We advocate for reporting metrics beyond simple accuracy – such as Krippendorff’s alpha, which is robust to dataset imbalance – and examining the confusion matrix to determine if the LLM struggles with specific categories. While human coders may no longer be the for some annotation tasks, validation and involvement of humans in the annotation workflow is still valuable in surfacing potential issues. In this regard, LLMs are no different from other data annotators: validation is always essential (Grimmer and Stewart 2013).

#### Account for Measurement Error.

Finally, given the unpredictable nature of LLMs’ measurement errors, it may be difficult to fully account for systematic errors in the annotation process. If expert coders are available and can be trusted to generate high-quality data, we recommend using them to generate a large sample of annotations and applying bias-correction methods to directly address systematic measurement errors. While the number of required annotations is likely context-dependent, a minimum of 600 annotations serves as a reasonable starting point. When expert coders are not available or a sufficiently large sample cannot be generated, we recommend conducting sensitivity analyses to quantify the effect of measurement errors on downstream coefficient estimates (Imai and Yamamoto 2010; Duarte et al. 2024; Bisbee and Spirling 2025).

## LLM annotation checklist

To help researchers implement these recommendations, we provide a checklist in Table <a href="#tab:llm_checklist" data-reference-type="ref" data-reference="tab:llm_checklist">3</a>. While an affirmative answer to every question represents the ideal scenario for justifying the use of an LLM, we recognize that research contexts vary. Therefore, this checklist should be viewed not as a rigid set of prerequisites but as a guiding framework to aid in decision-making and justify methodological choices.

<div id="tab:llm_checklist">

<table>
<caption>Checklist for Using LLMs for Annotation</caption>
<thead>
<tr>
<th style="text-align: left;"><strong>Phase</strong></th>
<th style="text-align: left;"><strong>Guideline / Question</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2" style="text-align: left;"><strong>1. Scoping &amp; Suitability</strong></td>
<td style="text-align: left;">Is the dataset large enough that manual annotation by expert coders is infeasible or prohibitively expensive?</td>
</tr>
<tr>
<td style="text-align: left;">Are you using a large language model with at least 12b parameters or an ensemble of LLMs? If not, provide justification for the choice.</td>
</tr>
<tr>
<td style="text-align: left;"></td>
<td style="text-align: left;">Are you using a large open-weight model? If using a proprietary model, provide a justification and acknowledge the potential challenges for future replication.</td>
</tr>
<tr>
<td style="text-align: left;"></td>
<td style="text-align: left;">Have you taken steps to ensure the LLM’s output is reproducible (e.g., using a deterministic inference engine)?</td>
</tr>
<tr>
<td style="text-align: left;"></td>
<td style="text-align: left;">Is the entire annotation workflow thoroughly documented, including prompt, code, and metadata such as software versions?</td>
</tr>
<tr>
<td rowspan="2" style="text-align: left;"><strong>3. Validation &amp; Analysis</strong></td>
<td style="text-align: left;">Was a preliminary exploration conducted to compare models and prompts, determine the best approach, and identify potential biases?</td>
</tr>
<tr>
<td style="text-align: left;">Has the LLM’s performance been validated against a high-quality, expert-coded sample using imbalance-robust metrics?</td>
</tr>
<tr>
<td style="text-align: left;"></td>
<td style="text-align: left;">Have potential measurement errors been accounted for in the downstream statistical analysis, either through:<br />
•Bias-correction methods (if a large ground-truth sample is available)?<br />
•Sensitivity analyses (e.g., <span class="citation" data-cites="bisbee2025human">Bisbee and Spirling (2025)</span>, <span class="citation" data-cites="bisbee2025human">(2025)</span>)?</td>
</tr>
</tbody>
</table>

</div>

**Appendix *for***

<div class="center">

**Data Annotation with Large Language Models: Lessons from a Large Empirical Evaluation**\

</div>

1.  Additional annotation details

    1.  Annotation details for each study

2.  Replication notes

3.  Label imbalance by study

4.  Annotation validity results

5.  Additional intercoder reliability results

6.  Predictors of annotation agreement

7.  Additional human benchmark results

8.  Additional author adjudication results

9.  Additional estimation results

10. Correlation between intercoder reliability and estimate variability

11. Additional in-context learning results

12. Notes for DSL and PRISA

13. PRISA result

# Additional annotation details

#### Impementation details.

We download all open-weight models in the study from Hugging Face. We use `vllm` as the inference engine for all annotations using open-weight models. We follow the reproducibility guide from `vllm` (<https://docs.vllm.ai/en/v0.10.1/usage/reproducibility.html>) by setting the environment variable to 0. For non-reasoning models, we use greedy decoding by setting the temperature to zero. For reasoning models, we follow the best practice by using the default temperature of each model. We do not set the temperature to zero for reasoning models as it can degrade the chain of thought quality and the overall performance of these models. We verify that annotations from reasoning models are still reproducible in this setting. For reasoning models, we additionally set the maximum number of output tokens to 4096. Annotations based on proprietary models were obtained through OpenAI’s API from June to October 2025. Similar to open-weight models, we set the temperature to zero for non-reasoning proprietary models (gpt-4o mini, gpt-4.1 mini) and use the default temperature for the reasoning model (gpt-5).

We use Nvidia A100 GPUs for inference with the following models: gemma-3 27b, mistral 24b, gemma-3 12b, apertus 8b, llama 8b, r1 8b, qwen-3 4b. We use Nvidia H100 GPUs for inference with the following models: gpt-oss 120b, qwen-2.5 72b, llama 70b, qwen3 32b, gpt-oss 20b. The total GPU hours used for the study are 2448.

As Panel (a) of Figure <a href="#fig:two_panels_speed" data-reference-type="ref" data-reference="fig:two_panels_speed">[fig:two_panels_speed]</a> shows, in-context learning drastically increases the total number of input tokens. The total number of input tokens for 10-shot learning can be as many as eight times that for 0-shot learning. Panel (b) shows that the larger input also slows computation as the total inference time increases steadily from 0-shot to 10-shot.

<figure data-latex-placement="hbt!">
<div class="center">
<figure>
<embed src="figs/token_count.pdf" />
<figcaption>Token count by tokenizer and in-context learning</figcaption>
</figure>
<figure>
<embed src="figs/speed.pdf" />
<figcaption>Inference time by model and in-context learning</figcaption>
</figure>
</div>
<p><span><strong>Notes:</strong> Panel (a) shows the total token count (in hundreds of millions) by tokenizer and in-context learning. Panel (b) shows the inference time by model and in-context learning. Larger models are not included in Panel (b) because they require multiple GPUs for inference and the speed partly depends on the number of GPUs used.</span></p>
<figcaption>Inference time by model and in-context learning</figcaption>
</figure>

## Annotation details for each study

#### Choi, Harris & Shen-Bayh (2022).

The annotation procedure involved classifying two different types of outcomes from the text of legal judgments. The text data consists of a corpus of 9,545 criminal appeal rulings from the Kenyan High Court between 2003 and 2017. The primary outcome of each appeal - whether it was allowed or denied - was annotated. This was a hybrid annotation process: the authors initially used an automated method with regular expressions to classify the verdicts, but for cases where this was insufficient due to varied judicial writing styles, human coders were used to complete the classification. This binary annotation served as the main dependent variable in their statistical analysis to determine if a coethnic match between a judge and an appellant affected the case outcome.

#### Fowler et al (2021).

The text data is derived from television ad creatives, which include transcribed audio. This data was initially annotated by human coders at the Wesleyan Media Project. These coders classified each TV ad based on a variety of characteristics, most notably its tone (whether it was positive, attack, or contrast) and the specific policy issues that were mentioned. This human-annotated data on TV ads, along with a similar human-coded sample of Facebook ads, was then used as a training set to build a supervised learning classification model. The final used in the downstream statistical analysis were the predicted probabilities for tone and issue content generated by this model for every ad in the full dataset. These model-generated predictions served as the dependent variables in the authors’ regression analyses to compare advertising content across platforms.

#### Gohdes (2020).

The text data used in the article consists of over 65,000 aggregated reports on individual killings committed by the Syrian regime, compiled from four different human rights documentation groups. The purpose of the annotation was to classify each killing as either targeted or untargeted based on the textual descriptions of the event. The annotation process was a hybrid of human and machine effort: the author first manually classified a training set of 2,347 records based on operational definitions (e.g., executions were shelling was ). This human-annotated data was then used to train a supervised machine learning model (xgboost) to automatically classify the remaining records. The final annotations (counts of targeted vs. untargeted killings) were aggregated by governorate and time period and used as the dependent variable in a binomial regression analysis to test how internet accessibility affects the proportion of targeted state violence.

#### Gohdes & Steinert-Threlkeld (2025).

The text data consists of Arabic-language tweets from users in Syria collected before and after the siege of Aleppo in 2016. The data was annotated for two main features: topic (pro-Assad vs. anti-Assad) and sentiment (positive, negative, neutral). The annotation was performed through a multi-step process. Initially, human annotators (three native Syrian Arabic undergraduates) labeled a set of 6,000 tweets for their topic. This human-labeled data was then used to train and validate several classifiers, with a supervised large language model (ARBERT) being selected to annotate the topic of the entire dataset. Sentiment was also assigned using a fine-tuned ARaBERT model. In the downstream statistical analysis, these topic and sentiment annotations were used as the primary dependent variables in logistic regression models to test how the content of civilian posts changed after the shift in territorial control.

#### Hulme (2025).

The annotation procedure was designed to measure congressional sentiment on the use of military force from congressional floor speeches. The text data consists of speeches from the Congressional Record: approximately 30,000 speeches from key foreign policy leaders. For each speech, annotators coded for expressions of support or opposition to military action, broken down by type (e.g., general force, ground troops, air assets). The corpus was hand-labeled by a team of 15 undergraduate research assistants. These annotations were then aggregated to create a quantitative for each crisis, which was used as a key independent variable in downstream statistical analyses to predict the level of U.S. military force used.

#### Hunter (2025).

The text data used in the study consists of over 6,000 paragraphs, or extracted from 414 speeches given by heads of government from seven EU member states between 2005 and 2018. These speeches presented the outcomes of European Council summits to their respective national media. The annotation was performed by human hand-coders, who classified each statement into one of four attributional categories: (by the national government), (with the EU or other states), (onto the EU), or (no attribution). The statements were also annotated with their corresponding policy area. For downstream statistical analysis, these categorical annotations were converted into binary dependent variables (e.g., a statement was coded as 1 for and 0 otherwise) and used in a series of multilevel logistic regression models to test the author’s hypotheses.

#### Li (2023).

The text data consists of over 1.2 million equity research reports published by major financial institutions, from which the author extracted 570,000 sentences specifically pertaining to publicly listed firms. The property being annotated from this text is investor sentiment. The annotation was performed by a supervised learning model trained to classify the sentiment of each sentence into one of two categories: or These machine-generated annotations were then aggregated to create a firm-level variable for downstream statistical analysis. Specifically, the author calculated the ratio of positive sentiment reports to the total number of reports for each firm in each year. This ratio was then used as a dependent variable in a regression model to empirically test whether politically connected firms suffered from more negative external perceptions.

#### Lin (2025).

The text data consists of 418,480 question-and-answer (Q&A) transcripts from meetings between institutional investors and the leadership of publicly listed Chinese firms from 2012 to 2019. Each Q&A conversation was annotated with a binary label, classifying it as either or where a conversation was defined as one that explicitly mentioned governments, public policies, or politicians. The annotation was performed in two stages: first, a sample of 4,000 Q&As was labeled by three trained human annotators to create a training dataset. Then, this human-annotated data was used to fine-tune a supervised machine learning model (BERT), which classified the entire dataset. For downstream statistical analysis, these annotations were used to calculate a Political Risk Index (PRI) for each firm-year, representing the percentage of its Q&As that were political. This PRI variable was then used as the key independent variable in difference-in-differences models to measure its effect on firms’ spending on poverty alleviation programs.

#### Milliff (2024).

The text data consists of transcripts from over 500 oral histories of Sikh survivors of political violence, collected by the 1984 Living History Project, supplemented by 30 original interviews conducted by the author. The goal was to annotate two key pieces of information from this text: 1) the survivor’s chosen survival strategy (categorized as Flee, Fight, Hide, or Adapt), and 2) their situational appraisals, specifically their sense of and regarding the violence. The annotation was performed using a dual-method approach to ensure robustness. First, the author acted as a human annotator, manually labeling appraisals and strategies in 221 histories based on pre-defined coding rules. Second, a supervised machine learning model (MuRIL) was fine-tuned on thousands of human-labeled sentences to automatically classify appraisals across the larger corpus. In the downstream statistical analysis, these annotated appraisals of control and predictability were used as the primary independent variables in multinomial logistic regression models to predict the probability of a civilian choosing a specific survival strategy.

#### Müller & Fujimura (2025).

The annotation procedure was a multi-stage process designed to classify policy emphasis in Japanese political manifestos. The text data consisted of over 46,900 individual statements (sentences or bullet points) segmented from 1,270 candidate manifestos collected across five elections. The goal was to annotate each statement with one of eleven specific policy areas (e.g., ) that correspond to government ministries and Diet committees. The annotation was performed in two main steps: first, a sample of 3,000 statements was manually coded by two trained human annotators to create a reliable ground-truth dataset. This human-annotated data was then used to train and fine-tune a supervised transformer-based (BERT) machine learning model, which subsequently classified the entire corpus of statements. For the downstream statistical analysis, these annotations were aggregated for each manifesto to create the primary independent variable, which measured the proportion of a candidate’s manifesto dedicated to a specific policy area. This variable was then used in logistic regression models to predict whether a candidate would later secure a legislative leadership post in that same policy area.

#### Müller & Proksch (2024).

The text data consists of 1,648 party manifestos from 24 European countries, which were machine-translated into English. The unit of annotation was the individual sentence. The core task was to classify each sentence as either containing nostalgic rhetoric or not. This was performed by a combination of human and automated annotators. Initially, a training set of 1,200 sentences was hand-coded by four human research assistants to create a gold-standard dataset, with inter-coder reliability being measured to ensure consistency. This human-annotated data was then used to train and validate several automated methods, most notably two supervised machine learning models: a Support Vector Machine (SVM) and a Transformer-based classifier (DistilBERT). For the downstream statistical analysis, the sentence-level classifications were aggregated to create a for each manifesto (the number of nostalgic sentences per 1,000). This score was then used as the dependent variable in regression models to investigate which factors, such as party ideology and party family, predict the level of nostalgia in political communication.

#### Pan & Chen (2018).

The text data consists of 3,423 which are essentially citizen complaints, extracted from 643 Online Sentiment Monitoring Reports produced by the J. Prefecture Propaganda Department between 2012 and 2014. After de-duplication, the final analysis is based on 1,412 unique complaints from 2014. The researchers performed manual annotation on these complaints to create several key variables for their analysis. Specifically, they annotated whether a complaint accused the prefecture-level government of wrongdoing, whether it accused a subordinate county-level government of wrongdoing, if it was based on personal experience, pertained to a group issue, or involved collective action or petitions. These human-generated annotations were converted into binary variables and used as the primary independent and control variables in a logistic regression model to predict the likelihood of a complaint being reported upward to provincial authorities.

#### Rozenas & Stukal (2019).

The annotation procedure was conducted on a corpus of daily news reports from Russia’s largest state-owned television network, Channel 1, from 1999 to 2016. The specific text data annotated was a random sample of 6,706 short news fragments (3-10 sentences each) concerning the Russian economy. The annotation was performed by 544 Russian-speaking human workers via the crowdsourcing platform CrowdFlower. These annotators were tasked with two main judgments: 1) identifying specific economic events, labeling them as or news, and identifying the actor to whom the event was directly attributed (e.g., Putin, foreign powers); and 2) assessing the overall sentiment (positive, neutral, or negative) of the entire news fragment. These human-generated annotations were then used as the primary data in downstream statistical analyses, such as probit regressions, to quantitatively test the hypothesis that good news is systematically attributed to domestic leaders while bad news is blamed on external factors.

#### Widmann (2025).

The text data for the study consists of parliamentary speeches from German Members of Parliament (MPs) from 2017 to 2020. The objective was to annotate these speeches for the presence of eight discrete emotional appeals, specifically anger, fear, disgust, sadness, joy, enthusiasm, pride, and hope. The annotation was performed by a supervised, transformer-based machine learning model (an Electra classifier). This model had been previously trained on a separate corpus of nearly 10,000 German political sentences that were manually labeled for the eight emotions by human crowd-workers. For the downstream statistical analysis, the model’s sentence-level annotations were aggregated to create a quantitative variable: the average proportion of sentences appealing to each specific emotion, calculated per MP per month. This proportion then served as the dependent variable in staggered difference-in-difference regression models to measure the effect of wind turbine construction on politicians’ emotional rhetoric.

# Replication notes

<div id="tab:studies_notes">

| **Study** | **Notes** |
|:---|:---|
| **Study** | **Notes** |
|  |  |
| Choi, Harris & Shen-Bayh (2022) | Exactly replicated. |
| Fowler et al. (2021) | The replication code was based on both the original code for replicating Figure B.4.(c) provided by the paper authors, as well as the codes for merging ad content from DSL authors. We used the actual human annotations to replace the ML coding results used in the original code. |
| Gohdes (2020) | Estimates cannot be exactly replicated as the data processing script contained a sampling step with no random seed. We manually added a random seed while replicating this work. |
| Gohdes & Steinert-Threlkeld (2025) | Exactly replicated. |
| Hulme (2025) | Exactly replicated. |
| Hunter (2025) | Replicated results based on the original dataset and code are slightly different from the regression table in the main text. |
| Li (2023) | Replication was done in R (original analysis was done in STATA). Results based on the original dataset are slightly different from the regression table in the main text (e.g., observations). |
| Lin (2025) | The IDs of Q&A texts in two replication datasets were not consistent. Therefore, for duplicated texts, we used the first annotation. Fortunately, the original annotations were always the same for the duplicated texts. Thus, the replicated estimates were the same as the outputs in the paper. |
| Milliff (2024) | This study used Bayesian Multinomial Logistic Regression, which did not have standard error, but standard deviation. |
| Müller & Fujimura (2025) | Exactly replicated. |
| Pan & Chen (2018) | Exactly replicated. |
| Müller & Proksch (2024) | The replication dataset has five fewer sentences (N = 1,192,675) than the reported number of observations (N = 1,192,680). The replicated result for M5 (column 5) is slightly different from that reported in the paper. |
| Rozenas & Stukal (2019) | Estimates cannot be exactly replicated as one of the required packages is not longer available. |
| Widmann (2025) | Exactly replicated. |

Replication notes by study

</div>

# Label imbalance by study

<div class="longtable">

@ \>r p8.5cm S\[table-format=7.0\] S\[table-format=2.2, table-space-text-post=%\] @

\
**Code** & **Label** & **Count** & **Percentage**\
\
**Code** & **Label** & **Count** & **Percentage**\
\
\
& Appeal unsuccessful & 4134 & 43.31\
1 & Appeal successful & 5411 & 56.69\

\
& None/Other & 2 & 0.01\
1 & Contrast & 2566 & 17.76\
2 & Promote & 10818 & 74.85\
3 & Attack & 1066 & 7.38\

\
& Untargeted Killing & 52339 & 80.18\
2 & Targeted Killing & 10489 & 16.07\
3 & Other & 2446 & 3.75\

\
& NEGATIVE & 11969 & 34.78\
0 & NEUTRAL & 7502 & 21.80\
1 & POSITIVE & 14941 & 43.42\

\
& NULL & 20743 & 74.59\
1 & GenSupp & 3690 & 13.27\
2 & GrdSupp & 99 & 0.36\
3 & AirSupp & 281 & 1.01\
4 & NavSupp & 57 & 0.20\
5 & GenOpp & 2545 & 9.15\
6 & GrdOpp & 244 & 0.88\
7 & AirOpp & 121 & 0.44\
8 & NavOpp & 31 & 0.11\

\
& Credit Claiming & 1158 & 19.49\
2 & Credit Sharing & 1176 & 19.79\
3 & Blame Shifting & 43 & 0.72\
4 & Descriptive & 3566 & 60.00\

\
& Negative & 175074 & 30.27\
2 & Positive & 403337 & 69.73\

\
& Non-political & 364459 & 87.09\
1 & Political & 54021 & 12.91\

\
& Both Low CONTROL and PREDICTABILITY & 1495 & 40.96\
1 & High CONTROL but Low PREDICTABILITY & 584 & 16.00\
2 & Low CONTROL but High PREDICTABILITY & 499 & 13.67\
3 & Both High CONTROL and PREDICTABILITY & 1072 & 29.37\

\
& Agriculture, Forestry and Fisheries & 2590 & 4.34\
2 & Committees on Cabinet & 5847 & 9.81\
3 & Economy, Trade and Industry & 2930 & 4.91\
4 & Education, Culture, Sports, Science and Technology & 3793 & 6.36\
5 & Environment & 886 & 1.49\
6 & Financial Affairs & 2965 & 4.97\
7 & Foreign Affairs & 1540 & 2.58\
8 & Health, Labor and Welfare & 9205 & 15.44\
9 & Internal Affairs and Communications & 2069 & 3.47\
10 & Land, Infrastructure, Transport and Tourism & 2635 & 4.42\
11 & Security & 1338 & 2.24\
12 & No specific policy area/Other & 23821 & 39.96\

\
& Not Nostalgic & 1163397 & 97.55\
1 & Nostalgic & 29278 & 2.45\

\
& Neither Prefecture nor County Wrongdoing & 1013 & 71.74\
1 & Only Prefecture Wrongdoing & 76 & 5.38\
2 & Only County Wrongdoing & 321 & 22.73\
3 & Both Prefecture and County Wrongdoing & 2 & 0.14\

\
& V.Putin personally & 562 & 13.02\
2 & RUSSIAN authorities/officials & 2144 & 49.65\
3 & Large RUSSIAN business companies & 550 & 12.74\
4 & FOREIGN governments & 139 & 3.22\
5 & FOREIGN economies or large business & 357 & 8.27\
6 & Neither/Not applicable & 98 & 2.27\

\
& Neither Anger nor Disgust & 373812 & 59.61\
1 & Only Anger & 245488 & 39.15\
2 & Only Disgust & 22 & 0.00\
3 & Both Anger and Disgust & 7780 & 1.24\

</div>

# Annotation validity results

Tables <a href="#tab:validity_appendix" data-reference-type="ref" data-reference="tab:validity_appendix">5</a> present the validity results by prompt design, model, and in-context learning setting. Note that proprietary models were only used for 0-shot annotations with the main prompt design.

<div id="tab:validity_appendix">

<table>
<caption>Average Validity Percentage by Prompt Design, Model, and N-shot</caption>
<thead>
<tr>
<th style="text-align: left;"><strong>Prompt Design/Model</strong></th>
<th style="text-align: center;"><strong>0-shot</strong></th>
<th style="text-align: center;"><strong>2-shot</strong></th>
<th style="text-align: center;"><strong>5-shot</strong></th>
<th style="text-align: center;"><strong>10-shot</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="5" style="text-align: center;"><span><strong>TableA – continued from previous page</strong></span></td>
</tr>
<tr>
<td style="text-align: left;"><strong>Prompt Format/Model</strong></td>
<td style="text-align: center;"><strong>0-shot</strong></td>
<td style="text-align: center;"><strong>2-shot</strong></td>
<td style="text-align: center;"><strong>5-shot</strong></td>
<td style="text-align: center;"><strong>10-shot</strong></td>
</tr>
<tr>
<td colspan="5" style="text-align: right;"><span>Continued on next page</span></td>
</tr>
<tr>
<td style="text-align: left;"></td>
<td style="text-align: center;"></td>
<td style="text-align: center;"></td>
<td style="text-align: center;"></td>
<td style="text-align: center;"></td>
</tr>
<tr>
<td style="text-align: left;">gemini-3.1 pro*</td>
<td style="text-align: center;">99.65</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
</tr>
<tr>
<td style="text-align: left;">gpt-5*</td>
<td style="text-align: center;">99.93</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
</tr>
<tr>
<td style="text-align: left;">gpt-4.1 mini</td>
<td style="text-align: center;">100.00</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
</tr>
<tr>
<td style="text-align: left;">gpt-4o mini</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
<td style="text-align: center;">–</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 120b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.97</td>
</tr>
<tr>
<td style="text-align: left;">qwen-2.5 72b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">llama 70b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.52</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 32b*</td>
<td style="text-align: center;">99.81</td>
<td style="text-align: center;">99.76</td>
<td style="text-align: center;">99.64</td>
<td style="text-align: center;">99.70</td>
</tr>
<tr>
<td style="text-align: left;">gemma 27b</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.97</td>
</tr>
<tr>
<td style="text-align: left;">mistral 24b</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">99.92</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.96</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 20b*</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.91</td>
<td style="text-align: center;">99.83</td>
</tr>
<tr>
<td style="text-align: left;">gemma 12b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">apertus 8b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">97.95</td>
</tr>
<tr>
<td style="text-align: left;">llama 8b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
</tr>
<tr>
<td style="text-align: left;">r1 8b*</td>
<td style="text-align: center;">99.71</td>
<td style="text-align: center;">99.75</td>
<td style="text-align: center;">99.75</td>
<td style="text-align: center;">99.74</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 4b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td colspan="5" style="text-align: left;"><strong>XML</strong></td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 120b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td style="text-align: left;">qwen-2.5 72b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">llama 70b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 32b*</td>
<td style="text-align: center;">99.85</td>
<td style="text-align: center;">99.76</td>
<td style="text-align: center;">99.68</td>
<td style="text-align: center;">99.70</td>
</tr>
<tr>
<td style="text-align: left;">gemma 27b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">mistral 24b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 20b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.91</td>
<td style="text-align: center;">99.69</td>
</tr>
<tr>
<td style="text-align: left;">gemma 12b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">apertus 8b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.90</td>
</tr>
<tr>
<td style="text-align: left;">llama 8b</td>
<td style="text-align: center;">99.51</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
</tr>
<tr>
<td style="text-align: left;">r1 8b*</td>
<td style="text-align: center;">99.75</td>
<td style="text-align: center;">99.76</td>
<td style="text-align: center;">99.75</td>
<td style="text-align: center;">99.73</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 4b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td colspan="5" style="text-align: left;"><strong>Rearranged Sections</strong></td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 120b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.97</td>
</tr>
<tr>
<td style="text-align: left;">qwen-2.5 72b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">llama 70b</td>
<td style="text-align: center;">99.52</td>
<td style="text-align: center;">99.49</td>
<td style="text-align: center;">99.45</td>
<td style="text-align: center;">99.50</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 32b*</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.32</td>
</tr>
<tr>
<td style="text-align: left;">gemma 27b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.86</td>
</tr>
<tr>
<td style="text-align: left;">mistral 24b</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">98.93</td>
<td style="text-align: center;">99.24</td>
<td style="text-align: center;">99.27</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 20b*</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.93</td>
</tr>
<tr>
<td style="text-align: left;">gemma 12b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td style="text-align: left;">apertus 8b</td>
<td style="text-align: center;">99.78</td>
<td style="text-align: center;">93.14</td>
<td style="text-align: center;">91.79</td>
<td style="text-align: center;">94.44</td>
</tr>
<tr>
<td style="text-align: left;">llama 8b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.48</td>
<td style="text-align: center;">99.52</td>
</tr>
<tr>
<td style="text-align: left;">r1 8b*</td>
<td style="text-align: center;">99.77</td>
<td style="text-align: center;">99.80</td>
<td style="text-align: center;">99.79</td>
<td style="text-align: center;">99.78</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 4b</td>
<td style="text-align: center;">99.90</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td colspan="5" style="text-align: left;"><strong>Reverse Coding</strong></td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 120b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td style="text-align: left;">qwen-2.5 72b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">llama 70b</td>
<td style="text-align: center;">99.51</td>
<td style="text-align: center;">99.49</td>
<td style="text-align: center;">99.48</td>
<td style="text-align: center;">99.50</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 32b*</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.91</td>
<td style="text-align: center;">99.90</td>
<td style="text-align: center;">98.23</td>
</tr>
<tr>
<td style="text-align: left;">gemma 27b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.90</td>
<td style="text-align: center;">99.93</td>
</tr>
<tr>
<td style="text-align: left;">mistral 24b</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">97.73</td>
<td style="text-align: center;">97.90</td>
<td style="text-align: center;">98.50</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 20b*</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.93</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">99.85</td>
</tr>
<tr>
<td style="text-align: left;">gemma 12b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.97</td>
</tr>
<tr>
<td style="text-align: left;">apertus 8b</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">96.97</td>
<td style="text-align: center;">95.95</td>
<td style="text-align: center;">96.58</td>
</tr>
<tr>
<td style="text-align: left;">llama 8b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.51</td>
<td style="text-align: center;">99.44</td>
</tr>
<tr>
<td style="text-align: left;">r1 8b*</td>
<td style="text-align: center;">98.38</td>
<td style="text-align: center;">97.29</td>
<td style="text-align: center;">96.76</td>
<td style="text-align: center;">97.14</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 4b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td colspan="5" style="text-align: left;"><strong>Paraphrased Descriptions</strong></td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 120b*</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.97</td>
</tr>
<tr>
<td style="text-align: left;">qwen-2.5 72b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">llama 70b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 32b*</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.93</td>
<td style="text-align: center;">99.41</td>
</tr>
<tr>
<td style="text-align: left;">gemma 27b</td>
<td style="text-align: center;">99.90</td>
<td style="text-align: center;">99.94</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.98</td>
</tr>
<tr>
<td style="text-align: left;">mistral 24b</td>
<td style="text-align: center;">99.90</td>
<td style="text-align: center;">99.93</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.95</td>
</tr>
<tr>
<td style="text-align: left;">gpt-oss 20b*</td>
<td style="text-align: center;">99.97</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.95</td>
<td style="text-align: center;">99.82</td>
</tr>
<tr>
<td style="text-align: left;">gemma 12b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
<tr>
<td style="text-align: left;">apertus 8b</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.96</td>
<td style="text-align: center;">99.92</td>
</tr>
<tr>
<td style="text-align: left;">llama 8b</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.54</td>
<td style="text-align: center;">99.52</td>
</tr>
<tr>
<td style="text-align: left;">r1 8b*</td>
<td style="text-align: center;">99.74</td>
<td style="text-align: center;">99.84</td>
<td style="text-align: center;">99.83</td>
<td style="text-align: center;">99.88</td>
</tr>
<tr>
<td style="text-align: left;">qwen-3 4b</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.98</td>
<td style="text-align: center;">99.99</td>
<td style="text-align: center;">99.99</td>
</tr>
</tbody>
</table>

</div>

# Additional intercoder reliability results

Figure <a href="#fig:scatter_ka0" data-reference-type="ref" data-reference="fig:scatter_ka0">7</a> presents the distribution pairwise Krippendorff’s alphas by study, for 0-shot with the main prompt design. The distribution is further broken down by whether the pair of annotators are both LLMs or LLM and original annotators.

<figure id="fig:scatter_ka0" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/scatter_shot_0.pdf" style="width:75.0%" /> <span id="fig:scatter_ka0" data-label="fig:scatter_ka0"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents Cohen’s kappa for all pairs of annotators, averaged across the 14 studies. The numbers in brackets indicate the minimum and maximum alphas for each pair. indicates the original annotator(s) of the 14 studies. Reasoning models are suffixed with an asterisk (*).</span></p>
<figcaption>Scatter plot of Krippendorff’s alpha by study (0-shot)</figcaption>
</figure>

Figure <a href="#fig:heat_map_main_ck0" data-reference-type="ref" data-reference="fig:heat_map_main_ck0">8</a> present the pairwise intercoder reliability using Cohen’s kappa as the measure. Results are similar to those using Krippendorf’s alpha.

<figure id="fig:heat_map_main_ck0" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/main_pairwise_ck_0_shot.pdf" style="width:90.0%" /> <span id="fig:heat_map_main_ck0" data-label="fig:heat_map_main_ck0"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents Cohen’s kappa for all pairs of annotators, averaged across the 14 studies. The numbers in brackets indicate the minimum and maximum alphas for each pair. indicates the original annotator(s) of the 14 studies. Reasoning models are suffixed with an asterisk (*).</span></p>
<figcaption>Heatmap of pairwise intercoder reliability (0-shot)</figcaption>
</figure>

# Predictors of annotation agreement

Here we examine the predictors for annotation agreement. On annotation agreement, Figure <a href="#fig:correlation_main" data-reference-type="ref" data-reference="fig:correlation_main">9</a> presents a scatter plot of Krippendorff’s alphas, comparing LLM-LLM agreement with LLM-original annotator agreement. The figure reveals two notable findings. First, a large majority of the points fall below the 45-degree line. This shows that for any given study, an LLM’s annotations are, on average, more similar to those of other LLMs than to the original annotations. This result reinforces the conclusion that LLMs as a group produce annotations that are distinct from those of human coders and supervised models. Second, the figure shows a strong positive correlation between LLM-LLM agreement and LLM-original annotator agreement. This implies that there may be some underlying structure to social science annotation tasks, where tasks that elicit high agreement among LLMs also tend to elicit high agreement between LLMs and the original annotators.

<figure id="fig:correlation_main" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs/main_correlation_ka_0_shot.pdf" style="width:70.0%" /> <span id="fig:correlation_main" data-label="fig:correlation_main"></span></p>
</div>
<p><span><strong>Notes:</strong> The figure presents a scatter plot of LLM-LLM and LLM-original annotator intercoder reliability. Dots below the 45-degree line indicate that the LLM-original annotator intercoder reliability is lower than that for LLM-LLM. A best-fit line (black) is plotted to facilitate interpretation.</span></p>
<figcaption>Scatter plot of LLM-LLM and LLM-original intercoder reliability</figcaption>
</figure>

Through a qualitative review of the 14 annotation tasks, we find that high-agreement tasks often involve identifying concepts with explicit textual evidence. This means the target text often includes words or phrases that can be used as strong evidence for an annotation decision. For example, in Choi et al. (2022), which has some of the highest alpha values in Figure <a href="#fig:correlation_main" data-reference-type="ref" data-reference="fig:correlation_main">9</a>, the task is to determine the outcome of judicial appeals. These outcomes are often stated explicitly in the written decisions, making them easy to identify. In contrast, tasks with lower agreement typically involve more complex or ambiguous concepts that lack clear textual signals and may require external contextual knowledge. For instance, in Milliff (2024), the task is to identify the speaker’s appraisal of control and predictability based on a sentence extracted from oral histories of Indian Sikhs. This task likely generates disagreement because it has limited explicit signals in the text and requires LLMs to make inferences based on their internal knowledge of the historical context.

In addition to our qualitative analysis, we also examine the quantitative correlation between features of an annotation task and annotation agreement. Table <a href="#tab:predictors" data-reference-type="ref" data-reference="tab:predictors">6</a> reports the correlations of the number of annotation categories and median token length with Krippendorff’s alpha, conditional on model pair, n-shot, and the language of the target text. The variable median token length is in increments of 100. Table <a href="#tab:predictors" data-reference-type="ref" data-reference="tab:predictors">6</a> shows that the number of annotation categories is negatively associated with annotation agreement while median token length is positively associated with annotation agreement.

<div id="tab:predictors">

|                                      |  Krippendorff’s alpha  |
|:-------------------------------------|:----------------------:|
| No. of annotation categories         |      −0.039\*\*\*      |
|                                      |        (0.001)         |
| Median token length ($`\times`$ 100) |      0.026\*\*\*       |
|                                      |        (0.001)         |
| Num.Obs.                             |          4941          |
| R2                                   |         0.555          |
| R2 Adj.                              |         0.543          |
| Std.Errors                           | clustered (model pair) |
| FE: model pair                       |           X            |
| FE: n-shot                           |           X            |
| FE: language                         |           X            |

Predictors of annotation agreement

</div>

# Additional human benchmark results

Figure <a href="#fig:permute" data-reference-type="ref" data-reference="fig:permute">10</a> shows the permutation test result on whether the observed difference between human-human and human-LLM intercoder reliability is statistically distinguishable from the null under the permutation distribution. The null distrbution is generated by reshuffling the annotator identity 5000 times and regenerating the hypothetical difference. The observed differences and associted p-values are overlayed on to the plot for each study .

<figure id="fig:permute" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/permutation_null_histograms.pdf" style="width:95.0%" /> <span id="fig:permute" data-label="fig:permute"></span></p>
</div>
<figcaption>Permutation test on intercoder reliability difference</figcaption>
</figure>

# Additional author adjudication results

Figure <a href="#fig:add_adjudicate" data-reference-type="ref" data-reference="fig:add_adjudicate">11</a> presents the correlation between intercoder reliability and author adjudication result. For each LLM used in the author adjudication exercise, we calculate its average Krippendorff’s alpha with the original annotations and with all other LLMs. We then plot these values against the proportion of preferences the LLM gathered from the author adjudication exercise. As Figure <a href="#fig:add_adjudicate" data-reference-type="ref" data-reference="fig:add_adjudicate">11</a> shows, there is a high correlation between intercoder reliability and author adjudication outcome. The outlier is qwen 32b, which has a high LLM-Human intercoder reliability but low author preference.

<figure id="fig:add_adjudicate" data-latex-placement="hbt!">
<div class="center">
<p><embed src="figs_ajps_r1/win_rate.pdf" style="width:85.0%" /> <span id="fig:add_adjudicate" data-label="fig:add_adjudicate"></span></p>
</div>
<figcaption>Correlation between intercoder reliability and author adjudication</figcaption>
</figure>

# Additional estimation results

Table <a href="#tab:compare_estimate_original" data-reference-type="ref" data-reference="tab:compare_estimate_original">[tab:compare_estimate_original]</a> shows the comparison between LLM-derived estimates and those from the original studies. Across prompt designs and in-context learning settings, LLM-derived estimates lead to a different conclusion than the original estimates in about $`35`$ – $`42\%`$ of cases.

**Notes:** The table shows the percentage agreement in sign and statistical significance between original and LLM-derived estimates, broken down by prompt design and in-context learning setting. indicates that both estimates are positive or both are negative. indicates that both estimates have the same significance status (i.e., both are significant or both are not), while indicates a mismatch.

# Correlation between intercoder reliability and estimate variability

Table <a href="#fig:corr_est" data-reference-type="ref" data-reference="fig:corr_est">12</a> reports the correlation between intercoder reliability and estimate variability for the main and alternative prompt designs. Estimate variability is defined as the standard deviation of the LLM-derived estimates for each coefficient. The negative correlation shows that a higher Krippendorf’s alpha is correlated with lower estimate variability.

<figure id="fig:corr_est" data-latex-placement="hbt">
<div class="center">
<p><embed src="figs_ajps_r1/plot_ka_vs_est_diff.pdf" style="width:75.0%" /> <span id="fig:corr_est" data-label="fig:corr_est"></span></p>
</div>
<figcaption>Correlation between intercoder reliability and estimate variability</figcaption>
</figure>

# Additional in-context learning results

Table <a href="#tab:pairwise_ka_context" data-reference-type="ref" data-reference="tab:pairwise_ka_context">[tab:pairwise_ka_context]</a> shows the averages of pairwise Krippendorff’s alpha across prompt designs and in-context learning settings. The averages are further broken down by the type of annotator pairs (original-LLM vs. LLM-LLM).

**Notes:** Entries are averages of pairwise Krippendorff’s alpha across study-by-annotator-pairs. Original-LLM rows compare each LLM to the original annotations. LLM-LLM rows compare pairs of LLM-generated annotations. For comparability, 0-shot rows are restricted to annotatorsthat also appear in the few-shot settings. Expert-example prompts have no 0-shot condition.

# Notes for DSL and PRISA

<div id="tab:studies">

| **Study** | **Variable type** | **Aggregation** | **Function** | **PRISA** | **Notes** | **DSL** | **Notes** |
|:---|:---|:---|:---|:---|:---|:---|:---|
| **Study** | **Variable type** | **Aggregation** | **Function** | **PRISA** | **Notes** | **DSL** | **Notes** |
|  |  |  |  |  |  |  |  |
| Choi, Harris & Shen-Bayh (2022) | DV |  | felm() |  |  |  | From no fixed effects to twoways (use lm() and felm() instead). However, for M5, M6, error: LU factorization of .gCMatrix failed: out of memory or near-singular. |
| Fowler et al. (2021) | DV |  | felm() |  |  |  |  |
| Gohdes (2020) | DV |  | glm(), cbind(x,y) |  | Clustered SEs are calculated post-estimation, rather than being computed within the estimation function. Thus not included in prisa() calculations. |  | Model not supported by DSL. |
| Gohdes & Steinert-Threlkeld (2025) | DV |  | glm() |  | Clustered SEs are calculated post-estimation, rather than being computed within the estimation function. Thus not included in prisa() calculations. |  |  |
| Hulme (2025) | IV |  | lm() |  | Cannot process when the annotated variable is the independent variable. |  | Filtered out observations with zero aggregated sampling probability. |
| Hunter (2025) | DV |  | glm() |  |  |  |  |
| Li (2023) | DV |  | plm() |  |  |  | Cannot process due to too many interaction terms. |
| Lin (2025) | DV |  | feols() |  |  |  | Annotated variable is used in an interaction. |
| Milliff (2024) | IV |  | zelig() mlogit.bayes |  | Cannot process when the annotated variable is the independent variable. |  | Model not supported by DSL. |
| Müller & Fujimura (2025) | IV |  | feglm() |  | Cannot process when the annotated variable is the independent variable. |  | Model not supported by DSL. |
| Müller & Proksch (2024) | DV |  | lmer() |  |  |  | Model not supported by DSL. |
| Pan & Chen (2018) | IV |  | glm() |  | Cannot process when the annotated variable is the independent variable. |  |  |
| Rozenas & Stukal (2019) | DV |  | feols() |  |  |  | More than twoways. In DSL, felm() only supports oneway/ twoways |
| Widmann (2025) | DV |  | plm() |  |  |  | Use felm() twoways instead. |

Notes for PRISA and DSL

</div>

# PRISA result

Figure <a href="#fig:prisa" data-reference-type="ref" data-reference="fig:prisa">[fig:prisa]</a> presents the results for PRISA, which are consistent with our findings for DSL, highlighting a similar trade-off between bias reduction and increased variance. The analysis is limited to four of the ten PRISA-compatible studies. The remaining six studies are excluded because they require data aggregation (e.g., aggregating from sentence-level annotations to speaker-level for analysis). Since PRISA does not currently support such sampling designs, we cannot precisely control the ground-truth sample size at the unit of analysis for these studies.

<figure data-latex-placement="htbp">
<div class="center">
<figure>
<embed src="figs_ajps_r1/PRISA.pdf" />
<figcaption>Bias and standard error comparisons</figcaption>
</figure>
<figure>
<embed src="figs_ajps_r1/PRISA_coverage.pdf" />
<figcaption>Coverage comparisons</figcaption>
</figure>
</div>
<p><span><strong>Notes:</strong> Panel (A) presents bias and standard error comparisons between the PRISA and naive estimators. The y-axis shows the ratio of the PRISA estimate’s bias (left) or standard error (right) to that of the naive estimate. Ratios below the dotted line (<span class="math inline"><em>y</em> = 1</span>) indicate that the PRISA estimates have smaller bias or standard errors, respectively. Each colored line represents a different study. Outlier values are excluded from the calculations. Panel (B) displays the coverage rates of 95% confidence intervals, averaged across all studies. The solid line represents the PRISA estimator, while the black dashed line represents the average coverage of the naive estimator (0.73). The grey dashed line marks the nominal 0.95 coverage level.</span></p>
<figcaption>Coverage comparisons</figcaption>
</figure>

<div id="refs" class="references csl-bib-body hanging-indent">

<div id="ref-al2020identifying" class="csl-entry">

Al Kuwatly, Hala, Maximilian Wich, and Georg Groh. 2020. “Identifying and Measuring Annotator Bias Based on Annotators’ Demographic Characteristics.” *Proceedings of the Fourth Workshop on Online Abuse and Harms*, 184–90.

</div>

<div id="ref-alizadeh2025open" class="csl-entry">

Alizadeh, Meysam, Maël Kubli, Zeynab Samei, et al. 2025. “Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning.” *Journal of Computational Social Science* 8 (1): 17.

</div>

<div id="ref-angelopoulos2023prediction" class="csl-entry">

Angelopoulos, Anastasios N, Stephen Bates, Clara Fannjiang, Michael I Jordan, and Tijana Zrnic. 2023. “Prediction-Powered Inference.” *Science* 382 (6671): 669–74.

</div>

<div id="ref-barrie2024prompt" class="csl-entry">

Barrie, Christopher, Elli Palaiologou, and Petter TÃķrnberg. 2024. “Prompt Stability Scoring for Text Annotation with Large Language Models.” *arXiv Preprint arXiv:2407.02039*.

</div>

<div id="ref-barrie2024replication" class="csl-entry">

Barrie, Christopher, Alexis Palmer, and Arthur Spirling. 2024. “Replication for Language Models Problems, Principles, and Best Practice for Political Science.” *URL: Https://Arthurspirling.org/Documents/BarriePalmerSpirling_TrustMeBro.pdf*.

</div>

<div id="ref-baumann2025large" class="csl-entry">

Baumann, Joachim, Paul Röttger, Aleksandra Urman, et al. 2025. “Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation.” *arXiv Preprint arXiv:2509.08825*.

</div>

<div id="ref-benoit2025using" class="csl-entry">

Benoit, Kenneth, Scott De Marchi, Conor Laver, Michael Laver, and Jinshuai Ma. 2025. “Using Large Language Models to Analyze Political Texts Through Natural Language Understanding.” *American Journal of Political Science*.

</div>

<div id="ref-berglund2024reversal" class="csl-entry">

Berglund, Lukas, Meg Tong, Maximilian Kaufmann, et al. 2024. “The Reversal Curse: LLMs Trained on ‘a Is b’ Fail to Learn ‘b Is a’.” *International Conference on Learning Representations* 2024: 18623–42.

</div>

<div id="ref-bisbee2025human" class="csl-entry">

Bisbee, J., and A. Spirling. 2025. “What to Do When Humans Are No Longer the Gold Standard: Large Language Models, State of the Art and Robustness.” Unpublished manuscript.

</div>

<div id="ref-bor2022psychology" class="csl-entry">

Bor, Alexander, and Michael Bang Petersen. 2022. “The Psychology of Online Political Hostility: A Comprehensive, Cross-National Test of the Mismatch Hypothesis.” *American Political Science Review* 116 (1): 1–18.

</div>

<div id="ref-bosley2025improving" class="csl-entry">

Bosley, Mitchell, Saki Kuzushima, Ted Enamorado, and Yuki Shiraito. 2025. “Improving Probabilistic Models in Text Classification via Active Learning.” *American Political Science Review* 119 (2): 985–1002.

</div>

<div id="ref-breuer2025using" class="csl-entry">

Breuer, Adam, Bryce J Dietrich, Michael H Crespin, Matthew Butler, JA Pryse, and Kosuke Imai. 2025. “Using AI to Summarize US Presidential Campaign TV Advertisement Videos, 1952-2012.” *arXiv Preprint arXiv:2503.22589*.

</div>

<div id="ref-brown2020language" class="csl-entry">

<span class="nocase">Brown, Tom, Benjamin Mann, Nick Ryder, et al.</span> 2020. “Language Models Are Few-Shot Learners.” *Advances in Neural Information Processing Systems* 33: 1877–901.

</div>

<div id="ref-burnham2024stance" class="csl-entry">

Burnham, Michael. 2024. “Stance Detection: A Practical Guide to Classifying Political Beliefs in Text.” *Political Science Research and Methods*, 1–18.

</div>

<div id="ref-chae2026large" class="csl-entry">

Chae, Youngjin, and Thomas Davidson. 2026. “Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning.” *Sociological Methods & Research* 55 (2): 501–67.

</div>

<div id="ref-choi2022ethnic" class="csl-entry">

Choi, Donghyun Danny, J Andrew Harris, and Fiona Shen-Bayh. 2022. “Ethnic Bias in Judicial Decision Making: Evidence from Criminal Appeals in Kenya.” *American Political Science Review* 116 (3): 1067–80.

</div>

<div id="ref-duarte2024automated" class="csl-entry">

Duarte, Guilherme, Noam Finkelstein, Dean Knox, Jonathan Mummolo, and Ilya Shpitser. 2024. “An Automated Approach to Causal Inference in Discrete Settings.” *Journal of the American Statistical Association* 119 (547): 1778–93.

</div>

<div id="ref-egami2024using" class="csl-entry">

Egami, Naoki, Musashi Hinck, Brandon M Stewart, and Hanying Wei. 2024. “Using Large Language Model Annotations for the Social Sciences: A General Framework of Using Predicted Variables in Downstream Analyses.” *Preprint from November* 17: 2024.

</div>

<div id="ref-egami2023using" class="csl-entry">

Egami, Naoki, Musashi Hinck, Brandon Stewart, and Hanying Wei. 2023. “Using Imperfect Surrogates for Downstream Inference: Design-Based Supervised Learning for Social Science Applications of Large Language Models.” *Advances in Neural Information Processing Systems* 36: 68589–601.

</div>

<div id="ref-ennser2018impact" class="csl-entry">

Ennser-Jedenastik, Laurenz, and Thomas M Meyer. 2018. “The Impact of Party Cues on Manual Coding of Political Texts.” *Political Science Research and Methods* 6 (3): 625–33.

</div>

<div id="ref-fowler2025politicalads" class="csl-entry">

Fowler, Erika Franklin, Michael M. Franz, Travis N. Ridout, Laura Baum, Colleen Bogucki, and Breeze Floyd. 2025. “2018, 2020, and 2022 Political TV Advertising.” Versions 1.0, 1.0 & 1.0. The Wesleyan Media Project, Wesleyan University.

</div>

<div id="ref-gilardi2023chatgpt" class="csl-entry">

Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. 2023. “ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks.” *Proceedings of the National Academy of Sciences* 120 (30): e2305016120.

</div>

<div id="ref-grimmer2013text" class="csl-entry">

Grimmer, Justin, and Brandon M Stewart. 2013. “Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts.” *Political Analysis* 21 (3): 267–97.

</div>

<div id="ref-guo2017calibration" class="csl-entry">

Guo, Chuan, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. “On Calibration of Modern Neural Networks.” *International Conference on Machine Learning*, 1321–30.

</div>

<div id="ref-halterman2026codebook" class="csl-entry">

Halterman, Andrew, and Katherine A Keith. 2026. “Codebook Llms: Evaluating Llms as Measurement Tools for Political Science Concepts.” *Political Analysis* 34 (2): 188–204.

</div>

<div id="ref-hulme2025war" class="csl-entry">

Hulme, M Patrick. 2025. “War and Responsibility.” *American Political Science Review*, 1–24.

</div>

<div id="ref-imai2010causal" class="csl-entry">

Imai, Kosuke, and Teppei Yamamoto. 2010. “Causal Inference with Differential Measurement Error: Nonparametric Identification and Sensitivity Analysis.” *American Journal of Political Science* 54 (2): 543–60.

</div>

<div id="ref-krippendorff2018content" class="csl-entry">

Krippendorff, Klaus. 2018. *Content Analysis: An Introduction to Its Methodology*. Sage publications.

</div>

<div id="ref-le2025positioning" class="csl-entry">

Le Mens, Gaël, and Aina Gallego. 2025. “Positioning Political Texts with Large Language Models by Asking and Averaging.” *Political Analysis* 33 (3): 274–82.

</div>

<div id="ref-lin2025using" class="csl-entry">

Lin, Gechun. 2025. “Using Cross-Encoders to Measure the Similarity of Short Texts in Political Science.” *American Journal of Political Science*.

</div>

<div id="ref-mellon2024ais" class="csl-entry">

Mellon, Jonathan, Jack Bailey, Ralph Scott, James Breckwoldt, Marta Miori, and Phillip Schmedeman. 2024. “Do AIs Know What the Most Important Issue Is? Using Language Models to Code Open-Text Social Survey Responses at Scale.” *Research & Politics* 11 (1): 20531680241231468.

</div>

<div id="ref-miller2020active" class="csl-entry">

Miller, Blake, Fridolin Linder, and Walter R Mebane Jr. 2020. “Active Learning Approaches for Labeling Text: Review and Assessment of the Performance of Active Learning Approaches.” *Political Analysis* 28 (4): 532–51.

</div>

<div id="ref-milliff2024making" class="csl-entry">

Milliff, Aidan. 2024. “Making Sense, Making Choices: How Civilians Choose Survival Strategies During Violence.” *American Political Science Review* 118 (3): 1379–97.

</div>

<div id="ref-raleigh2010introducing" class="csl-entry">

Raleigh, Clionadh, Rew Linke, Håvard Hegre, and Joakim Karlsen. 2010. “Introducing ACLED: An Armed Conflict Location and Event Dataset.” *Journal of Peace Research* 47 (5): 651–60.

</div>

<div id="ref-sap2021annotators" class="csl-entry">

Sap, Maarten, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A Smith. 2021. “Annotators with Attitudes: How Annotator Beliefs and Identities Bias Toxic Language Detection.” *arXiv Preprint arXiv:2111.07997*.

</div>

<div id="ref-timoneda2025memory" class="csl-entry">

Timoneda, Joan C, and Sebastián Vallejo Vera. 2025. “Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks.” *arXiv Preprint arXiv:2503.04874*.

</div>

<div id="ref-widmann2025politicians" class="csl-entry">

Widmann, Tobias. 2025. “Do Politicians Appeal to Discrete Emotions? The Effect of Wind Turbine Construction on Elite Discourse.” *The Journal of Politics* 87 (1): 335–46.

</div>

<div id="ref-wu2024concept" class="csl-entry">

Wu, Patrick Y, Jonathan Nagler, Joshua A Tucker, and Solomon Messing. 2024. “Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models.” *2024 IEEE International Conference on Big Data (BigData)*, 7232–41.

</div>

<div id="ref-ziems2024can" class="csl-entry">

Ziems, Caleb, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. “Can Large Language Models Transform Computational Social Science?” *Computational Linguistics* 50 (1): 237–91.

</div>

</div>

[^1]: We distinguish LLMs from supervised machine learning models by their ability to annotate without training () on the specific annotation task. Our definition differs from some existing studies. In particular, we consider models like BERT to be supervised models rather than LLMs.

[^2]: See e.g., Mellon et al. (2024; Breuer et al. 2025; Le Mens and Gallego 2025; Lin 2025).

[^3]: We define an LLM as if its trained parameters (weights) are publicly available for anyone to download and use. In contrast, proprietary models are those for which the weights are not publicly available.

[^4]: Recent LLMs are trained on tens of trillions of tokens, where a token is the smallest unit (typically a sub-word) that LLMs use to represent text.

[^5]: In addition to textual data, some LLMs can also take other media, such as audio, image, and video data, as input. In this paper, we focus on LLMs’ textual abilities.

[^6]: For example, reasoning models from Google and OpenAI achieved gold medal-level performance at the 2025 International Mathematical Olympiad (IMO). See <https://www.axios.com/2025/07/21/openai-deepmind-math-olympiad-ai>.

[^7]: <https://docs.vllm.ai/>.

[^8]: <https://github.com/soichiroy/prisa>

[^9]: Of the 14 studies in our sample, four reported Krippendorff’s alpha, with values of $`0.56`$, $`0.78`$, $`0.84`$, and $`0.84`$.

[^10]: A permutation test (SM <a href="#sec:human_bench" data-reference-type="ref" data-reference="sec:human_bench">14</a>) over annotator identities confirms that this difference is unlikely to arise by chance: the observed human-human minus human-LLM gap is positive in every study and statistically distinguishable from the permutation distribution in all six studies ($`p \le 0.026`$).

[^11]: This is achieved by sampling text-annotation pairs with replacement.

[^12]: <https://docs.sglang.ai/>
