Research
Publications
Abstract
We develop specialized language models for political conflict analysis that outperform general-purpose LLMs like Gemma 2, Llama 3.1, and Qwen 2.5 in accuracy, precision, and recall while being hundreds of times faster.
Abstract
We introduce ConflLlama, a specialized variant of Llama 3.1 fine-tuned for political conflict classification, demonstrating superior performance in event coding and conflict analysis tasks compared to traditional approaches.
Abstract
We analyze how health advocacy groups adapted their Medicare-For-All messaging on Twitter during the COVID-19 pandemic, revealing distinct approaches to public engagement and narrative adaptation.
Abstract
This paper examines patterns of executive dominance and expansion in the Netherlands, a least-likely case of executive aggrandizement given its strong coalitional governance and consensus culture. Drawing on a legislative dataset, our findings show that executive dominance over the legislative agenda and the overall volume of delegated lawmaking are longstanding and stable: roughly 75% of legislation takes the form of secondary executive acts, a level that does not increase over our study period. What has expanded is the legal technique underpinning this dominance: framework laws have become an increasingly cited basis for transferring regulatory authority to individual ministers, with citations of framework instruments in new executive acts rising markedly since 2017. Weak oversight tools such as motions and parliamentary questions remain the opposition’s primary (and historically stable) mode of engagement, but coalition MPs have grown markedly more vocal in using these same tools since 2017, reflecting tensions within fragmented governing coalitions. While large, diverse coalitions create internal checks against outright aggrandizement, the growing reliance on framework laws as a delegation device represents a gradual but consequential shift in how, rather than how much, executive dominance is exercised.
Under Review
Abstract
How does regime type shape a state’s intervention in its information environment? We argue that democracies and non-democracies differ not only in the quantity but, more importantly, in the institutional channel of content removal. To that end, we develop a political agency model and show that a democratic government would refrain from direct executive takedown and instead delegate to the court. This is because electoral accountability disciplines politicians’ behavior by providing incentives for reputation-building. We test our hypotheses using data from Google transparency reports. Exploiting quasi-experimental variation in the timing of constitutionally scheduled elections, we provide supporting evidence that compared to non-democracies, takedown requests from democratic governments decline significantly as elections approach. This reputational discipline effect is not observed in authoritarian regimes or other types of requests.
Two Types of Censorship? An Assessment of the Informational Autocracy Thesis
Abstract
That executive aggrandizement proceeds through legal, procedural channels is well established. This paper asks whether governments that withdraw legislative scrutiny do so indiscriminately, or target the legislation that expands their own power. I develop a theory of strategic procedural aggrandizement for parliamentary systems in which scrutiny is discretionary to the government, and test it by linking 653 Indian parliamentary bills (2009–2026) to the full text of the statutes they became, each scored with a 316-phrase dictionary of executive aggrandizement. Under coalition government, scrutiny tracked stakes: power-expanding bills were referred to committee more often and passed more slowly. Under the single-party majority since 2014, this relationship inverted. Committee referral fell from 74% to 17%, median passage time from 255 to 18 days, and the most aggrandizing statutes cleared parliament in days. The pattern is institutional, not ideological: scrutiny recovered when the majority shrank in 2024. The mechanism generalizes to Westminster-heritage parliaments.
Abstract
Researchers in computational social science increasingly face a consequential choice when adopting natural language processing tools: build a domain-specific model from scratch, borrow and adapt an existing one, or simply fine-tune a general-purpose model on task data? Each approach occupies a different point on the spectrum of performance, cost, and required expertise, yet the discipline has offered little empirical guidance on how to navigate this trade-off. This paper provides such guidance. Using conflict event classification as a test case, I fine-tune ModernBERT on the Global Terrorism Database (GTD) to create Confli-mBERT and systematically compare it against ConfliBERT, a domain-specific pretrained model that represents the current gold standard. Confli-mBERT achieves 75.46% accuracy compared to ConfliBERT’s 79.34%. Critically, the four-percentage-point gap is not uniform: on high-frequency attack types such as Bombing/Explosion and Kidnapping, the models are nearly indistinguishable. Performance differences concentrate in rare event categories comprising fewer than 2% of all incidents. I use these findings to develop a practical decision framework applicable to any NLP-assisted classification task: when does the research question demand a specialized model, and when does an accessible fine-tuned alternative suffice? The marginal value of domain-specific pretraining, I show, is a predictable function of class prevalence, error tolerance, and available resources. The model, training code, and data are publicly available on Hugging Face.
Abstract
On 2 February 2021, the musician Rihanna tweeted about the Indian farmer protests to her 100 million followers, producing the largest attention spike in the movement’s yearlong history. Using 1.02 million tweets classified by a large language model, I show that the intervention displaced the movement’s own discourse rather than diluting it with newcomers. Benchmarked against 277 placebo dates, the celebrity share of conversation rose 33 percentage points while the policy share fell 8 points, and a within-user decomposition attributes most of the shift to established participants changing their own speech. The shift carried costs: participants who pivoted hardest toward celebrity content were half as likely to remain active a month later, and a coordinated pro-government campaign penetrated the movement’s hashtags for the first time. Comparison with six domestic events shows the pattern is specific to the celebrity shock. I call this dynamic the attention trap.
Abstract
India accounts for more internet shutdowns than any other country, yet the public record of these events is strikingly uneven: some shutdowns dominate national news coverage while others vanish from it entirely. Drawing on a combined panel of 607 state-level internet shutdowns (2013–2024) and 1.5 million Times of India articles, I measure per-event media amplification and ask which shutdowns reach the national press agenda. I find a sharp, structural amplification gap that follows India’s securitized periphery and is deepest in Jammu and Kashmir, where shutdowns receive roughly 2.7 times fewer articles than comparable shutdowns elsewhere, conditional on duration, geographic scope, year, and data source. The gap survives consolidating fragmented event records, fixing the coverage window, and randomization inference on the periphery indicator; about half of it is explained by the newspaper’s generally thin attention to the region, and the remainder is specific to shutdowns. The gap is fully formed before the August 2019 abrogation of Article 370 and is not explained by state ruling party. Shutdowns that disrupt legible civic activities attract disproportionate coverage. I argue that this pattern reflects not a failure of journalistic capacity, which existing work shows endures under shutdowns, but a structural failure of amplification: the national English-language press retains the freedom to report but systematically declines to place certain shutdowns on the national agenda. I discuss implications for how press freedom is conceptualized in democracies exhibiting backsliding.
Working Papers
Abstract
Does the European Commission’s enlargement reporting emphasise executive strength over institutional checks on executive power, and does that emphasis track later democratic backsliding? Building on Meyerrose’s argument that international democracy support can strengthen executives at the expense of the institutions meant to constrain them, we use large language models to content-analyse the Commission’s progress reports for 23 candidate countries (105,971 sentences across 207 country-years, 1998 to 2023), separating what each report emphasises (compliance and state capacity versus judicial, legislative, civic, and media checks) from how it assesses each domain. Relating both measures to the V-Dem Liberal Democracy Index, we find that the Commission’s assessments track subsequent democratic outcomes while its relative emphasis does not, a contrast that survives bootstrap inference and a coder-based correction for measurement error. Emphasis is nonetheless associated with later decline where judicial constraints are weak.
Abstract
What determines whether an expert advisory body issues a critical opinion on a legislative proposal? This paper investigates the conditions under which the Dutch Council of State (Raad van State) issues critical advisory opinions, testing five competing theoretical mechanisms: institutional learning within coalition governments, legislative experience of the proposing actor, the coalition position of the sponsoring party, ideological distance from the governing majority, and the type of legislative instrument under review. Drawing on 2,898 advisory opinions issued between 2004 and 2025, comprising 2,111 opinions on laws and 785 on executive decrees, we find that instrument type is the strongest predictor of opinion severity: executive decrees receive dramatically milder treatment than laws, with 69% receiving no objections compared to just 13% for laws. Among laws, junior coalition partners attract significantly less critical opinions than leading parties, consistent with anticipatory compliance. Ideological distance from the cabinet median is positively associated with opinion severity, though the effect is modest. These findings suggest that the Council’s most consequential effects operate through two channels: institutional deference to delegated legislation, and the anticipatory self-discipline of politically exposed actors.
Abstract
Three decades of research on parliamentary questions have mapped who asks them, and almost never what askers get back. I turn to the answers, and to a puzzle: the Dutch government replies to its fiercest critics fastest. Opposition and populist members are answered as quickly as the coalition’s own backbenchers, or quicker, and three answers in four miss the 21-day deadline for everyone alike. Timing the answers makes the government look even-handed. Reading them does not. Scoring more than 130,000 question–answer pairs (2010–2025) with a validated large-language-model measure of evasion, I find that critics are answered no more slowly, yet 0.20 to 0.47 standard deviations more evasively, net of cabinet, topic, and ministry. The emptiness is aimed: within a single reply, the substantive sub-questions that demand an account are evaded most, and more so when the asker is a critic. Two probes trace the pattern to its incentives. When a question arrives with a news story attached, critics are answered faster still, yet no more fully; and evasive answers draw no more follow-up questions than responsive ones, so evasion costs a minister nothing. One logic runs through all of it, and it is visibility. Delay is public and punished; evasion is buried, unread, and unpriced. A government that treats some askers worse than others does so through the channel where the discrimination will not be seen. Ministers answer on time, and empty the answer instead.
Abstract
Many constructs in political analysis are realised at corpus scale across thousands of legal artefacts, and the substantive phenomenon often spans the relationship between documents rather than residing within any single one. Computational text analysis nonetheless inherits document- and sentence-level units from earlier hand-coding traditions, leaving the unit choice as an unexamined default. I work through the consequences for executive aggrandizement, the incremental and legally formal expansion of executive power. Applied to Dutch legislation between 2007 and 2026 (3,443 bills, 12,884 parent-implementing edges, 488 chains with consensus labels), I compare four methods spanning the design space: a keyword density, a directional grid, a twelve-flag composite combining textual and procedural signal, and a multi-model LLM committee that scores delegation chains on a five-dimension schema. Document-level and chain-level scores are weakly negatively correlated, and the divergence is asymmetric: 28% of chains with a top-decile parent law receive a low chain rating, while only 2.3% show the reverse pattern. Document-level methods narrow the universe of candidate concerns; chain-level methods calibrate within it. The two answer different questions, and conflating them produces conclusions the data do not support.
Abstract
When working with real-world text, researchers often inherit corpora and annotations with costly human judgments but were not collected, organized, or annotated for modern text-as-data workflows. This is especially common in social science domains where web-scraped documents are noisy, incomplete, or misaligned with the unit of analysis. We propose codebook attention: a codebook-guided turn-by-turn extractive summarization pipeline that reuses valuable annotations without requiring full re-annotation. For each victim in a corpus of disappearance reports from Mexico, the model processes documents sequentially, extracts verbatim evidence spans for each codebook category, and updates a running structured summary to produce a clean and evidence-traceable document-level representation. Using the codebook as the task schema focuses model attention on the fields that matter for downstream classification and reconnects prior human annotations to contemporary language model workflows. We evaluate the resulting summaries by classifying and comparing the predicted labels with human annotations across multiple open-source models including Llama 3.1, Gemma 3 and Ministral 3. Results suggest that hours of computation can approximate weeks of human coding effort while producing evidence-traceable summaries suitable for downstream human rights and event data research.
Event Horizon: Revolutionizing Data Annotation with Reinforcement Learning Model
Dissertation
Digital Sovereignty: The Political Economy of Internet Governance Slides