New paper in Nature. The more a government controls its domestic media, the more it dominates AI training data, the more pro-regime outputs we get from AI. By scraping the open web, LLMs are unwittingly laundering state-coordinated narratives into seemingly objective answers.
I wrote about my experience tinkering with AI for social science research: open.substack.com/pub/ey211/p/...
Tinkering with AI for social science research
I spent two weekends figuring out how much I can integrate AI in my own research.
open.substack.com
New paper out at AJPS: "The limits of AI for authoritarian control." The more repression there is, the less information exists in AI's training data, and the worse the AI performs.
We've updated the localLLM package on CRAN (cran.r-project.org/package=loca...). It allows you to run LLMs locally and natively in R. Everything is reproducible. And it's free. Some functionalities on reproducibility and validation🧵👇
localLLM: Running Local LLMs with 'llama.cpp' Backend
Provides R bindings to the 'llama.cpp' library for running large language models. The package uses a lightweight architecture where the C++ backend library is downloaded at runtime rather than bundled...
cran.r-project.org
Russia, Venezuela, Iran, China, the Sahel region, the United States ... Want to know why state agents carry out brutal repression — or participate in illegal coups? Our new book "Making a Career in Dictatorship" provides answers — it just got published by @academic.oup.com: tinyurl.com/ystwm3tf
So, randomization is not a *sufficient* condition for good research. Far from it. The best experimental social science being done is that work where either the theory or the operationalization (or both) are the emphasis. Randomization is "easy" - the challenge is what you randomize and why.
Can Large Multimodal Models (LMMs) extract features of urban neighborhoods from street-view images? New with Paige Bollen (OSU) and @joehigton.bsky.social (NYU): Sometimes, but the models better recover national assessments that local ones, even w/additional prompting (which can make things worse!)
Currently in FirstView: In “Nationally Representative, Locally Misaligned: The Biases of Generative Artificial Intelligence in Neighborhood Perception,” Paige Bollen, @joehigton.bsky.social, and @msands.bsky.social test which populations Generative AI is most representative of.
New paper: LLMs are increasingly used to label data in political science. But how reliable are these annotations, and what are the consequences for scientific findings? What are best practices? Some new findings from a large empirical evaluation. Paper: eddieyang.net/research/llm_annotation.pdf
Great analogy to connect AI to many canonical political science questions. Political behavior has led the way in studying AI. Excited to see institutions catch up😀
My "AI as Governance" piece is now out at @annualreviews.bsky.social of Political Science. It should be free access to everyone and I'm very happy with how it worked out (the second half is an extended spin of Applied Gopnikism to political science @alisongopnik.bsky.social @cshalizi.bsky.social
If no resource constraint, what open-weight LLM would you use in your research (for data labeling, coding etc.)?
Awesome work! Love to see different approaches to this problem.
1/9 We are excited to share our new working paper: arxiv.org/abs/2502.12323 If you use ML predictions (like remote-sensed data) as outcomes, the resulting regression coefficients can be biased by measurement error. With @megan-ayers.bsky.social @mdgordo.bsky.social @eliana-stone.bsky.social
Really interesting read. Refreshing perspective.
1. @alisongopnik.bsky.social, Cosma Shalizi, James Evans and myself have a new piece in Science on "AI" Large Models, pushing back against much of the collective wisdom about what they can and can't do. Official below, unpaywalled at henryfarrell.net/large-ai-mod... . So why this now?