Alexander Hoyle
@alexanderhoyle.bsky.social
Currently a postdoctoral fellow at ETH AI Center, working on Computational Social Science + NLP. Following the postdoc, will join TU Wien and Complexity Science Hub Vienna as an Assistant Professor. PhD in CS from UMD. alexanderhoyle.com
A (very small) silver lining of writing a metareview with dozens of spammed LLM rebuttals is that occasionally the authors' unproofed generated response gives the game away by saying something like "Your criticism is spot on—it points to a serious oversight that requires a major revision."
DH people: a student is working on an interdisciplinary art history / ML project. With #DH2026 going on, do you view it as a “terminal” venue, like in CS, or is it more like IC2S2 where it’s a stepping stone to other venues? If so, what are they?
Under the current policy of fixed acceptance rates, the signal of a published paper is going to go to ~zero (negative for slop) I think we need to take a longer view: when the cost of producing a paper is so low, what do we want a publication to mean, and how do we encourage that meaning?
Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?
one benefit of the rise of ai writing is that i feel more personal liberty to let my voice come through in my writing. colloquialisms, asides, jokes, mixing of registers, etc
folks, don't use claude for slides. the style's immediately recognizable, it looks good, but it slides right over people's brains. in an era where sounding professional doesn't cost nothing to nobody, just some tokens, we're entirely free to be non-professional . embrace dreamcore wordart or smth
Very delighted to announce the next step in my career! After my postdoc at ETH, I will begin a joint appointment at TU Wien and the Complexity Science Hub Vienna as an Assistant Professor in NLP. I'm so grateful to all who helped me along the way And yes, I’m hiring! Details on PhD positions below
We are implementing a similar policy at @arxiv.bsky.social. If there is incontrovertible evidence of LLM slop in a paper, this means the authors did not take the time to read the LLM output and we can't trust anything else in the paper. Penalty is 1 year ban from arXiv followed by...
Computational approaches to media narrative analysis either miss nuanced storytelling patterns through coarse-grained analysis, or require domain-specific taxonomies that limit scalability. We show joint event and character modeling can address this gap. Details in our #ACL2026 (Main) paper. 🧵1/10
This paper is getting a lot of (deserved) attention; I think this paper serves as a nice complement arxiv.org/abs/2602.18710
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent teams testing the same h...
arxiv.org
What can we learn from automating an entire quantitative social science paper, from prompt to finished product? Thread about ongoing work with @natewilmers.bsky.social 1/12 Paper: osf.io/preprints/so...
What’s the best way to analyze online discourse on any given topic? Is there a right way to use NLP tools to sift through massive datasets? To find out, we tested several tools across different collaboration settings and report findings in an #ACL2026 (Main) paper: arxiv.org/abs/2408.09030 🧵1/7
This article has been making the rounds, but someone else pointed out that the "evidence" is based on LLM-simulated users. I have extremely low faith in the validity of these results, especially given the established stickiness of political beliefs and the known issues with in-silica simulation
Testing 61 policy questions, John Burn-Murdoch's FT analysis found that major AI chatbots consistently pull users away from fringe views. Grok nudged responses center-right, while GPT, Gemini, and DeepSeek pulled them center-left. This is compared against social media where extreme views dominate.
I wrote a blog post on my experience using AI for slide generation Basic idea: write your lecture notes first, then prompt the LLM to produce corresponding slides in reveal.js (h/t @chenhaotan.bsky.social). I'm picky about my slides but was happy with the results! alexanderhoyle.com/posts/ai-sli...
[corrected link] LLMs are often used for text annotation in social science. In some cases, this involves placing text items on a scale: eg, 1 for liberal and 9 for conservative There are a few ways to handle this task. Which work best? Our new EMNLP paper has some answers🧵 arxiv.org/abs/2509.03116
LLMs are often used for text annotation, especially in social science. In some cases, this involves placing text items on a scale: eg, 1 for liberal and 9 for conservative There are a few ways to accomplish this task. Which work best? Our new EMNLP paper has some answers🧵 arxiv.org/pdf/2507.00828
Computer Science is no longer just about building systems or proving theorems--it's about observation and experiments. In my latest blog post, I argue it’s time we had our own "Econometrics," a discipline devoted to empirical rigor. doomscrollingbabel.manoel.xyz/p/the-missin...
Accepted to EMNLP (and more to come 👀)! The camera ready version is now online---very happy with how this turned out arxiv.org/abs/2507.01234
New preprint! Have you ever tried to cluster text embeddings from different sources, but the clusters just reproduce the sources? Or attempted to retrieve similar documents across multiple languages, and even multilingual embeddings return items in the same language? Turns out there's an easy fix🧵
this looks terrific, very excited to read
LLMs introduce a huge range of new capabilities for research, but also make it possible for researchers to "hack" their results in new ways by how they chose to use models for annotation This is a useful pass at quantifying some of the risk, and some mitigation strategies arxiv.org/pdf/2509.08825
I am delighted to share our new #PNAS paper, with @grvkamath.bsky.social @msonderegger.bsky.social and @sivareddyg.bsky.social, on whether age matters for the adoption of new meanings. That is, as words change meaning, does the rate of adoption vary across generations? www.pnas.org/doi/epdf/10....
At #ACL2025 this week! Please reach out if you want to chat :) We have two lovely posters: Tues Session 2, 10:30-11:50 — Large Language Models Struggle to Describe the Haystack without Human Help Wed Session 4 11:00-12:30 — ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
🗣️ Excited to share our new #ACL2025 Findings paper: “Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks” with Jad Kabbara and Deb Roy. Arxiv: arxiv.org/abs/2507.15821 Read about our findings ⤵️
Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks
LLM use in annotation is becoming widespread, and given LLMs' overall promising performance and speed, simply "reviewing" LLM annotations in interpretive tasks can be tempting. In subjective annotatio...
arxiv.org
The precursor to this paper "The Incoherence of Coherence" had our most-watched paper video ever, so I thought we had to surpass it somehow ... so we decided to do a song parody (of Roxanne, obviously): youtu.be/87OBxEM8a9E
ProxAnn (Sting Parody): An AI Alternative to NPMI [Research]
YouTube video by Jordan Boyd-Graber
youtu.be
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
New preprint! Have you ever tried to cluster text embeddings from different sources, but the clusters just reproduce the sources? Or attempted to retrieve similar documents across multiple languages, and even multilingual embeddings return items in the same language? Turns out there's an easy fix🧵
Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵
Michael Roth's recent outspokenness has made me proud to be a Wes alum. A decade of NYT op-ed handwringing about "free speech" on campuses has only provided ammunition for bad faith attacks on academia (Perhaps I should be better at responding to those fundraising emails)
www.insidehighered.com/news/faculty...
I for one am grateful for the opportunity to meditate on the meaning of “scientific artifact” at 2:15am
Do folks really find the ARR checklist valuable enough to justify that a paper submission takes this much effort?
Heartbreaking and evil. International students have always been treated like an indentured underclass, but we’ve moved from byzantine indifference to deliberate terrorizing. Unforgivable Are there mutual aid networks for international students? What can we as citizens do here?
A dozen Johns Hopkins students have had their visas revoked, for unspecified reasons. This confirms rumors swirling around campus yesterday. www.thebaltimorebanner.com/education/hi...
this holiday season I am thankful that, rather than fixing the literally decade-old problem of multi-file search, Overleaf instead implemented the world's worst writing assistance tool
BERTopic users: how do you retrieve the documents most associated with a given topic? I can see some possible options from the documentation, but I'm most interested in standard practice (NB: please don't take this question as a tacit endorsement of BERTopic, I'm just trying to evaluate it fairly)