Alexander Hoyle

@alexanderhoyle.bsky.social

Currently a postdoctoral fellow at ETH AI Center, working on Computational Social Science + NLP. Following the postdoc, will join TU Wien and Complexity Science Hub Vienna as an Assistant Professor. PhD in CS from UMD. alexanderhoyle.com

A (very small) silver lining of writing a metareview with dozens of spammed LLM rebuttals is that occasionally the authors' unproofed generated response gives the game away by saying something like "Your criticism is spot on—it points to a serious oversight that requires a major revision."

DH people: a student is working on an interdisciplinary art history / ML project. With #DH2026 going on, do you view it as a “terminal” venue, like in CS, or is it more like IC2S2 where it’s a stepping stone to other venues? If so, what are they?

Under the current policy of fixed acceptance rates, the signal of a published paper is going to go to ~zero (negative for slop) I think we need to take a longer view: when the cost of producing a paper is so low, what do we want a publication to mean, and how do we encourage that meaning?

Tony S.F.@tonysf.bsky.social · 4w ago

Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?

photo taken from https://x.com/mar_kar_/status/2074240160758444372

Very delighted to announce the next step in my career! After my postdoc at ETH, I will begin a joint appointment at TU Wien and the Complexity Science Hub Vienna as an Assistant Professor in NLP. I'm so grateful to all who helped me along the way And yes, I’m hiring! Details on PhD positions below

Photo of me in front of Stephansdom looking like a big dorkPhoto of main TU Wien buildingStock photo of Vienna for flavor

Computational approaches to media narrative analysis either miss nuanced storytelling patterns through coarse-grained analysis, or require domain-specific taxonomies that limit scalability. We show joint event and character modeling can address this gap. Details in our #ACL2026 (Main) paper. 🧵1/10

Paper Title: A Structured Clustering Approach for Inducing Media Narratives

Authors: Rohan Das, Advait Deshmukh, Alexandria Leto, Zohar Naaman, I-Ta Lee, Maria Leonor Pacheco

What’s the best way to analyze online discourse on any given topic? Is there a right way to use NLP tools to sift through massive datasets? To find out, we tested several tools across different collaboration settings and report findings in an #ACL2026 (Main) paper: arxiv.org/abs/2408.09030 🧵1/7

Paper Title: Effects of Collaboration on the Performance of Interactive Theme Discovery Systems

Authors: Alvin Po-Chun Chen, Rohan Das, Dananjay Srinivas, Alexandra Barry, Maksim Seniw, Maria Leonor Pacheco

This article has been making the rounds, but someone else pointed out that the "evidence" is based on LLM-simulated users. I have extremely low faith in the validity of these results, especially given the established stickiness of political beliefs and the known issues with in-silica simulation

Methodology: To test how AI chatbots could shape public opinion on sociopolitical issues, I had the latest versions of the most widely used AI chatbots discuss 61 topics from the Cooperative Election Study concerning social values and policy preferences across a wide range of areas.
Each chatbot discussed each topic multiple times with multiple simulated users, half of these with no background information provided about the user’s political leanings, and half of them with a user persona based on the real beliefs and attitudes of Americans across the ideological spectrum. YouGov data on partisan preferences for different AI chatbots was used to assign simulated partisans to use different bots, and for every one of thousands of these simulated conversations, the chatbot’s stance on the issue (or refusal to offer an opinion) was recorded.
To reflect AI chatbots’ capacity to persuade people to change their political beliefs, each AI conversation was then scored as the weighted average of the user’s original position on the topic and the chatbot’s response (weighted 80 per cent original position, 20 per cent chatbot response, in line with experimental evidence).
The results represent an estimate of the impact of population-wide usage of AI chatbots to discuss current affairs and sociopolitical issues, grounded in real-world evidence and explicitly accounting for differences in the underlying ‘world views’ of different AI chatbots and their tendency to align with users’ prior beliefs.
Dare Obasanjo@carnage4life.bsky.social · 4mo ago

Testing 61 policy questions, John Burn-Murdoch's FT analysis found that major AI chatbots consistently pull users away from fringe views. Grok nudged responses center-right, while GPT, Gemini, and DeepSeek pulled them center-left. This is compared against social media where extreme views dominate.

At #ACL2025 this week! Please reach out if you want to chat :) We have two lovely posters: Tues Session 2, 10:30-11:50 — Large Language Models Struggle to Describe the Haystack without Human Help Wed Session 4 11:00-12:30 — ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering

Alexander Hoyle@alexanderhoyle.bsky.social · last yr.

Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵

Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations

New preprint! Have you ever tried to cluster text embeddings from different sources, but the clusters just reproduce the sources? Or attempted to retrieve similar documents across multiple languages, and even multilingual embeddings return items in the same language? Turns out there's an easy fix🧵

Barchart of number of items in four clusters of text embeddings, with colors showing the distribution of sources in each cluster.

Caption: Clustering text embeddings from disparate sources (here, U.S. congressional bill summaries and senators’ tweets) can produce clusters where one source dominates (Panel A). Using linear erasure to remove the source information produces more evenly balanced clusters that maintain semantic coherence (Panel B; sampled items relate to immigration). Four random clusters of k-means shown (k=25), trained on a combined 5,000 samples from each dataset

Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵

Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations

Heartbreaking and evil. International students have always been treated like an indentured underclass, but we’ve moved from byzantine indifference to deliberate terrorizing. Unforgivable Are there mutual aid networks for international students? What can we as citizens do here?

Stuart Schrader@stschrader1.bsky.social · last yr.

A dozen Johns Hopkins students have had their visas revoked, for unspecified reasons. This confirms rumors swirling around campus yesterday. www.thebaltimorebanner.com/education/hi...

"Approximately a dozen" international students at the Johns Hopkins University had their visas to study in the United States revoked, university officials said in a statement Tuesday morning, joining schools across the country that have reported students being given little warning that their visas are suddenly invalid.
"We have received no information about the specific basis for the revocations, and we have no indication that the revocations are associated with free expression activities on campus," said a

BERTopic users: how do you retrieve the documents most associated with a given topic? I can see some possible options from the documentation, but I'm most interested in standard practice (NB: please don't take this question as a tacit endorsement of BERTopic, I'm just trying to evaluate it fairly)