This morning at #ic2s2 I have the chance to present ongoing work on applying a SEM from survey methods to LLM text annotations: 👉 Evaluating LLM Text Annotations Without Ground Truth 📍10:45, Mansfield (210), Understanding LLMs w/ @maximiliankreutner.bsky.social Alex Cernat & @mstrohm.bsky.social
Manoel Horta Ribeiro
@manoelhortaribeiro.bsky.social
Assistant Professor @ Princeton Previously: EPFL 🇨🇭, UFMG 🇧🇷 Interests: Computational Social Science, Platforms, GenAI, Moderation
@inhwa.bsky.social and @mariannealq.bsky.social presenting our ongoing work on news knowledge and anxiety. Our internal joke is that this project is to figure out what can make people both ‘chill’ and informed!
@mariannealq.bsky.social and @isabelcorpus.bsky.social presenting their work on local variations on AI overviews! They find that AI overviews are much better in high information environments.
Let me start by changing my own practices! 🤦♂️ But I think the recent reporting checklist paper at NHB (led by @sfeuerriegel.bsky.social) is relevant to the discussion. www.nature.com/articles/s41...
A reporting checklist for large language models in behavioural science - Nature Human Behaviour
Large language models offer new opportunities for behavioural science, but their rapid evolution poses challenges for research rigour. We introduce a consensus-based reporting checklist to improve tra...
nature.com
(feel free to complain to me to) And for the #IC2S2 crowd, I'll also point to the work of the Coalition of Independent Technology Research in helping protect CSS research & researchers, and encourage you to join the organization. independenttechresearch.org
Home - Coalition for Independent Technology Research
independenttechresearch.org
I hope folks have a great time at #ic2s2 this week! So say hi to @davidlazer.bsky.social if you're interested in thinking about how to make computational social science more repeatable, and feel free to complain to me about any errors in our work: citizensandtech.org/2026/07/comm...
Attending #IC2S2!! Will be co-organizing a tutorial on Simulating Human Survey Responses with Large Language Models, today at 1.15PM, and presenting on Friday in the Understanding Large Language Models session, 10.45AM. Looking forward to catching up with old friends and meeting new ones!!
Headed to @ic2s2.bsky.social today!! Excited to meet old friends and make new ones! Come say hi! Also, bringing most of my lab with me for the first time! (The dog remained at @princetoncitp.bsky.social to supervise the remaining lab members) #IC2S2
❓❓Should LLMs have access to ACM’s Digital Library❓❓ I wrote a short essay arguing that "yes!" I try to argue that, even if you are concerned about their negative impact on science, denying them access is a poor way to address their harm! doomscrollingbabel.manoel.xyz/p/science-sh...
Science Should Be Open, For LLMs Too
Even if you think LLMs are bad for science, you should let them have access to research papers
doomscrollingbabel.manoel.xyz
1. We—@eduede.bsky.social, @mjcrockett.bsky.social, Kevin Gross, and I—have a new preprint on the arXiv today, based on ideas that emerged during an @sfiscience.bsky.social workshop in November 2024: The unintended consequences of large language models as a labor-augmenting technology in science.
The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at ...
arxiv.org
Hot take: ACM should just provide torrents called all_papers_{year_month_day}.tar.gz and let anyone download it!
Not a lot of info here about what kinds of agreements, with whom, and finances. But another example of data opening up for “AI” that wasn’t open for researchers. I fought many data wrangling battles at Semantic Scholar over missing ACM data. All research papers should be open access!
We are opening a consultation on the inclusion of ACM publications within AI licensing agreements for access and training purposes. We have published this piece as to why we believe it is the right time now to enter into these discussions. buff.ly/VTatuda
Just back from vacation, and very hyped for today's Brazil vs. Japan game. As a Brazilian, one of the most endearing things is seeing other countries (esp. other developing ones) show affection for our team! We don't deserve you all! Below is the least I expect from today's game :-)
I swear that never, ever again, will I take a connecting flight through Frankfurt
New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!
1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
Hot take: "AI slop" is a bad term and blurs the debate on the value of AI-generated content. doomscrollingbabel.manoel.xyz/p/ai-slop-isnt
AI Slop Isn’t
AI slop is a silly term, and we should abandon it
doomscrollingbabel.manoel.xyz
Please join our interview study 🙏
Are you a Bluesky user or developer? Princeton's HCI group is conducting an IRB-approved interview study on perspectives about generative AI (genAI) in decentralized social media. Participation includes: • 45–60 min Zoom interview • $30 USD digital gift card Signup: forms.gle/EDmcLnjbamtb...
New preprint! We introduce a new benchmark, SciConBench, with 9.11k scientific questions derived from Cochrane Systematic Reviews. We find evidence that frontier AI agents **cannot** synthesize scientific conclusions well. A thread 🧵 w/ @hayoungjung.bsky.social & others!
Quite a list of authors. See also Diyi Yang, @mariaa.bsky.social , M Sagalnik, @manoelhortaribeiro.bsky.social , &c
Very pleased to have been on the leadership team for this paper! LLMs are already being used all over behavioural science. But it is often pretty hard to work out exactly what has been done, and therefore how much confidence to place in the results. www.nature.com/articles/s41...
📝Excited to share our new preprint, “AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education”: arxiv.org/abs/2606.03095 A thread 🧵 1/8
🚨New preprint!🚨 arxiv.org/abs/2606.03095 In a randomized field experiment in a 300-level machine learning course, teaching assistants received AI-assisted feedback drafts after grading student submissions. Did AI help them provide more feedback when feedback was optional? A thread 🧵
“a town supervisor on Long Island had to debunk a rumor about a new data-center project after an inaccurate AI-generated search summary attracted so much attention that residents planned a protest (which they promoted with a flyer that itself appeared to be AI-generated).” @kait.bsky.social
The Rise of Anti-AI AI Slop
Anti-AI sentiment is genuine, but its online expression looks stranger and stranger.
theatlantic.com
In a new blog post, I reminisce about the Web I grew up in and argue that AI is only the latest blow to an already dying notion of "the Web." I argue that we shouldn't be overly reactionary because much of what AI threatens isn't worth saving. doomscrollingbabel.manoel.xyz/p/the-slow-u...
Social media is no longer social. Most of it is passive viewing of videos and pictures from people we've never met. But we're still studying social media like it's 2010. We've entered the post-social media era — and research needs to catch up. osf.io/preprints/so... preprint w/ Richard Rogers
Three things said about AI and social science: it's a topic to study, a thing to critique, a tool to use. My new Daedalus essay argues these aren't three conversations; they're one. A 🧵 on "Field Theory: AI as Social Science Question, Object, and Tool." www.amacad.org/publication/...
Field Theory: AI as Social Science Question, Object & Tool
Uses of advanced artificial intelligence are changing how societies organize labor, govern, produce knowledge, and make meaning. In light of these developments, this essay argues that AI models, tools...
amacad.org
In a new blog post, I argue that the anti-ai movement ought to distinguish between claims about the technology and the "project of AI," as defined by Vetsi et al. in their new paper. 🔗: doomscrollingbabel.manoel.xyz/p/the-anti-a...
My new post discusses the probable future of AI-assisted shopping! Drawing on our recent paper, I argue that if LLMs become the interface through which we search, compare, and buy things, we need a much sharper line between advertising and advice! doomscrollingbabel.manoel.xyz/p/when-the-a...
When the AI Assistant Becomes the Ad
This post describes our recent paper: Commercial Persuasion in AI-Mediated Conversations
doomscrollingbabel.manoel.xyz
In a post-CHI blog post, I talk about what I believe technology like the Anthropic Interviewer will mean for qualitative research. doomscrollingbabel.manoel.xyz/p/qualitativ...