🚨SciConBench accepted at #NeurIPS2026! Paper: arxiv.org/abs/2606.11337 Can AI *synthesize* scientific conclusions in health? We find that, on medical conclusions, we are far from it! Check out our live dashboard: sciconbench.cs.princeton.edu led by @hayoungjung.bsky.social
Manoel Horta Ribeiro
@manoelhortaribeiro.bsky.social
Assistant Professor @ Princeton Previously: EPFL 🇨🇭, UFMG 🇧🇷 Interests: Computational Social Science, Platforms, GenAI, Moderation
New blog post! I argue that the “stochastic parrots” metaphor dismisses LLM capabilities by reducing them to “haphazardly stitching together” text... ...but retreats to the trivially true claim that LLMs predict the next token when challenged. doomscrollingbabel.manoel.xyz/p/of-mottes-...
Of Mottes, Baileys, and Stochastic Parrots 🦜
Every six months, Twitter and Bluesky circle back to whether LLMs are “stochastic parrots,” though the debate has become increasingly confusing as proponents of the metaphor have started using it in v...
doomscrollingbabel.manoel.xyz
Do algorithms drive online polarization? A review by Duncan Watts and coauthors finds the evidence is inconclusive. What users seek out matters too, though platforms aren't off the hook. infodem.upenn.edu/research-com...
Research Compendium - Center on Media, Technology and Democracy
News, announcements, and insights from the Penn Center on Media, Technology, and Democracy.
infodem.upenn.edu
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
We're hiring in Computer Science at the University of Colorado Boulder! ☀️⛰️ Machine learning and NLP people, please apply!
Come join us in Boulder! Feel free to get in touch if you have any questions. jobs.colorado.edu/jobs/JobDeta...
Get to know Assistant Professor @tiziano.bsky.social in our New Faculty Q&A:
New faculty Q&A: Tiziano Piccardi
Get to know Tiziano Piccardi, who joins Johns Hopkins as an assistant professor of computer science and a member of the Data Science and AI Institute.
cs.jhu.edu
Maybe we should anthropomorphize AI systems (a tiny bit!) My latest blog post for Doomscrolling Babel is on dogs, parrots, Dennett, agent civilizations, and why human concepts can help generate useful hypotheses about how AI systems behave. doomscrollingbabel.manoel.xyz/p/in-favor-o...
What would it look like to move beyond normatively imposing a single set of values on everyone, and instead build generative AI systems that adapt to the values that matter most to each person in each situation? Excited that our work led by @rashonpoole.bsky.social will be at EMNLP!
Excited to share that our paper “Personalizing LLMs Through User Value Profiles” has been accepted to @emnlpmeeting.bsky.social 2026 Main Track! #emnlp2026
We got new research out today y'all--if you study TikTok, buckle up because we found something wilddddd... arxiv.org/abs/2608.09917 (accepted @ ICWSM 2027)
WhichTok? Comparing Three TikTok Data Acquisition Tools
TikTok's global growth has made it a prime platform for both entertainment and political discourse, prompting increased social science research. However, this rapidly evolving research field faces a f...
arxiv.org
Do our social media algorithms correctly reflect our values? Our new article published today in @pnas.org shows that the answer is often not, and that the content that gets promoted into their ranked feeds is often actively counter to our values.
What if AI is just normal technology? The human enterprise of science will continue with AI, as it continued with computers and search engines. We must adapt to this new reality since collective abstinence is an unlikely equilibrium!
The plagiarism machines break the chains of provenance of ideas and rest on a misconception of science as mere accumulation of objectively available facts. See: www.buzzsprout.com/2126417/epis... >>
This morning at #ic2s2 I have the chance to present ongoing work on applying a SEM from survey methods to LLM text annotations: 👉 Evaluating LLM Text Annotations Without Ground Truth 📍10:45, Mansfield (210), Understanding LLMs w/ @maximiliankreutner.bsky.social Alex Cernat & @mstrohm.bsky.social
@inhwa.bsky.social and @mariannealq.bsky.social presenting our ongoing work on news knowledge and anxiety. Our internal joke is that this project is to figure out what can make people both ‘chill’ and informed!
@mariannealq.bsky.social and @isabelcorpus.bsky.social presenting their work on local variations on AI overviews! They find that AI overviews are much better in high information environments.
Let me start by changing my own practices! 🤦♂️ But I think the recent reporting checklist paper at NHB (led by @sfeuerriegel.bsky.social) is relevant to the discussion. www.nature.com/articles/s41...
A reporting checklist for large language models in behavioural science - Nature Human Behaviour
Large language models offer new opportunities for behavioural science, but their rapid evolution poses challenges for research rigour. We introduce a consensus-based reporting checklist to improve tra...
nature.com
(feel free to complain to me to) And for the #IC2S2 crowd, I'll also point to the work of the Coalition of Independent Technology Research in helping protect CSS research & researchers, and encourage you to join the organization. independenttechresearch.org
Home - Coalition for Independent Technology Research
independenttechresearch.org
I hope folks have a great time at #ic2s2 this week! So say hi to @davidlazer.bsky.social if you're interested in thinking about how to make computational social science more repeatable, and feel free to complain to me about any errors in our work: citizensandtech.org/2026/07/comm...
Attending #IC2S2!! Will be co-organizing a tutorial on Simulating Human Survey Responses with Large Language Models, today at 1.15PM, and presenting on Friday in the Understanding Large Language Models session, 10.45AM. Looking forward to catching up with old friends and meeting new ones!!
Headed to @ic2s2.bsky.social today!! Excited to meet old friends and make new ones! Come say hi! Also, bringing most of my lab with me for the first time! (The dog remained at @princetoncitp.bsky.social to supervise the remaining lab members) #IC2S2
❓❓Should LLMs have access to ACM’s Digital Library❓❓ I wrote a short essay arguing that "yes!" I try to argue that, even if you are concerned about their negative impact on science, denying them access is a poor way to address their harm! doomscrollingbabel.manoel.xyz/p/science-sh...
Science Should Be Open, For LLMs Too
Even if you think LLMs are bad for science, you should let them have access to research papers
doomscrollingbabel.manoel.xyz
1. We—@eduede.bsky.social, @mjcrockett.bsky.social, Kevin Gross, and I—have a new preprint on the arXiv today, based on ideas that emerged during an @sfiscience.bsky.social workshop in November 2024: The unintended consequences of large language models as a labor-augmenting technology in science.
The unintended consequences of large language models as a labor-augmenting technology in science
As a labor-augmenting technology, large language models (LLMs) have the potential to accelerate scientific activity across the research pipeline. But even if LLMs perform on par with human experts at ...
arxiv.org
Hot take: ACM should just provide torrents called all_papers_{year_month_day}.tar.gz and let anyone download it!
Not a lot of info here about what kinds of agreements, with whom, and finances. But another example of data opening up for “AI” that wasn’t open for researchers. I fought many data wrangling battles at Semantic Scholar over missing ACM data. All research papers should be open access!
We are opening a consultation on the inclusion of ACM publications within AI licensing agreements for access and training purposes. We have published this piece as to why we believe it is the right time now to enter into these discussions. buff.ly/VTatuda
Just back from vacation, and very hyped for today's Brazil vs. Japan game. As a Brazilian, one of the most endearing things is seeing other countries (esp. other developing ones) show affection for our team! We don't deserve you all! Below is the least I expect from today's game :-)
I swear that never, ever again, will I take a connecting flight through Frankfurt
New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!
1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
Hot take: "AI slop" is a bad term and blurs the debate on the value of AI-generated content. doomscrollingbabel.manoel.xyz/p/ai-slop-isnt
AI Slop Isn’t
AI slop is a silly term, and we should abandon it
doomscrollingbabel.manoel.xyz