Manoel Horta Ribeiro

@manoelhortaribeiro.bsky.social

Assistant Professor @ Princeton Previously: EPFL 🇨🇭, UFMG 🇧🇷 Interests: Computational Social Science, Platforms, GenAI, Moderation

Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.

Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.

What would it look like to move beyond normatively imposing a single set of values on everyone, and instead build generative AI systems that adapt to the values that matter most to each person in each situation? Excited that our work led by @rashonpoole.bsky.social will be at EMNLP!

rashonpoole.bsky.social@rashonpoole.bsky.social · last mo.

Excited to share that our paper “Personalizing LLMs Through User Value Profiles” has been accepted to @emnlpmeeting.bsky.social 2026 Main Track! #emnlp2026

Do our social media algorithms correctly reflect our values? Our new article published today in @pnas.org shows that the answer is often not, and that the content that gets promoted into their ranked feeds is often actively counter to our values.

Bild

What if AI is just normal technology? The human enterprise of science will continue with AI, as it continued with computers and search engines. We must adapt to this new reality since collective abstinence is an unlikely equilibrium!

Emily M. Bender@emilymbender.bsky.social · 2mo ago

The plagiarism machines break the chains of provenance of ideas and rest on a misconception of science as mere accumulation of objectively available facts. See: www.buzzsprout.com/2126417/epis... >>

D

(feel free to complain to me to) And for the #IC2S2 crowd, I'll also point to the work of the Coalition of Independent Technology Research in helping protect CSS research & researchers, and encourage you to join the organization. independenttechresearch.org

Home - Coalition for Independent Technology Research

independenttechresearch.org

J. Nathan Matias@natematias.bsky.social · 2mo ago

I hope folks have a great time at #ic2s2 this week! So say hi to @davidlazer.bsky.social if you're interested in thinking about how to make computational social science more repeatable, and feel free to complain to me about any errors in our work: citizensandtech.org/2026/07/comm...

Attending #IC2S2!! Will be co-organizing a tutorial on Simulating Human Survey Responses with Large Language Models, today at 1.15PM, and presenting on Friday in the Understanding Large Language Models session, 10.45AM. Looking forward to catching up with old friends and meeting new ones!!

Bild

❓❓Should LLMs have access to ACM’s Digital Library❓❓ I wrote a short essay arguing that "yes!" I try to argue that, even if you are concerned about their negative impact on science, denying them access is a poor way to address their harm! doomscrollingbabel.manoel.xyz/p/science-sh...

Science Should Be Open, For LLMs Too

Even if you think LLMs are bad for science, you should let them have access to research papers

doomscrollingbabel.manoel.xyz

Not a lot of info here about what kinds of agreements, with whom, and finances. But another example of data opening up for “AI” that wasn’t open for researchers. I fought many data wrangling battles at Semantic Scholar over missing ACM data. All research papers should be open access!

Association for Computing Machinery@acm.org · 3mo ago

We are opening a consultation on the inclusion of ACM publications within AI licensing agreements for access and training purposes. We have published this piece as to why we believe it is the right time now to enter into these discussions. buff.ly/VTatuda

Just back from vacation, and very hyped for today's Brazil vs. Japan game. As a Brazilian, one of the most endearing things is seeing other countries (esp. other developing ones) show affection for our team! We don't deserve you all! Below is the least I expect from today's game :-)

New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!

Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
Teagan Johnson@teagrjohnson.bsky.social · 4mo ago

1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:

This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.