Lauren Klein

@laurenfklein.bsky.social

Digital humanities, data science, AI, eating, professor of Data & Decision Science and English. Coauthor #DataFeminism w/ @kanarinka. PI #AIAInetwork. Views my own.

Please join us at Carnegie Mellon University and online on October 2nd, 2026 for the Community-Centered AI Infrastructure Summit, a statewide convening to consider the opportunities, challenges, and governance of AI infrastructure development in Pennsylvania. RSVP here: communityaisummit.org

Flyer for the Community-Centered AI Infrastructure Summit on Friday October 2, 2026 from 9-5pm at Mellon Institute Auditorium, Carnegie Mellon University, Pittsburgh, PA. Join residents, community organizations, policymakers, researchers, technologists, energy experts, and advocates from across Pennslyvania for a public convening. Organized by CMU, Data & Society, PA Climate Equity Table, and Center for Coalfield Justice, with support from the Block Center.

Today, @wired.com is launching our latest issue: WIRED Women. This issue started at an office happy hour. What if our publication, which spent years putting powerful men on its covers, made an entire edition about women? We thought, we talked, we laughed. We got mad. And then we went for it...

Prospective PhD students: if you work on computational humanities, cultural analytics and/or cultural AI, pls apply to Duke English! We're recruiting in these areas & you'd join a vibrant community. Candidates with training in both lit studies + CS/sciences esp welcome. Email me if any questions!

The World Bank estimates that there are between 150 million to 430 million data workers in the planet. The US population is 342 million people. Let that sink in.

New from our lab! #COLM2026 When people generate stories, they don't just write one prompt. Instead, they explore narrative space via branching edits 🌱 We reconstruct 24k of these edit trees 🌳 from chat logs and map edit types, story formats, how they relate to tree depth, and more!

The Garden of Forking Prompts: How Users Explore Narrative
Space in Story Generation
Advait Deshmukh♣ Nora Benedict♠ Melanie Walsh♡ Maria Antoniak♣
♣University of Colorado Boulder ♠University of Georgia ♡University of Washington

Abstract

Large language models (LLMs) have changed the way people engage
with stories. Drawing on public chatbot logs, we can see that when users
generate stories, they iteratively edit their prompts to explore narrative
possibilities, adjusting characters, redirecting plots, and swapping fictional universes. As aggregated data, these prompts represent rich traces of creative preference at scale. Yet story generation evaluation benchmarks rely on static, one-shot prompts that cannot capture this exploratory behavior. In this work, we study how users revise consecutive story prompts in the wild. Using a dataset of naturally occurring user-chatbot conversations, we construct WildStories, a sample of 275,635 story generation prompts (labeled with story format, prompt components, and explicitness), and WildEdits, a collection of 24,291 edit trees that model how users iteratively edit base story prompts and explore branching story possibilities. From these trees we develop a framework of edit types crossing four directions (adding, removing, changing, and extending) with fourteen targets (e.g., plot, character, genre). We then use our datasets and this framework to
analyze user behavior in navigating narrative space via LLMs. Finally, we
show how automated permutations based on the framework can be used
for story generation benchmarking. Content Warning: This paper works with “wild” chatbot logs, which often include toxic and sexually explicit themes.
Advait Deshmukh@advaitdeshmukh.com · 2w ago

1/7 In 1941, Borges imagined an impossible novel that follows all narrative branches at once. Today, we can observe chatbot users exploring narrative possibilities as they repeatedly edit story prompts. In our COLM 2026 paper, we studied this branching exploration via the WildChat dataset. 🧵

Paraphrased edit trees from the dataset

A good excuse to preorder books! B&N is offering 25% off all preorders for Rewards Members through Sept. 11 (+10% for Premium Members). That includes The War Over Rent, Data by Design, and The Father of the Man. Preorders matter to authors. Code: PREORDER25 ❤️📚

Bild

🎂 SAVE THE DATE for Feb 11, 2027! 🎂 Douglass Day will feature Transcribe the Early Black Press! 💻 New Crowdsourcing Project 📽️ Live Broadcast 🍰 Great Douglass Day Bake Off 📣 Special guests! Host an event or join us online! Help us spread the word? Register now: DougalssDay.org

Postcard graphic with large text saying "Save the Date for Douglass Day" Below that, text reads "Celebrating the 200th Anniversary of the Black Press" and the date "2.11.2027" followed by the web address DouglassDay.org.

NB: The first lesson is about the data center in Virginia that is hosting our cloud compute, and the last one is about the actual people who do RLHF for OpenAI and Meta (our class will be doing its own). Along the way, the core NLP techniques that it turns out LLMs depend upon.

Lauren Klein@laurenfklein.bsky.social · last mo.

After far too many AI-generated final projects, and my own uncertainty about what students really need to know about NLP these days, I decided to overhaul my Text as Data class so that it's structured around a single project: training up a C19 LLM from scratch. Here we go! github.com/laurenfklein...