Lucy Li

@lucy3.bsky.social

assistant prof at Wisconsin CS. #NLProc, computational social science, cultural analytics, responsible AI. she/her. Previously at UW, Berkeley, Ai2, MSR, Stanford. lucy3.github.io

Starting off my first lecture later this week as a prof with a little history of NLP! Warren Weaver (known for writing a memorandum that ignited machine translation in 1949, before "artificial intelligence" was even coined as a term) did his BS and PhD at UW-Madison 🦡🧀

Spent a good amount of this summer digging through a 20+ person team of teachers' free-text annotations, learning what "scaffolding" and "push for rigor" means, and iterating on this pipeline. We find that AI tutors, by default, frequently over-scaffold and rarely push for rigor.

Ai2@ai2.bsky.social · last mo.

Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵

I used to joke that going to Stanford for college w/ palm trees and manicured landscapes was like being on vacation with exams -- summers at UW-Madison have a similar feel. Some evenings are so pretty that these ducks must be paid actors to decorate these lakes also the ice cream is very creamy

ducks and a dock and a sunset on lake mendota

Wrote a brief reflection on poor conceptualizations in AI work and why they matter, something I believe is so startlingly neglected. TL;DR - Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims.

Back-to-basics: on poor conceptualizations in AI work

TL;DR — Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims. (And, no, your metric is not your construct.)

rigor-in-ai.leaflet.pub

I have two projects involving Chinese speakers this year, and so I'm saying "I'm going to ask my mom" in project meetings more often than usual

me: "please draw a wisconsin badger hugging a jar of hazelnut spread except instead of the logo being nutella because it's trademarked it is the logo of my new research group" human 👨‍🎨:

concept art for a badger hugging a jar with a rough sketch of the wisconsin social language and ai logo as the brand label

New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!

Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
Teagan Johnson@teagrjohnson.bsky.social · 3mo ago

1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:

This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.