Lucy Li

@lucy3.bsky.social

Postdoc at UW NLP 🏔️. #NLProc, computational social science, cultural analytics, responsible AI. she/her. Previously at Berkeley, Ai2, MSR, Stanford. Incoming assistant prof at Wisconsin CS. lucy3.github.io

I used to joke that going to Stanford for college w/ palm trees and manicured landscapes was like being on vacation with exams -- summers at UW-Madison have a similar feel. Some evenings are so pretty that these ducks must be paid actors to decorate these lakes also the ice cream is very creamy

ducks and a dock and a sunset on lake mendota

Wrote a brief reflection on poor conceptualizations in AI work and why they matter, something I believe is so startlingly neglected. TL;DR - Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims.

Back-to-basics: on poor conceptualizations in AI work

TL;DR — Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims. (And, no, your metric is not your construct.)

rigor-in-ai.leaflet.pub

I have two projects involving Chinese speakers this year, and so I'm saying "I'm going to ask my mom" in project meetings more often than usual

me: "please draw a wisconsin badger hugging a jar of hazelnut spread except instead of the logo being nutella because it's trademarked it is the logo of my new research group" human 👨‍🎨:

concept art for a badger hugging a jar with a rough sketch of the wisconsin social language and ai logo as the brand label

New work from my lab! @teagrjohnson.bsky.social built a 12-dimensional narrative framework, annotated Dolma (no small feat given its extreme diversity), and analyzed narrative features across pretraining subsections. Highlight: pretraining data space displays strong narrative organization!

Figure 1: A web passage scored across our 12 narrative dimensions. Agency and setting dimensions are rated on a 1–5 Likert scale, temporal sequencing and causal density are passage-level proportions (0–1), and event density is the rate of event triggers per token. This passage scores high on agency and event features but low on setting, a “narrative profile” commonly seen across first-person web narratives.Figure 6: UMAP reduction of SBERT embeddings for 20,000 randomly sampled NARRADOLMA documents, colored by PC1 score (interiority). Labels are based on manual examination. Overlays for all three PCs appear in Fig. A5.
Teagan Johnson@teagrjohnson.bsky.social · 2mo ago

1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:

This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.

Looking for a emergency reviewer for a paper related to human-AI interaction/cooperation for ACL ARR (we've been iterating on getting replacement reviewers for this paper for awhile now and fingers crossed) -- if you're able to help me out with this, lmk. Happy to return the favor in the future.

What are the worst ethical disasters in NLP history? (I'm teaching "ethics of NLP" tomorrow and history is good for teaching this topic.) Most are data breaches/releases (AOL search logs, OKCupid profiles, Finnish therapy records...) but what others? I'll put some other examples in thread --> 1/n