Jessy Li

@jessyjli.bsky.social

https://jessyli.com Associate Professor, UT Austin Linguistics. Part of UT Computational Linguistics https://sites.utexas.edu/compling/ and UT NLP https://www.nlp.utexas.edu/

The full BBS treatment from me and @futrell.bsky.social on "How linguistics learned to stop worrying and love the LMs" is now out, with all the commentaries and our response. If you "Save PDF", it will give you the whole target article + commentary + response pdf: www.cambridge.org/core/journal...

How linguistics learned to stop worrying and love the language models | Behavioral and Brain Sciences | Cambridge Core

How linguistics learned to stop worrying and love the language models - Volume 49

cambridge.org

I'm on a new committee reviewing ARR's Responsible NLP Checklist and looking to potentially make changes. I'd love to hear others' thoughts on what is working well or needs revised, especially given it might be fresh in memory from the recent ARR cycle. 1/

We are thrilled to present a detailed report describing the system built for the AAAI-26 AI review pilot, the survey results, and a new benchmark that was created to assess the capabilities of the system. Read the full article: arxiv.org/pdf/2604.13940

arxiv.org

New opinion piece on the interface between research on concepts and categories in minds vs. in neural network LMs! I take the position that there is much to be learned from this interface (e.g., learning about concepts from language alone) and outline some directions for future.

Title page of "Semantic Cognition for and from Language Models" followed by a figure showing tests that target conceptual structure and content vs. those that target function.

QUDs going multimodal! With MQUD, we can train models to generate scientific questions❓that are inquisitive and insightful enough to be answered in the scientific paper! Huge thanks to the many paper authors who contributed to our data. Check out @yatingwu.bsky.social’s work:

Yating Wu@yatingwu.bsky.social · 3mo ago

What does a scientific figure make you wonder? 📊 We introduce MQUD: multimodal Questions Under Discussion for scientific figures. With 1,250 author-annotated questions over 245 figures from 56 papers, MQUD asks what scientific question a figure raises in context.

Nice research! You may be interested in the small scale ($4 budget) verification performed by my personal Opus agent here: muninn.austegard.com/blog/this-tr... in which we also introduced a framing-resistant prompt to see how much that would mitigate the effetcs. 1/3

This Treatment Works, Right? Testing Framing Resistance in Medical QA

A rapid replication testing whether a framing-resistant prompt can mitigate LLM sensitivity to question phrasing in medical contexts.

muninn.austegard.com

Check out @asher-zheng.bsky.social's work on quantifying strategic language in dialogue, just appeared in the Dialogue and Discourse journal. We study non-cooperative moves that are subtle to capture, where modern AI still have trouble comprehending. Work w/ David_Beaver

BildBild
Asher Zheng@asher-zheng.bsky.social · 6mo ago

Very excited to share that the paper w/ @jessyjli.bsky.social @DavidBeaver "Strategic Dialogue Assessment: The Crooked Path to Innocence" (used to have the name COBRA) was accepted by Dialogue and Discourse Vol 17 No.1. Check it out! 👉https://journals.uic.edu/ojs/index.php/dad/article/view/14503

“All bears have a property”, “Some bears have a property”, “Bears have a property” are different in terms of how the property is generalized to a specific bear – a great example of how language constrains thought! This holds for kids, adults, and according to our new work, (V)LMs! 🧵

Title page of our paper: "Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences"

New work to appear @ TACL! Language models (LMs) are remarkably good at generating novel well-formed sentences, leading to claims that they have mastered grammar. Yet they often assign higher probability to ungrammatical strings than to grammatical strings. How can both things be true? 🧵👇

Screenshot of a figure with two panels, labeled (a) and (b). The caption reads: "Figure 1: (a) Illustration of messages (left) and strings (right) in toy domain. Blue = grammatical strings. Red = ungrammatical strings. (b) Surprisal (negative log probability) assigned to toy strings by GPT-2."

Delighted Sasha's (first year PhD!) work using mech interp to study complex syntax constructions won an Outstanding Paper Award at EMNLP! Also delighted the ACL community continues to recognize unabashedly linguistic topics like filler-gaps... and the huge potential for LMs to inform such topics!

aclanthology.org

Sasha Boguraev@sashaboguraev.bsky.social · last yr.

A key hypothesis in the history of linguistics is that different constructions share underlying structure. We take advantage of recent advances in mechanistic interpretability to test this hypothesis in Language Models. New work with @kmahowald.bsky.social and @cgpotts.bsky.social! 🧵👇!