Can we map out gaps in LLMs’ cultural knowledge? Check out our #EMNLP2025 talk: Culture Cartography 🗓️ 11/5, 11:30 AM 📌 A109 (CSS Orals 1) Compared to traditional benchmarking, our mixed-initiative method finds more knowledge gaps even in reasoning models like R1! Paper: arxiv.org/pdf/2510.27672
Caleb Ziems
@calebziems.com
PhD student at Stanford NLP. Working on Social NLP and CSS. Previously at GaTech, Meta AI, Emory. 📍Palo Alto, CA 🔗 calebziems.com
AI always calling your ideas “fantastic” can feel inauthentic, but what are sycophancy’s deeper harms? We find that in the common use case of seeking AI advice on interpersonal situations—specifically conflicts—sycophancy makes people feel more right & less willing to apologize.
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
AI companions aren’t science fiction anymore 🤖💬❤️ Thousands are turning to AI chatbots for emotional connection – finding comfort, sharing secrets, and even falling in love. But as AI companionship grows, the line between real and artificial relationships blurs.
Introducing CAVA: The Comprehensive Assessment for Voice Assistants A new benchmark for evaluating the capabilities required for speech-in-speech-out voice assistants! - Latency - Instruction following - Function calling - Tone awareness - Turn taking - Audio Safety TalkArena.org/cava
Comprehensive Assessment for Voice Assistants
CAVA is a new benchmark for assessing how well Large Audio Models support voice assistant capabilities.
talkarena.org
Reward models for LMs are meant to align outputs with human preferences—but do they accidentally encode dialect biases? 🤔 Excited to share our paper on biases against African American Language in reward models, accepted to #NAACL2025 Findings! 🎉 Paper: arxiv.org/abs/2502.12858 (1/10)
EgoNormia (egonormia.org) exposes a major gap in Vision-Language Models understanding of the social world: they don't know how to behave when norms about the physical world *conflict* ⚔️ (<45% acc.) But humans are naturally quite good at this (>90% acc.) Check it out! ➡️ arxiv.org/abs/2502.20490
EgoNormia: A Benchmark for Visual Frontier Models' Normative Reasoning
A large scale video dataset and a benchmark for evaluating frontier models' understanding of physical social norms through videos.
egonormia.org
We are getting closer to have agents operating in the real physical world. However, can we trust frontier models to make embodied decisions 🎮 aligned with human norms 👩⚖️ ? With EgoNormia, a 1.8k ego-centric video 🥽 QA benchmark, we show that this is surprisingly challenging!
There's been a lot of work on "culture" in NLP, but not much agreement on what it is. A position paper by me, @dbamman.bsky.social, and @ibleaman.bsky.social on cultural NLP: what we want, what we have, and how sociocultural linguistics can clarify things. Website: naitian.org/culture-not-... 1/n
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)
Bill Labov died this morning. I'm not coherent enough to talk about how important and influential and brilliant he was. I am very sad. I was so lucky to know him, and I am grateful every day that he (and Gillian, and Walt, etc) built an academic field where kindness is expected.
With an increasing number of Large *Audio* Models 🔊, which one do users like the most? Introducing talkarena.org — an open platform where users speak to LAMs and receive text responses. Through open interaction, we focus on rankings based on user preferences rather than static benchmarks. 🧵 (1/5)
Maybe some starter packs for the Dyirbal noun classes? 1. most animate objects, men 2. women, water, fire, violence, and exceptional animals 3. edible fruit and vegetables 4. miscellaneous (includes things not classifiable in the first three)
Hi Bluesky! You get to be the very first internet people to see my standup comedy debut. Because I know you’ll be nicer to me than the 12 year olds on TikTok. youtu.be/KqL2ahOvAgg?...
AI is not the GOAT. (Uh oh, your professor is attempting stand up comedy.)
YouTube video by Casey Fiesler
youtu.be
I noticed a lot of starter packs skewed towards faculty/industry, so I made one of just NLP & ML students: go.bsky.app/vju2ux Students do different research, go on the job market, and recruit other students. Ping me and I'll add you!
I'm recruiting 1-2 PhD students to work with me at the University of Colorado Boulder! Looking for creative students with interests in #NLP and #CulturalAnalytics. Boulder is a lovely college town 30 minutes from Denver and 1 hour from Rocky Mountain National Park 😎 Apply by December 15th!
Repost if you’ve participated in a Summer Institute in Computational Social Science. Let’s get #SICSS Bluesky going!
I'm sharing materials from my academic job search last year! Includes research, teaching, and diversity statements, plus my UMD cover letter and job talk slides. I applied for a mix of iSchool, data sci, CS, and linguistics positions). Feel free to share! juliamendelsohn.github.io/resources/
resources | Julia Mendelsohn
Materials that some people might find helpful
juliamendelsohn.github.io
All the ACL chapters are here now: @aaclmeeting.bsky.social @emnlpmeeting.bsky.social @eaclmeeting.bsky.social @naaclmeeting.bsky.social #NLProc
I wanted to contribute to "Starter Pack Season" with one for Stanford NLP+HCI: go.bsky.app/VZBhuJ5 Here are some other great starter packs: - CSS: go.bsky.app/GoEyD7d + go.bsky.app/CYmRvcK - NLP: go.bsky.app/SngwGeS + go.bsky.app/JgneRQk - HCI: go.bsky.app/p3TLwt - Women in AI: go.bsky.app/LaGDpqg
Ready for another Computational Social Science Starter Pack? Here is number 2! More amazing folks to follow! Many students and the next gen represented! go.bsky.app/GoEyD7d
🤖🧠 I'll be considering applications for postdocs & PhD students to start at Yale in Fall 2025! If you are interested in the intersection of linguistics, cognitive science, and AI, I encourage you to apply! Postdoc link: rtmccoy.com/prospective_... PhD link: rtmccoy.com/prospective_...
Hello! I'm an internet linguist! I wrote a book called Because Internet about how we use language online gretchenmcculloch.com/book I make @lingthusiasm.bsky.social, a podcast that's enthusiastic about linguistics And I maintain a linguistics starter pack here: go.bsky.app/UUM7Gcx
The AI Interdisciplinary Institute at the University of Maryland (AIM) is hiring 40 new faculty members in all areas of AI, particularly: - accessibility, - sustainability, - social justice, and - learning; building on computational, humanistic, or social scientific approaches to AI. >
🎓 Fully funded PhD Fellowship in Interpretable NLP at the University of Copenhagen & Pioneer Centre for AI available! 📆 Application deadline: 15 Jan 2025 👥 Supervisors: Pepa Atanasova & me 🤝 Reasons to apply: www.copenlu.com/post/why-ucph/ 📝 Apply here: employment.ku.dk/phd/?show=16... #NLProc #XAI
Papers at #EMNLP2024 #1 Statistical Uncertainty in Word Embeddings: GloVe-V Neural models, from word vectors through transformers, use point estimate representations. They can have large variances, which often loom large in CSS applications. Tue Nov 12 15:15-15:30 Flagler
I guess I should cross post here too. I'm recruiting one (1) PhD student to work on multimodal embodied agents. VLMs + RL. Please apply to the UCSD CSE app by Dec 15 and mark my name as a faculty of interest. More lab info at pearls.ucsd.edu
📣 I am recruiting 1-2 PhD students for Fall 2025 at the University of Maryland College of Information. Consider applying if you're interested in language, society/politics, and computers! Deadline Dec 3: ischool.umd.edu/academics/ph... And pls share with anyone who may be interested!
Doctor of Philosophy in Information Studies (PhD) - College of Information (INFO)
This doctoral program prepares students to address the hardest social and technical problems of today and tomorrow.
ischool.umd.edu
Academic job market post! 👀 I’m a CS Postdoc at Stanford in the Stanford HCI group. I develop ways to improve the online information ecosystem by designing better social media feeds & improving Wikipedia. I work on AI, Social Computing, and HCI. piccardi.me 🧵
Our paper was accepted to EMNLP findings 😌 Can’t wait to talk about it at EMNLP and meet old friends and make new ones!
What differentiates in-group speech from out-group speech? I've been pondering this question for most of my PhD, and the final chapter of my dissertation tackles this question in a super interesting domain: comments from NFL🏈 team subreddits on live game threads. Our insight was to frame...🧵[1/7]