Computational Linguistics @UPF

@colt-upf.bsky.social

Gemma Boleda, Marco Baroni, Thomas Brochhagen, Iria de Dios Flores | Computational Linguistics and Linguistic Theory Universitat Pompeu Fabra. upf.edu/web/colt Barcelona

Presenting this at #ICML with @rjantonello.bsky.social and Aditya Vaidya✨ Why do 𝙢𝙞𝙙𝙙𝙡𝙚 layers in LLMs and speech-audio models best predict brain responses to language? We show a peak in the dimensionality of 🤖 activations (left) to track high 🧠 predictivity (right) 🧵(cross-posted from X)

Bild

Many forces have been argued to shape natural language lexica, and there are different ways they can be operationalized and interact. We study which out of a set of forces and their interactions best fit cross-linguistic data. Now out in Cognitive Science: onlinelibrary.wiley.com/doi/10.1111/...

Assessing Pressures Shaping Natural Language Lexica

Human languages balance communicative informativity with complexity, conveying as much as needed through the simplest means required to do so. Yet, these concepts—informativity and complexity—have be...

onlinelibrary.wiley.com

Evaluating topic models (and document clustering methods) is hard. In fact, since our paper critiquing standard evaluation practices four years ago, there hasn't been a good replacement metric That ends today (we hope)! Our new ACL paper introduces an LLM-based evaluation protocol 🧵

Screenshot of first page of paper. It is here: https://arxiv.org/pdf/2507.00828

Abstract: Topic model and document-clustering evaluations either use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. We design a scalable human evaluation protocol and a corresponding automated approximation that reflect practitioners' real-world usage of models. Annotators -- or an LLM-based proxy -- review text items assigned to a topic or cluster, infer a category for the group, then apply that category to other documents. Using this protocol, we collect extensive crowdworker annotations of outputs from a diverse set of topic models on two datasets. We then use these annotations to validate automated proxies, finding that the best LLM proxies are statistically indistinguishable from a human annotator and can therefore serve as a reasonable substitute in automated evaluations

Today at UPF Campus de la Ciutadella at 2:30 pm! Come slightly earlier to check in! Sala Polivalent 24S18 maps.app.goo.gl/n1hBxiviKcLW...

Computational Linguistics @UPF@colt-upf.bsky.social · last yr.

⭐ Registration open til May 27th! ⭐ Website: www.upf.edu/web/colt/sym... June 2nd, UPF 𝗦𝗽𝗲𝗮𝗸𝗲𝗿 𝗹𝗶𝗻𝗲𝘂𝗽: Arianna Bisazza (language acquisition with NNs) Naomi Saphra (emergence in LLM training dynamics) Jean-Rémi King (TBD) Louise McNally (pitfalls of contextual/formal accounts of semantics)

📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 UPF Campus de la Ciutadella **Sala Polivalent 24.S18** Thank you for bearing with us!

Computational Linguistics @UPF@colt-upf.bsky.social · last yr.

Last day to sign up for the COLT Symposium! Register: tinyurl.com/colt-register 📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 June 2nd, 14:30 - 19:00 UPF Campus de la Ciutadella Room 40.101 maps.app.goo.gl/1216LJRsWmTE...

Last day to sign up for the COLT Symposium! Register: tinyurl.com/colt-register 📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 June 2nd, 14:30 - 19:00 UPF Campus de la Ciutadella Room 40.101 maps.app.goo.gl/1216LJRsWmTE...

Computational Linguistics @UPF@colt-upf.bsky.social · last yr.

⭐ Registration open til May 27th! ⭐ Website: www.upf.edu/web/colt/sym... June 2nd, UPF 𝗦𝗽𝗲𝗮𝗸𝗲𝗿 𝗹𝗶𝗻𝗲𝘂𝗽: Arianna Bisazza (language acquisition with NNs) Naomi Saphra (emergence in LLM training dynamics) Jean-Rémi King (TBD) Louise McNally (pitfalls of contextual/formal accounts of semantics)

Please find us at #ICLR2025! We will present our work on intrinsic dimension as a cue for stages of language processing in LLMs. Saturday morning, Poster session 5 Hall 3 + Hall2B #563 iclr.cc/virtual/2025... Arxiv: arxiv.org/abs/2405.15471

Bild
Emily Cheng@emcheng.bsky.social · 2y ago

Here's our work accepted to #ICLR2025! We look at how intrinsic dimension evolves over LLM layers, spotting a universal high-dimensional phase. This ID peak is where: - linguistic features are built - different LLMs are most similar, with implications for task transfer 🧵 1/6

📢 Upcoming Seminar Words are weird? On the role of lexical ambiguity in language 🗣 Gemma Boleda (Universitat Pompeu Fabra, Spain) Why is language so ambiguous? Discover how ambiguity balances cognitive simplicity and communicative complexity through large-scale studies. 📍 UniMiB, Room U6-01C, Milan

Bild

Here's our work accepted to #ICLR2025! We look at how intrinsic dimension evolves over LLM layers, spotting a universal high-dimensional phase. This ID peak is where: - linguistic features are built - different LLMs are most similar, with implications for task transfer 🧵 1/6

Bild

Hello🌍! We're a computational linguistics group in Barcelona headed by Gemma Boleda, Marco Baroni & Thomas Brochhagen We do psycholinguistics, cogsci, language evolution & NLP, with diverse backgrounds in philosophy, formal linguistics, CS & physics Get in touch for postdoc, PhD & MS openings!

⚡Postdoc opportunity w/ COLT Beatriu de Pinós contract, 3 yrs, competitive call by Catalan government. Apply with a PI (Marco Gemma or Thomas) Reqs: min 2y postdoc experience outside Spain, not having lived in Spain for >12 months in the last 3y. Application ~December-February (exact dates TBD)