Emily Cheng

@emcheng.bsky.social

https://generalstrikeus.com/ PhD student in computational linguistics at UPF chengemily1.github.io Previously: MIT CSAIL, ENS Paris Barcelona

Presenting this at #ICML with @rjantonello.bsky.social and Aditya Vaidya✨ Why do 𝙢𝙞𝙙𝙙𝙡𝙚 layers in LLMs and speech-audio models best predict brain responses to language? We show a peak in the dimensionality of 🤖 activations (left) to track high 🧠 predictivity (right) 🧵(cross-posted from X)

Bild

Our paper "Prediction Hubs are Context-Informed Frequent Tokens in LLMs" has been accepted at ACL 2025! Main points: 1. Hubness is not a problem when language models do next-token prediction. 2. Nuisance hubness can appear when other comparisons are made.

Bild

Last day to sign up for the COLT Symposium! Register: tinyurl.com/colt-register 📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 June 2nd, 14:30 - 19:00 UPF Campus de la Ciutadella Room 40.101 maps.app.goo.gl/1216LJRsWmTE...

Computational Linguistics @UPF@colt-upf.bsky.social · last yr.

⭐ Registration open til May 27th! ⭐ Website: www.upf.edu/web/colt/sym... June 2nd, UPF 𝗦𝗽𝗲𝗮𝗸𝗲𝗿 𝗹𝗶𝗻𝗲𝘂𝗽: Arianna Bisazza (language acquisition with NNs) Naomi Saphra (emergence in LLM training dynamics) Jean-Rémi King (TBD) Louise McNally (pitfalls of contextual/formal accounts of semantics)

🧵 Excited to share our paper "Unique Hard Attention: A Tale of Two Sides" with Selim, Jiaoda, and Ryan, where we show that the way transformers break ties in attention scores has profound implications on their expressivity! And it got accepted to ACL! :) The paper: arxiv.org/abs/2503.14615

Unique Hard Attention: A Tale of Two Sides

Understanding the expressive power of transformers has recently attracted attention, as it offers insights into their abilities and limitations. Many studies analyze unique hard attention transformers...

arxiv.org

🌍📣🥳 I could not be more excited for this to be out! With a fully automated pipeline based on Universal Dependencies, 43 non-Indoeuropean languages, and the best LLMs only scoring 90.2%, I hope this will be a challenging and interesting benchmark for multilingual NLP. Go test your language models!

Jaap Jumelet@jumelet.bsky.social · last yr.

✨New paper ✨ Introducing 🌍MultiBLiMP 1.0: A Massively Multilingual Benchmark of Minimal Pairs for Subject-Verb Agreement, covering 101 languages! We present over 125,000 minimal pairs and evaluate 17 LLMs, finding that support is still lacking for many languages. 🧵⬇️

The DOGE firings have nothing to do with “efficiency” or “cutting waste.” They’re a direct push to weaken federal agencies perceived as liberal. This was evident from the start, and now the data confirms it: targeted agencies overwhelmingly those seen as more left-leaning. 🧵⬇️

Scatterplot titled “Empirical Evidence of Ideological Targeting in Federal Layoffs: Agencies seen as liberal are significantly more likely to face DOGE layoffs.”
	•	The x-axis represents Perceived Ideological Leaning of federal agencies, ranging from -2 (Most Liberal) to +2 (Most Conservative), based on survey responses from over 1,500 federal executives.
	•	The y-axis shows Agency Size (Number of Staff) on a logarithmic scale from 1,000 to 1,000,000.

Each point represents a federal agency:
	•	Red dots indicate agencies that experienced DOGE layoffs.
	•	Gray dots indicate agencies with no layoffs.

Key Observations:
	•	Liberal-leaning agencies (left side of the plot) are disproportionately represented among red dots, indicating higher layoff rates.
	•	Notable targeted agencies include:
	•	HHS (Health & Human Services)
	•	EPA (Environmental Protection Agency)
	•	NIH (National Institutes of Health)
	•	CFPB (Consumer Financial Protection Bureau)
	•	Dept. of Education
	•	USAID (U.S. Agency for International Development)
	•	The National Nuclear Security Administration (DOE), despite its conservative leaning (+1 on the scale), is an exception among targeted agencies.
	•	A notable outlier: the Department of Veterans Affairs (moderately conservative) also faced layoffs despite its size.

Takeaway:

The figure visually demonstrates that DOGE layoffs disproportionately targeted liberal-leaning agencies, supporting claims of ideological bias. The pattern reveals that layoffs were not driven by agency size or budget alone but were strongly associated with perceived ideology.

Source: Richardson, Clinton, & Lewis (2018). Elite Perceptions of Agency Ideology and Workforce Skill. The Journal of Politics, 80(1).

🚨BREAKING. From a program officer at the National Science Foundation, a list of keywords that can cause a grant to be pulled. I will be sharing screenshots of these keywords along with a decision tree. Please share widely. This is a crisis for academic freedom & science.

list of banned keywords

Here's our work accepted to #ICLR2025! We look at how intrinsic dimension evolves over LLM layers, spotting a universal high-dimensional phase. This ID peak is where: - linguistic features are built - different LLMs are most similar, with implications for task transfer 🧵 1/6

Bild

I think some people hear “grants” and think that without them, scientists and government workers just have less stuff to play with at work. But grants fund salaries for students, academics, researchers, and people who work in all areas of public service. “Pausing” grants means people don’t eat.

White House pauses all federal grants, sparking confusion

The Trump administration has put a hold on all federal financial grants and loans, affecting tens of billions of dollars in payments.

washingtonpost.com

⚡Postdoc opportunity w/ COLT Beatriu de Pinós contract, 3 yrs, competitive call by Catalan government. Apply with a PI (Marco Gemma or Thomas) Reqs: min 2y postdoc experience outside Spain, not having lived in Spain for >12 months in the last 3y. Application ~December-February (exact dates TBD)

Hello🌍! We're a computational linguistics group in Barcelona headed by Gemma Boleda, Marco Baroni & Thomas Brochhagen We do psycholinguistics, cogsci, language evolution & NLP, with diverse backgrounds in philosophy, formal linguistics, CS & physics Get in touch for postdoc, PhD & MS openings!