@moritzplenz.bsky.social

How and when do multilingual LMs achieve cross-lingual generalization during pre-training? And why do later, supposedly more advanced checkpoints, lose some language identification abilities in the process? Our #ACL2025 paper investigates.

Probing classifier performance comparison between early and late checkpoint across layers. While the early checkpoint shows uniformly high performance, the later checkpoint exhibits relatively high variance across layers.

Debates aren’t always black and white—opposing sides often share common ground. These partial agreements are key for meaningful compromises Presenting “Perspectivized Stance Vectors” (PSVs) — an interpretable method to identify nuanced (dis)agreements 📜 arxiv.org/abs/2502.09644 🧵 More details below

Bild