Desmond Elliott

@delliott.bsky.social

Regular reminder that we have an alt-ARR slack workspace where ACs and SACs can support each other through the sometimes confusing process of the ARR cycle! Post or DM me a good email address for a Slack invitation and I will add you. #EMNLP2026

⚠️ Submitting to #EMNLP2026? Make sure to review our newly published Paper Integrity Policy first! It includes important updates on Generative Assistance in Authorship, thinly sliced contributions, and unverifiable references. 🔗 Details: 2026.emnlp.org/paper-integr...

EMNLP 2026 Paper Integrity Policy

EMNLP 2026, together with ARR, are taking actions against growing concerns of irresponsible AI use in paper submissions that take up valuable reviewer/AC resources and make little contributions to the...

2026.emnlp.org

Submitting to ARR for #EMNLP2026? We're running an opt-in AI Reviewing Experiment. Help us test AI-generated reviews during your ARR submission. 🤖 ✅ Reviewers, ACs, and SACs will not be able to see it ✅ Will not affect decisions 🔗 Read more: 2026.emnlp.org/ai-reviewing...

EMNLP 2026 AI Reviewing Experiment

EMNLP 2026 is running an AI Reviewing Experiment to collect feedback from authors about the quality of AI reviews of their submissions. This experiment is taking place on an opt-in basis, in which aut...

2026.emnlp.org

I have been thinking about some of the consequences of closed vs open research. Closed research can slow down scientific progress and concentrate knowledge, which results in what I call “model archaeology”. I discuss this idea in my ICLR 2026 Blogpost. Short thread 🧵and link👇

🚨New paper Are visual tokens going into an LLM interpretable 🤔 Existing methods (e.g. logit lens) and assumptions would lead you to think “not much”... We propose LatentLens and show that most visual tokens are interpretable across *all* layers 💡 Details 🧵

Bild

💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR. But the set of constraints and verifier functions is limited and most models overfit on IFEval. We introduce IFBench to measure model generalization to unseen constraints.

Bild

Following #CVPR2025, #ICCV2025 implemented a new policy targeting accountability and integrity. PCs identified 25 highly irresponsible reviewers, resulting in the desk rejection of 29 associated papers, including 12 submissions that otherwise would have been accepted.

Huge thanks to everyone that attended the Copenhagen NLP Symposium last week. Thanks for our wonderful speakers @kylelo.bsky.social, @najoung.bsky.social, Yohei Oseki, @mziizm.bsky.social, and @loubnabnl.hf.co! @mariaa.bsky.social did a great job of summarizing the talks in these liveposts (quoted).

People finding their seats before the event started
Maria Antoniak@mariaa.bsky.social · last yr.

Finally, @kylelo.bsky.social on the Olmo Cookbook! Pretraining data curation practices, challenges and frustrations, practical advice for your data curation. Starts with an overview: data acquisition, transformation, experimentation

Announcing our recent work “Multilingual Pretraining for Pixel Language Models”! We introduce PIXEL-M4, a pixel language model pretrained on four visually & linguistically diverse scripts: English, Hindi, Ukrainian & Simplified Chinese. #NLProc

Bild

I am excited to announce our latest work 🎉 "Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory". We review recent works on culture in VLMs and argue for deeper grounding in cultural theory to enable more inclusive evaluations. Paper 🔗: arxiv.org/pdf/2505.22793

Paper title "Cultural Evaluations of Vision-Language Models
Have a Lot to Learn from Cultural Theory"

Has anyone written anything about *scraping and text processing* for internet pretraining data? Practical details, which tools are used, which webpage elements are considered, how HTML to text conversion is done? (I know about work on quality filters, relevant but not quite what I'm looking for)

Had fun talking at the Spurious Correlations & Shortcut Learning at ICLR! One example I brought up, which I think provides an uncommon perspective: a case where spurious shortcuts can improve generalization... even to out-of-distribution sets where the spurious feature doesn't generalize! Thread:

I'm recruiting a postdoc on an 18-month contract candidate.hr-manager.net/ApplicationI.... The position is about deploying LLMs in the Danish public sector. This is an interdisciplinary project that touches on technical, ethical, and legal aspects of LLM usage. Apply by 1 May 2025.

Postdoctoral Researcher in Natural Language Processing

Postdoc in Natural Language Processing, Department of Computer Science, Faculty of Science, University of Copenhagen The Natural Language Process

candidate.hr-manager.net

Very excited to release Kaleidoscope—a multilingual, multimodal evaluation set for VLMs, built as part of our open-science initiative! 🌍 18 languages (high-, mid-, low-) 📚 21k questions (55% require image understanding) 🧪 STEM, social science, reasoning, and practical skills

Bild