Leshem (Legend) Choshen @EMNLP

@lchoshen.bsky.social

🥇 LLMs together (co-created model merging, BabyLM, textArena.ai) 🥈 Spreading science over hype in #ML & #NLP Proud shareLM💬 Donor @IBMResearch & @MIT_CSAIL

Do you think Anthropic's PR were worried they miss out the positive hype of openAI breaking laws and attacking harmfully two companies. So they ran to look for places to say they also do it? (and pretend remorseful) "Everything you can do I can do" better? apnews.com/article/anth...

Anthropic says its AI models hacked 3 organizations during testing

Anthropic says its AI models hacked into three organizations during testing. This comes just days after OpenAI said its AI models went rogue and hacked into another company.

apnews.com

OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts) using repeated budget-matched runs, strong baselines, and statistical hypothesis testing.

Bild

Is Gemma 4 multilingual? Not really🤖 A true multilingual LLM should share language-invariant skills like spatial understanding across languages. It doesnt😱 We had LLMs play 2D board games against themselves in different languages🕹️ and found 2D performance is language dependent

Bild

We pitted LLMs against themselves in different languages across games🎮 Usually when a LLM is worse at a task in some language, we go "it's a language issue". But isn't that circular? If the task is made of language, how do you tell if the model lacks skills, or just struggles to read the question? 🧵

Bild

While language is obviously solved, Google Home can't understand a word I speak to it. Text or speech to text language and cultural issues are countless. This calls for you to join, to co author and to show those model providers they should care about your language and culture as well.

Multilingual Representation Workshop @ EMNLP 2026@mrl-workshop.bsky.social · 2mo ago

After the enthusiasm of the shared task last year, we are running another shared task to create a community-made, culturally relevant multilingual benchmark! The deadline to contribute is August 1 AoE. See more details below.

I've just got hundreds of new citations in Google Scholar Apparently, you can search your name and first letter + last name and find tons of badly parsed versions of your papers. (Yes, citations don't matter, but making camera ready is boring, and I needed an escape,...)

Bild

At last, a way to unlock mulitlingual knowledge sharing in LLMs! 🌍 ​By pretraining an English/Arabic model and swapping Arabic for a word-wise, 1-to-1 translation mapping to English, we saw a massive boost in cross-lingual knowledge transfer 🚀

Bild