Julia Kreutzer

@juliakreutzer.bsky.social

NLP & ML research @cohereforai.bsky.social 🇨🇦

🤨"But why linguistics" is the most common question when talking about linguistic reasoning benchmarks. Last year we organized a shared task at WMT...and no one participated 🤣 🤯Let me change your mind why this is one of the most challenging, focused and best reasoning benchmarks right now.

Cohere Labs@cohereforai.bsky.social · last mo.

For the launch of the 1-month open challenge that precedes the in-person Olympiad, @juliakreutzer.bsky.social is hosting @danmirea.bsky.social & Eduardo Sanchez to discuss the appeal of these problems, solution challenges, and potential insights from AI-human evaluation. Join us: luma.com/jk8uv7zs

💭We need more research that focuses on aspects beyond accuracy, especially in multilingual AI. 👉Help us explore the importance of culture in building and testing AI, with a few minutes of your time. Happy to have a chat as well with anyone who's interested in that space!

Cohere Labs@cohereforai.bsky.social · 5mo ago

Does AI truly understand different cultures and languages? We’re surveying cultural awareness in real-world AI use. ✨ When cultural awareness matters in real-world AI use 💡 Whether AI reflects diverse norms, communication styles & knowledge 🫥Where AI falls short in cultural understanding

How well do LLMs handle multilinguality? 🌍🤖 🔬We brought the rigor from Machine Translation evaluation to multilingual LLM benchmarking and organized the WMT25 Multilingual Instruction Shared Task spanning 30 languages and 5 subtasks.

Bild

🌍Most multilingual instruction data starts as English and translation can’t capture cultural nuance or linguistic richness What if we optimized prompts instead of completions? That’s the focus of our most recent work on prompt space optimization for multilingual synthetic data🗣️

Bild

We’re not your average lab. We’re a hybrid research environment dedicated to revolutionizing the ML space. And we’re hiring a Senior Research Scientist to co-create with us. If you believe in research as a shared, global effort — this is your chance.

Bild

Breaking into AI research is harder than ever, and early-career researchers face fewer chances to get started. Entry points matter. We started the Scholars Program 3 years ago to give new researchers a real shot — excited to open applications for year 4✨

Cohere Labs@cohereforai.bsky.social · 12mo ago

Applications are now open for the next cohort of the Cohere Labs Scholars Program! 🌟 This is your chance to collaborate with some of the brightest minds in AI & chart new courses in ML research. Let's change the spaces breakthroughs happen. Apply by Aug 29.

While effective for chess♟️, Elo ratings struggle with LLM evaluation due to volatility and transitivity issues. New post in collaboration with AI Singapore explores why Elo falls short for AI leaderboards and how we can do better.

Bild

🍋 Squeezing the most of few samples - check out our LLMonade recipe for few-sample test-time scaling in multitask environments. Turns out that standard methods miss out on gains on non-English languages. We propose more robust alternatives. Very proud of this work that our scholar Ammar led! 🚀

Cohere Labs@cohereforai.bsky.social · last yr.

Can we improve the performance of LLMs during inference without the need for extensive sampling OR special reward models? 🤔 Our latest work introduces a new inference time scaling recipe that is sample-efficient, multilingual, and suitable for multi-task requirements. 🍋

1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵

Bild

🤓MT eyes on multilingual LLM benchmarks 👉 Here's a bunch of simple techniques that we could adopt easily, and in total get a much richer understanding of where we are with multilingual LLMs. 🍬Bonus question: how can we spur research on evaluation of evaluations?

Cohere Labs@cohereforai.bsky.social · last yr.

🚀🌍The rapid advancement of multilingual large language models (mLLMs) is exciting, but are we evaluating them effectively? Our new paper explores how we can improve generative evaluations for mLLMs by learning from machine translation (MT) evaluation practices. 🔎

Tired of messy non-replicable multilingual LLM evaluation? So were we. In our new paper, we experimentally illustrate common eval. issues and present how structured evaluation design, transparent reporting, and meta-evaluation can help us to build stronger models.

Julia Kreutzer@juliakreutzer.bsky.social · last yr.

📖New preprint with Eleftheria Briakou @swetaagrawal.bsky.social @mziizm.bsky.social @kocmitom.bsky.social! arxiv.org/abs/2504.11829 🌍It reflects experiences from my personal research journey: coming from MT into multilingual LLM research I missed reliable evaluations and evaluation research…

Screenshot of the paper header with title and author list and affiliations

🚀 We are excited to introduce Kaleidoscope, the largest culturally-authentic exam benchmark. 📌 Most VLM benchmarks are English-centric or rely on translations—missing linguistic & cultural nuance. Kaleidoscope expands in-language multilingual 🌎 & multimodal 👀 VLMs evaluation

Bild

☀️ Summer internship at Cohere! Are you excited about multilingual evaluation, human judgment, or meta-eval? Come help us explore how a rigorous eval really looks like while questioning the status quo in LLM evaluation. I’m looking for an intern (EU timezone preferred), are you interested? Ping me!

Command🅰️ technical report is out. Information-dense. Detailed. Pretty. Simply A+! 💎: cohere.com/research/pap...

Command A: An Enterprise-Ready Family of Large Language Models

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command

cohere.com

Max Bartolo@maxbartolo.bsky.social · last yr.

I'm excited to share the tech report for our @cohere.com @cohereforai.bsky.social Command A and Command R7B models. We highlight our novel approach to model training including self-refinement algorithms and model merging techniques at scale. Read more below! ⬇️

A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏

Bild