We're excited to be sponsoring and co-organizing IOL-AI 2026, a new open challenge with the International Linguistics Olympiad. 🌎💬 🌐 Open-science 🗓️ One-month competition 🎯 Targeting AI reasoning with challenging, unseen test problems from the 2026 olympiad.
Cohere Labs
@cohereforai.bsky.social
@Cohere.com's non-profit research lab and open science initiative that seeks to solve complex machine learning problems. Join us in exploring the unknown, together. https://cohere.com/research
I’m at Papers In The Park! All the good people are here. We’re covering CALIBER: Calibrating Confidence Before and After Reasoning in Language Models. Thanks to Alvin and @cohereforai.bsky.social for sponsoring… and the paper! Shoutout to Finlay, Kurien, Dash, Fadaee, and Ermis!
It’s S4E4 of Papers In The Park. Javad Rajabi is here to walk us through his paper on SEGA. arxiv.org/abs/2605.22668 Thanks to @cohereforai.bsky.social for sponsoring, to Anthony and Alvin for organizing, and Javad for walking us.
S2E2 of Papers In The Park! This week: DFlash: Block Diffusion for Flash Speculative Decoding arxiv.org/abs/2602.06036
AI is getting better at math. Better at code. But is it getting better at understanding cultural nuances? 🤔 Join us for “Cultural Awareness in AI — From Knowledge Tests to Social Norms and Beyond”, a conversation on what it means to build AI systems that work at global scale.
Does AI truly understand different cultures and languages? We’re surveying cultural awareness in real-world AI use. ✨ When cultural awareness matters in real-world AI use 💡 Whether AI reflects diverse norms, communication styles & knowledge 🫥Where AI falls short in cultural understanding
1) what? Cohere is here?!!!! 2) this is crazy
Introducing ✨Tiny Aya✨, a family of massively multilingual small language models built to run where people actually are. Tiny Aya delivers strong multilingual performance in 70+ global languages in a 3.35B parameter model, efficient enough to run locally, even on a phone.
Woo hoo, who would have thought Canada would produce efficient massively multicultural models
1) what? Cohere is here?!!!! 2) this is crazy
🌱Very proud of our team's latest release 😊 meet Tiny Aya, a massively multilingual model with 3.35B parameters. Tech report here: github.com/Cohere-Labs/...
github.com
Introducing ✨Tiny Aya✨, a family of massively multilingual small language models built to run where people actually are. Tiny Aya delivers strong multilingual performance in 70+ global languages in a 3.35B parameter model, efficient enough to run locally, even on a phone.
Introducing ✨Tiny Aya✨, a family of massively multilingual small language models built to run where people actually are. Tiny Aya delivers strong multilingual performance in 70+ global languages in a 3.35B parameter model, efficient enough to run locally, even on a phone.
Excited to have two of our papers featured in @j-novikova-nlp.bsky.social's @wiair.bsky.social podcast, as part of the NeurIPS reflection. ✨ Learn more / subscribe here women-in-ai-research.github.io and check out this thread 🧵 for our features...
Women in AI Research Podcast
Celebrating the remarkable contributions of female AI researchers from around the globe
women-in-ai-research.github.io
What an incredible week it’s been at #NeurIPS2025! 🎉 Today is our last one at the booth. We've had a great week connecting with our community in San Diego. Join our community to continue to connect with our research team: https://cohere.com/research/open-science/application
What's the story of your legend? Join ML researchers building their legends with 40 cards that capture our shared journey—explore and build yours: https://lab-legends.vercel.app/ 🎯
Just 1 day left until #NeurIPS2025 kicks off! The Cohere and Cohere Labs teams are ready to dive into a packed week of research, conversations, and community at the San Diego Convention Center✨ Come visit our booth — we’d love to chat and send you home with some swag!
How well do LLMs handle multilinguality? 🌍🤖 🔬We brought the rigor from Machine Translation evaluation to multilingual LLM benchmarking and organized the WMT25 Multilingual Instruction Shared Task spanning 30 languages and 5 subtasks.
River, Yinhong and I will all be in person and we look forward to the discussions!
Cohere Labs x EMNLP 2025 "When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning" Congrats to authors Yijiang River Dong, @tiancheng.bsky.social, Yinhong Liu, Ahmet Üstün, Nigel Collier. 📜 arxiv.org/abs/2502.19158
We’re thrilled to announce that some of our research will be presented at @emnlpmeeting.bsky.social next week! 🥳 If you’re attending the conference, don’t miss the chance to explore our work and connect with our team.
“Individually, we are one drop. Together, we are an ocean.” - Ryunosuke Satoro ✨ Cohere Labs is excited to announce Connect - a 3-day virtual conference celebrating the power of collaboration in open science!
🌍Most multilingual instruction data starts as English and translation can’t capture cultural nuance or linguistic richness What if we optimized prompts instead of completions? That’s the focus of our most recent work on prompt space optimization for multilingual synthetic data🗣️
🚀 Global MMLU Lite is now live on Kaggle Benchmarks! Developed by @cohereforai.bsky.social, it spans 16 languages with both Culturally Sensitive & Agnostic samples - helping researchers uncover cultural & linguistic biases in multilingual evaluation.
Global MMLU Lite Leaderboard | Kaggle
Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.
kaggle.com
Global AI deserves reproducible and transparent evaluation. 🌎 With Global MMLU Lite now part of @kaggle.com Benchmarks, you can track the multilingual performance of top models as well as test your own! Check out the leaderboard and notebook linked below.
This month, we've been very excited to welcome Joelle Pineau, @cohere.com's new Chief AI Officer. We look forward to working together on frontier research - advancing the science of building models that are robust, capable, and impactful in the real world.
Keynote talk: Optimizing Multilinguality Post Training. Can multilingual ability be boosted at post training? Julia Kreutzer from @cohereforai.bsky.social explores RL, test-time scaling & data distillation to improve open-ended tasks across languages. 🌍✨ #MELTWorkshop2025
Let's do the venue justice. Very excited for today's multilingual workshops at #COLM2025 💙
Today at COLM, Cohere Labs Sr Research Scientist, @juliakreutzer.bsky.social will be presenting at 2 workshops. First, the Multilingual Data Quality Signals workshop, bringing together researchers across disciplines to discuss & present research on data quality signals in multilingual data.
Today at COLM, Cohere Labs Sr Research Scientist, @juliakreutzer.bsky.social will be presenting at 2 workshops. First, the Multilingual Data Quality Signals workshop, bringing together researchers across disciplines to discuss & present research on data quality signals in multilingual data.
Today at COLM, we are excited to share our work Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation, during Poster Session 4, 4:30 - 6:30pm. Come connect with paper authors @juliakreutzer.bsky.social and @kocmitom.bsky.social.
💡A collaborative➕diverse team is key. In real life as in the LLM world 💪🦾 Check out our latest work that builds on this insight. 👇
Is Best-of-N really the best use of your inference compute? Introducing Fusion-of-N: a simple and powerful way to advance inference and distillation beyond Best-of-N.
Is Best-of-N really the best use of your inference compute? Introducing Fusion-of-N: a simple and powerful way to advance inference and distillation beyond Best-of-N.
We’re not your average lab. We’re a hybrid research environment dedicated to revolutionizing the ML space. And we’re hiring a Senior Research Scientist to co-create with us. If you believe in research as a shared, global effort — this is your chance.
What if the way we verify synthetic code is limiting model performance? In our latest work we uncover the Verification Ceiling Problem: strict “all tests must pass” rules throw away useful data, while weak tests let errors through.