Sara Hooker

@sarahooker.bsky.social

I lead Cohere For AI. Formerly Research Google Brain. ML Efficiency, LLMs, @trustworthy_ml.

1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵

Bild

This has been a topic close to my heart for a long time. We have an awesome lineup of speakers who have made deep contributions to open-source in ML, e.g. @sarahooker.bsky.social , @chrisrackauckas.bsky.social, Matt Johnson, Tri Dao, @stellaathena.bsky.social, Evan Shelhamer.

Frank Schneider@fsschneider.bsky.social · last yr.

Tired of your open-source ML work not getting the academic recognition it deserves? 🤔 Submit to the first-ever CodeML workshop at #ICML2025! It focuses on new libraries, improvements to established ones, best practices, retrospectives, and more. codeml-workshop.github.io/codeml2025/

Today we are releasing Kaleidoscope 🎉 A comprehensive multimodal & multilingual benchmark for VLMs! It contains real questions from exams in different languages. 🌍 20,911 questions and 18 languages 📚 14 subjects (STEM → Humanities) 📸 55% multimodal questions

Bild

We're particularly proud to release Aya Vision 8B - it's compact 🐭 and efficient 🐎, outperforming models up to 11x its size 📈. Releasing open weights helps to make breakthroughs in VLMs accessible to the research community.

Bild

Introducing ✨ Aya Vision ✨ - an open-weights model to connect our world through language and vision Aya Vision adds breakthrough multimodal capabilities to our state-of-the-art multilingual 8B and 32B models. 🌿

Aya Expanse, our open-weight 32B model, outperforms drastically larger models including Claude, Mistral Large 2, & Llama 405B on Scale's Private Multilingual Protocol. We are proud to work on global AI that is efficient and accessible 🔥

Bild

In this cross-institutional work, we introduce technical governance for AI and 100+ 🔢 open technical problems 🔧. We provide a taxonomy of open problem areas in TAIG organized by governance capacities and governance targets. 📜https://arxiv.org/pdf/2407.14981

Bild

The C4AI Research Grant program is proud to have supported a project focused on building LLM tools for teachers 🧑‍🏫 This project focused on adapting educational materials to students’ skill levels, ensuring more effective and responsible AI integration in classrooms.

Bild

Many people have asked me about the France Action Summit. I think a summit is typically most valuable as a catalyst, not as a solution in itself. But, will share some observations.

Bild

On Scale AI's private multilingual protocol, Aya Expanse is indexed as the best open-weights model. Additionally, in some languages we're outperforming: 🔒proprietary models 🐘larger models ⛰️models built by more researchers with more infrastructure Lots to be proud of today.

Bild

As the @cohereforai.bsky.social joins the Bluesky family — we will be sharing paper gems from when we first started as a lab. This paper is part of a larger research agenda where we have focused on how to better represent the long tail = making AI work for almost all real world distributions.

Cohere Labs@cohereforai.bsky.social · 2y ago

How can we mitigate the disparate effect of compression 🗜️on model performance for low-resource languages 💬? Check out our cross-institutional collaboration discusses intriguing & previously unknown generalisation properties of compression. 📜Learn more: arxiv.org/abs/2211.02738