Sasikanth Kotti

@ksasi.bsky.social

🚨New blog: The AI Evaluation Chart Crisis 📝 From misleading bar heights to missing error bars, recent model launches have sparked debate on AI evals. In our new blogpost, we dig into what’s broken, why it matters and how they should be presented 👇 evalevalai.com/documentatio...

The AI Evaluation Chart Crisis

Charts used to showcase performance demonstrate broader issues in the AI evaluation ecosystem: a lack of balance between competitive benchmarking and statistical rigor.

evalevalai.com

The baseline is around 65% accuracy with a simple fine-tuned BERT and ~95% with a LLM like GPT4o-mini.... your mission is to beat it while using the least energy possible! Check out the dataset here: huggingface.co/collections/... And stay tuned for the last task later this week!🔥

Frugal AI Challenge Tasks - a frugal-ai-challenge Collection

Find the 3 datasets for the Frugal AI Challenge in this Collection! 🌎 Find all the details of the challenge at https://frugalaichallenge.org/

huggingface.co

Today we’re launching the text challenge of the Frugal AI Challenge for the AI Action Summit: detecting climate misinformation in the media (Press, TV, Radio), sponsored by the French non-profit QuotaClimat. The goal of the task is to detect climate-based misinformation and to categorize its type 📃

Bild

Announcing Global-MMLU - an improved MMLU Open dataset with evaluation coverage across 42 languages. The result of months of work with the goal of advancing Multilingual LLM evaluation. Built together with the community and amazing collaborators at Cohere4AI, MILA, MIT, and many more.

Bild

Day 2 of smol course and the community is building something here. 👷 If you want to get involved, you can do this: - read (and star) the repo - check out our new discord channel - open a PR to submit an exercise on module 1 - open an issue to improve the course - review another submission 🧵

Just FYI because it seems relevant and I've seen it misstated a few times, LAION retrained their dataset and provided a diff to migrate over to the filtered Re-LAION-5B dataset HuggingFace, like most platforms that handle user content, do checks for CSAM too. laion.ai/blog/relaion...

Releasing Re-LAION 5B: transparent iteration on LAION-5B with additional safety fixes | LAION

<p>Today, following <a href="https://laion.ai/notes/laion-maintenance/">a safety revision procedure</a>, we announce Re-LAION-5B, an updated version of LAION...

laion.ai