Daniel Vila

@dvilasuero.hf.co

Everything datasets and human feedback for AI at Hugging Face. Prev: co-founder and CEO of Argilla (acquired by Hugging Face)

💫 Generate RAG data with the Synthetic Data Generator to improve your RAG system! 1️⃣ Generate from your documents, dataset, or dataset description. 2️⃣ Configure it. 3️⃣ Generate the synthetic dataset. 4️⃣ Fine-tune the retrieval and reranking models. 5️⃣ Build a RAG pipeline.

New chapter in the Hugging Face NLP course! 🤗 🚀 We've added a new chapter about the very basics of Argilla to the Hugging Face NLP course. Learn how to set up an Argilla instance, load & annotate datasets, and export them to the Hub.  Any feedback for improvements welcome!

Screenshot of the Introduction to Argilla in Chapter 10 of the Hugging Face NLP course

🎉 50,000+ annotations reached! The FineWeb2-C community is helping build better language models on annotation at a time. 📊 Current stats: - 115 languages represented - 419 amazing contributors - 24 languages with complete datasets But we're not done yet! 🧵

Screenshot of this text:   Total annotations submitted: 50,035  Languages with annotations: 115  Total contributors: 419

Was 2024 the year of datasets? Is 2025 the year for community-built datasets? It's exciting to see the progress of many languages in FineWeb-C: - Total annotations submitted: 41,577 - Languages with annotations: 106 - Total contributors: 363

💥 Ending 2024: A full data annotation journey on the Hugging Face Hub—from raw data to training-ready datasets! With Argilla 2.6.0, push your data to the Hub from the UI Let’s make 2025 the year anyone can build more transparent and accountable AI—no coding or model skills needed.

🔥 We got great feedback on this: "Synthetic Data Generator" A no-code tool to create datasets with LLMs, making it a breeze, allowing ANYONE to create datasets and models in minutes and without any code. Blog: https://buff.ly/4gybyoT GitHub: https://buff.ly/49IDSmd Space: https://buff.ly/3Y1S99z

Introducing the Synthetic Data Generator - Build Datasets with Natural Language

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

buff.ly

Help shape the future of multilingual Open Source AI! Join the FineWeb 2 Community Annotation Sprint to create an open training dataset with full transparency and human validation in many languages. Review datasets in your language and help identify the best sources for training.

Bild

👐 Open Image Preferences is an Apache 2.0 licensed dataset for text-to-image generation by the @hf.co community. This dataset contains 10K text-to-image preference pairs across image generation categories, using different model families and prompt complexities. Blog: huggingface.co/blog/image-p...

Open Preference Dataset for Text-to-Image Generation by the 🤗 Community

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Open Image Preferences released! 🚀 - Open-source dataset for text2image - 10K samples manually evaluated by the HF community. - Binarized format for SFT, DPO, or ORPO. It comes with a nice blog post explaining the steps to pre-process and generate the data, along with the results.

Bild

Announcing Global-MMLU - an improved MMLU Open dataset with evaluation coverage across 42 languages. The result of months of work with the goal of advancing Multilingual LLM evaluation. Built together with the community and amazing collaborators at Cohere4AI, MILA, MIT, and many more.

Bild

We're about to launch the biggest collaboration effort since the Open Assistant. Let's get the highest quality data for open foundation models with all the nuances & diversity of each language, all with data provenance and transparency Join us as language lead: docs.google.com/forms/d/10XI...

Language Lead sign-up

At Hugging Face 🤗, we're launching a big community initiative to improve LLM training for many languages. We're looking for Language Leads to help us cultivate specific languages during this initiativ...

docs.google.com

Next week we're launching a collaborative annotation effort to build a big multilingual dataset, so you can have high-quality data in your language. We are really close to getting leads for 100 languages! Can you help us cover the remaining 200?

Screenshot of a dashboard showing the number of languages with a lead and languages without a lead

[SATURDAY THREAD] ☕️ 🧑‍🎓 In case you spent the week reading GDPR legislation and missed everything. It’s all about vision language models and image preference datasets. >> 🧵 Here are the models and datasets you can use in your projects.

Recently, I added a feature to #Argilla to optimize plugin loading 🎉. It removes unnecessary code, improves readability, and lets future plugins load automatically. 🚀 Check out the PR 👇 and make your first contribution to our repo. github.com/argilla-io/a... #dev_experience #clean_code

🔥 Improve plugins loaders by damianpumar · Pull Request #5697 · argilla-io/argilla

Remove duplicated names Improve the way to load plugins and extensions Auto loading new plugins Convert to typescript Delete unused directive

github.com

🚀 We’re excited to announce Argilla v2.5.0, which includes: * Argilla webhooks, * A new design for the datasets home page. * Python 3.13 and Pydantic v2 support. 📙 Read here 👇 the full release notes github.com/argilla-io/a...

Release v2.5.0 · argilla-io/argilla

🔆 Release highlights Webhooks You can now create and manage webhooks to support your workflows! Webhooks allow you to submit real-time information to other applications whenever a specific event oc...

github.com

A dataset of 1 million or 2 million Bluesky posts is completely irrelevant to training large language models. The primary usecase for the datasets that people are losing their shit over isn't ChatGPT, it's social science research and developing systems that improve Bluesky.

Jeremy Howard @howard.fm · 2y ago

Did you know that 99% of email today is spam? Your inbox isn’t 99% spam because AI is used to filter it. The same 99% will happen here too, but if AI researchers continue to get perma-banned for making available the datasets needed to filter it, it’s going to make this platform unusable.

The best path forward in AI requires technologists to be reflective/self-critical about how their work impacts society. Transparency helps this. Appreciate Bsky for flagging AI ethics &my colleague’s response. Let’s make informed consent a real thing. More later; Recommend: bsky.app/profile/cfie...

Daniel van Strien@danielvanstrien.bsky.social · 2y ago

I've removed the Bluesky data from the repo. While I wanted to support tool development for the platform, I recognize this approach violated principles of transparency and consent in data collection. I apologize for this mistake.

I've removed the Bluesky data from the repo. While I wanted to support tool development for the platform, I recognize this approach violated principles of transparency and consent in data collection. I apologize for this mistake.

Daniel van Strien@danielvanstrien.bsky.social · 2y ago

First dataset for the new @huggingface.bsky.social @bsky.app community organisation: one-million-bluesky-posts 🦋 📊 1M public posts from Bluesky's firehose API 🔍 Includes text, metadata, and language predictions 🔬 Perfect to experiment with using ML for Bluesky 🤗 huggingface.co/datasets/blu...