We built code datasets, English datasets, and now it’s time for math! 🚀 Check out Anton’s thread to learn how we curated the best public math pre-training dataset.
Introducing 📐FineMath: the best open math pre-training dataset with 50B+ tokens! Math remains challenging for LLMs and by training on FineMath we see considerable gains over other math datasets, especially on GSM8K and MATH. 🤗 huggingface.co/datasets/Hug... Here’s a breakdown 🧵