@preslavnakov.bsky.social

Building on FineWeb’s global deduplication findings, we introduce a strategic upsampling recipe which outperforms FineWeb using TxT360. Full details are in the Upsampling Experiment section of the release blog.

Bild

🪟🛠️LLM360 is committed to making open source AI accessible, transparent, and reproducible. High-quality data is the first step toward better open source models...and we are excited to join the party contributing the first globally deduplicated dataset containing 5.7T tokens!

Bild