I want to reshare @brandfonbrener.bsky.social's @NeurIPSConf 2024 paper on CoLoR-Filter: A simple yet powerful method for selecting high-quality data for language model pre-training! With @hlzhang109.bsky.social @schwarzjn.bsky.social @shamkakade.bsky.social
David Brandfonbrener
@brandfonbrener.bsky.social
Research scientist at Meta on the llama team Thinking about language models Past: PhD at NYU, fellow at Harvard’s Kempner Institute
I’m heading to NeurIPS Wednesday through Sunday. DM me if you want to meet up!
NEW: we have an exciting opportunity for a tenure-track professor at the #KempnerInstitute and the John A. Paulson School of Engineering and Applied Sciences (SEAS). Read the full description & apply today: academicpositions.harvard.edu/postings/14362 #ML #AI
How does test loss change as we change the training data? And how does this interact with scaling laws? We propose a methodology to approach these questions by showing that we can predict the performance across datasets and losses with simple shifted power law fits.