Happy to share that the unprocessed results and code for fitting scaling laws and plotting are now available at: github.com/facebookrese...
GitHub - facebookresearch/compute-optimal-tokenization: The repository contains raw data results and code for scaling laws fitting and visualization used in "Compute Optimal Tokenization" paper.
The repository contains raw data results and code for scaling laws fitting and visualization used in "Compute Optimal Tokenization" paper. - facebookresearch/compute-optimal-tokenization
github.com
We present Compute Optimal Tokenization! 🔠 Common in LLM scaling works stick to one tokenizer, sweeping data/model size. But what happens when we control the tokenizer’s compression rate (bytes/token)? Here we sweep tokenizers, params, and data across compute budgets: [1/N]