Tomasz Limisiewicz

@tomlim.bsky.social

Postdoc at Meta and university of Washington in NLP. Before: PhD from Charles University (Prague 🏰). Interested in going into the inner workings of neural networks 🔍, multilinguality 🌍, tokenization 🔡 and fairer NLP ⚖️ (he/him)

We present Compute Optimal Tokenization! 🔠 Common in LLM scaling works stick to one tokenizer, sweeping data/model size. But what happens when we control the tokenizer’s compression rate (bytes/token)? Here we sweep tokenizers, params, and data across compute budgets: [1/N]