People should probably think about more ways to modify BPE tokenisers. I made this graphical representation of all the algorithms we have for this in the current literature (included in the ReBPE paper, aclanthology.org/2026.finding...)
ir. Thomas Bauwens
@bauwenst.bsky.social
PhD researcher in NLP at KU Leuven. Creator and maintainer of TkTkT, the largest tokenisation library in the world. Dancer, photographer and videographer in my free time (IG: @thomas__bailes)
BPE-knockout just got outperformed by an algorithm that modifies BPE tokenisers in a feedback loop to make them absorb more and more constraints. It doesn't even need more data to do that. It uses the tokeniser itself as a dataset. 🧵
The gadfly of {my immediate surroundings} has landed on Bluesky.