In my experimentation this week, I found that I was right about one thing and wrong about another. I was right that there is an ideal average token size for our workload. I was wrong that it was a single byte. elijahpotter.dev/articles/byt...
Bytes Are Not Big Enough
I am trying to create a strong and small base model. A model that can semantically understand small passages of text, which we can then retrain to do a variety of different things.
elijahpotter.dev