Nowadays ML projects feel like they need to be compressed into a few months. Its refreshing to be able to work on something for a few years! But also a slog.
Ok, so I can finally talk about this! We spent the last year (actually a bit longer) training an LLM with recurrent depth at scale. The model has an internal latent space in which it can adaptively spend more compute to think longer. I think the tech report ...🐦⬛