🚨 New @iclr_conf paper! 🚨 Learning is Forgetting: LLM Training As Lossy Compression by @henryconklin.bsky.social, @tomhosking.bsky.social, Yi-Chern Tan, Julian Gold, Jonathan Cohen, @cocoscilab.bsky.social, @maxbartolo.bsky.social and @seraphinagt.bsky.social Arxiv: arxiv.org/abs/2604.075... 🧵
@seraphinagt.bsky.social
Today (two weeks after model launch 🔥) we're releasing a technical report of how we made Command A and R7B 🚀! It has detailed breakdowns of our training process, and evaluations per capability (tools, multilingual, code, reasoning, safety, enterprise, long context)🧵 1/3.