Announcing 🔭Hubble, a suite of open-source LLMs to advance the study of memorization! Pretrained 1B/8B param models, with controlled insertion of texts designed to emulate key memorization risks: copyright (e.g., book passages), privacy (e.g., synthetic biographies), and test set contamination
Ameya Godbole
@ameyagodbole.bsky.social
PhD student USC NLP working on generalization and reasoning, prev UMassAmherst, IITG (he/him)
🤔 We know what people are using LLMs for, but do we know how they collaborate with an LLM? 🔍 In a recent paper we answered this by analyzing multi-turn sessions in 21 million Microsoft Copilot for consumers and WildChat interaction logs: arxiv.org/abs/2505.16023