Inference-time scaling is great, but it only works if your model actually explores. RL training — by default — kills that exploration. If you are curious to learn why and what we did about it in about our recent ICLR 2026 paper, check out this 🧵!
Germán Kruszewski
@germank.bsky.social
Senior Scientist @Naver Labs Europe. MSCA Postdoctoral Research @UPF (COLT). Just trying out bsky for now...
🚨 Excited about ML/NLP and looking for a research internship on controlled text generation? Come work with us on advanced constraint processing in large language models at NAVER LABS Europe! ✨ Learn more and apply here: europe.naverlabs.com/job/2-7/
Internship: Advanced Constraint Processing in LLMs
The ability to control the outputs of Large Language Models is a central topic in NLP and Machine Learning, with applications to safety, trustworthiness, reasoning, etc. Following […]
europe.naverlabs.com