Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed. Today, we introduce a simple algorithm that enables the model to learn from any rich feedback! And then turns it into dense supervision. (1/n)
Frederike Lübeck
@rikelue.bsky.social
PhD student in Machine Learning at ETH Zurich & Max Planck Institute
Clinical notes are messy, inconsistent, and unstructured—yet they hold some of the most valuable signals in real-world clinical practice. Join us today at ICML at the Foundation Models for Structured Data workshop to see how we can make sense of these notes! 📍 West Ballroom D