Andreas Krause

@arkrause.bsky.social

Professor at ETH Zurich

SDPO enables RL agents to learn from rich feedback (i.e., not only whether an attempt failed, but why it failed, such as error messages). Even without such rich feedback, SDPO can reflect on past attempts and outperform GRPO. SDPO also accelerates solution discovery at test time!

Jonas Hübotter@jonhue.bsky.social · 6mo ago

Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed. Today, we introduce a simple algorithm that enables the model to learn from any rich feedback! And then turns it into dense supervision. (1/n)

ICML reaffirms its support to the community and standards of conduct: - We do not tolerate harassment or other improper conduct; - Academic integrity is paramount; - We redouble our support to peer review, with more incentives for reviewers & financial support for OpenReview icml.cc/public/blog#...

Bild

EPFL, ETH Zurich, and CSCS today released Apertus, Switzerland's first large-scale, multilingual language model (LLM). As a fully open LLM, it serves as a building block for developers and organizations to create their own applications. ethz.ch/en/news-and-...

Apertus: a fully open, transparent, multilingual language model

EPFL, ETH Zurich and the Swiss National Supercomputing Centre (CSCS) released Apertus today, Switzerland’s first large-scale, open, multilingual language model — a milestone in generative AI for trans...

ethz.ch

Clinical notes are messy, inconsistent, and unstructured—yet they hold some of the most valuable signals in real-world clinical practice. Join us today at ICML at the Foundation Models for Structured Data workshop to see how we can make sense of these notes! 📍 West Ballroom D

Bild

In our ICML paper, we study fine-tuning a generalist policy for multiple tasks. We ask, provided a pre-trained policy, how can we maximize multi-task performance with a minimal number of additional demonstrations? 📌 We are presenting a possible solution on Wed, 11am to 1.30pm at B2-B3 W-609!

Bild

We've released our lecture notes for the course Probabilistic AI at ETH Zurich, covering uncertainty in ML and its importance for sequential decision making. Thanks a lot to @jonhue.bsky.social for his amazing effort and to everyone who contributed! We hope this resource is useful to you!

Jonas Hübotter@jonhue.bsky.social · last yr.

I'm very excited to share notes on Probabilistic AI that I have been writing with @arkrause.bsky.social 🥳 arxiv.org/pdf/2502.05244 These notes aim to give a graduate-level introduction to probabilistic ML + sequential decision-making. I'm super glad to be able to share them with all of you now!

🚨 New reinforcement learning algorithms 🚨 Excited to announce MaxInfoRL, a class of model-free RL algorithms that solves complex continuous control tasks (including vision-based!) by steering exploration towards informative transitions. Details in the thread 👇

We’re presenting our work “When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL” today at NeurIPS. Come join us in West at poster #6604 from 16:30-19:30! Joint work with my fantastic collaborators Bhavya Sukhija, Yarden As, Florian Dörfler, @arkrause.bsky.social