Jonas Hübotter

@jonhue.bsky.social

PhD student at ETH Zurich jonhue.github.io

Training LLMs with verifiable rewards uses 1bit signal per generated response. This hides why the model failed. Today, we introduce a simple algorithm that enables the model to learn from any rich feedback! And then turns it into dense supervision. (1/n)

Bild

In our ICML paper, we study fine-tuning a generalist policy for multiple tasks. We ask, provided a pre-trained policy, how can we maximize multi-task performance with a minimal number of additional demonstrations? 📌 We are presenting a possible solution on Wed, 11am to 1.30pm at B2-B3 W-609!

Bild

We've released our lecture notes for the course Probabilistic AI at ETH Zurich, covering uncertainty in ML and its importance for sequential decision making. Thanks a lot to @jonhue.bsky.social for his amazing effort and to everyone who contributed! We hope this resource is useful to you!

Jonas Hübotter@jonhue.bsky.social · last yr.

I'm very excited to share notes on Probabilistic AI that I have been writing with @arkrause.bsky.social 🥳 arxiv.org/pdf/2502.05244 These notes aim to give a graduate-level introduction to probabilistic ML + sequential decision-making. I'm super glad to be able to share them with all of you now!

Assume that the nodes of a social network can choose between two alternative technologies: B and X. A node using B receives a benefit with respect to X, but there is a benefit to using the same tech as the majority of your neighbors. Assume everyone uses X at time t=0. Will they switch to B?

Spread of innovation in a small world network.