Willem Röpke

@willemropke.bsky.social

PhD student | Interested in all things decision-making and learning

How can I stop ChatGPT from talking to me with emojis, this is just the worst update I've ever experienced. I've put it in its memory, in my details, and I even repeat it in the chat but it's just replying like 👉🥺👈

The fact that in the year 2025 we are still dealing with the stupid "make the paper fit in an arbitrary format for the camera ready submission" minigame is killing me. Either let me group authors or let me put acknowledgements after the main text. This isn't hard.

I think this is the best paper I’ve ever read: arxiv.org/abs/2404.03715 A strong emphasis on theoretically principled algorithms for RLHF followed by motivated practical implementations. Well-written and a clear overview of the relevant background and related work. 10/10 no comments

Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

This paper studies post-training large language models (LLMs) using preference feedback from a powerful oracle to help a model iteratively improve over itself. The typical approach for post-training L...

arxiv.org

I realise I'm woefully unqualified on this topic, but can someone please explain why we still don't have personal carrier drones? This seems like an obvious next step in transportation and given the state of our tech tree shouldn't be that hard?

I'm having a weird problem with training DQN on minatar (specifically the gymnax version). In space invaders and breakout, my eval metrics are extremely unstable while my train metric is very smooth. See an example of space invaders below (eval left, train right). Any ideas of what went wrong?

BildBild