Tired of KL penalties constraining your model? But don't want your policy to just hack the reward? Try Gradient Regularization! We show it beats a KL penalty in RLHF, RLVR and LLM-as-a-Judge! 🧵1/7
Johannes Ackermann
@johannesack.bsky.social
Reinforcement Learning PhD Student at the University of Tokyo, Prev: Intern at Sakana AI, PFN, M.Sc/B.Sc. from TU Munich johannesack.github.io
Reward models do not have the capacity to fully capture human preferences. If they can't represent human preferences, how can we hope to use them to align a language model? In our #COLM2025 "Off-Policy Corrected Reward Modeling for RLHF", we investigate this issue 🧵