My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.
Amir-massoud Farahmand
@sologen.bsky.social
Research Goal: Understanding the computational and statistical principles required to design AI/RL agents. Associate Professor at Polytechnique Montréal and Mila. 🇨🇦 academic.sologen.net
📚 Interested in Continual RL but not sure where to start? 👇 Dive in: sites.google.com/view/continu... ♾️ We've curated a resource hub with key papers, benchmarks, codebases, tutorials, and more to help you get up to speed quickly. #ContinualRL #ReinforcementLearning #MachineLearning #AI
This platform will not replace Twitter/X for us (scientists, researchers, profs, etc.). It might be (noticeably) better in several aspects, but (1) it is not a disruptive innovation, and (2) has the third mover disadvantage. P.S: I'll stay active here for now.
It is interesting that we still don't have a completely clear picture of why/when (Optimisitic) Policy Iteration + Monte Carlo estimate works, especially with every-visit update model (which can be biased though consistent, BTW).
Temporal Difference Learning for Diffusion Models (ICML 2026) arxiv.org/abs/2606.15048 By Yangchen Pan (my former PhD student) and co-authors. It reformulates diffusion training as a Markov reward process and introduces a TD objective to encourage temporal consistency across denoising steps.
Temporal Difference Learning for Diffusion Models
Diffusion models are typically trained with objectives that focus on local denoising targets at individual time steps (or adjacent pairs), which do not enforce consistency between predictions along th...
arxiv.org
Sad to hear the passing of Dimitri Bertsekas (1942- 2026). His work has been very influential to me and shaped the way I think about RL. I am sure this is the case for many others in the RL, Control, and Optimization communities.
Good news, RL Community! The early registration deadline for RLC'26 has been extended to June 17th — don't miss the early rates! Register today! Full refunds for cancellations before July 14, 2026.
🚀 PhD position in #NeuroAI & neurodevelopment 🚀 Co-supervised by Sarah Lippé and myself, to investigate visual processing & cognition abnormalities in children with neurodevelopmental disorders in a neuroAI framework. Full project details and how to apply here: tinyurl.com/kbuyntpn 🧠🤖 📈
If you have a Reinforcement Learning paper accepted at TMLR, JMLR, JAIR, AIJ, or MLJ, you can use this wonderful opportunity to present your work and meet your peers at RLC in Montreal, Canada. 🇨🇦
Got a great TMLR paper but missed the RLC deadline? Following last year’s success, @RL_Conference is back with a Journal-to-Conference track! Accepted TMLR papers within scope are invited to submit for consideration. Please submit here: docs.google.com/forms/d/e/1F...
📣 There's never a "best" time to share important updates, especially after sitting on this for so long. I'm joining the faculty @brighamyoungu.bsky.social this Summer as an Assistant Professor in the CS Dept, in preparation for the coming school year. Lots of excitement and a fair bit of nerves. 🧵
Do you use often use PPO, but wish you could use something just better? Try REPPO: Relative Entropy Pathwise Policy Optimization! Project Page: cvoelcker.de/projects/rep...
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
Hypothesis: People have been gradually shifting to write more like ChatGPT and alike. They use structures such as "This is not only X; but it is also Y". These struct. are natural part of lang, but either 1. they're becoming more prevalent, 2. I've become more sensitive to them.
We have a new PhD Candidate in town: @tylerkastner.bsky.social Looking forward to all the new work you will be doing on Distributional Reinforcement Learning.
We have the keynote speakers for RLC2026 now! Thrilled to welcome Rika Antonova, Sheila McIlraith, Marc G. Bellemare, Danijar Hafner, Balaraman Ravindran! Details: rl-conference.cc/index.html The RL community is coming together this August in Montréal, Québec, Canada. Hope you make it!
RLC 2026
rl-conference.cc
Palantir has student data, including immigration status, from the ed tech discussion platform Piazza. Palantir paid Piazza $916,000 for access to this data. www.sec.gov/Archives/edg... I blew the whistle on this in 2016 and the CEO contacted my employer.
Let’s also kick edtech out of schools because I promise you Palantir has that data too
Happy Norooz, the Persian new year 1405/2585, the equinox, and the beginning of spring!
Following advice by the always-wise @eugenevinitsky.bsky.social , I am trying to get back into the habit of blogging (again) ✏️! Featuring today's post: How to pick an RL algorithm for your problem cvoelcker.de/blog/2026/ch... Please share and give feedback! #reinforcementlearning
cookie monster is sitting at a table with a tray of food and the words choices written on it
Alt: cookie monster is sitting at a table with a tray of food and the words choices written on it
media.tenor.com
In light of the ongoing conflict in the Middle East, RLC decided to remove the abstract deadline: rl-conference.cc/callforpaper... The only deadline is for the full paper: Mar 5(AOE) openreview.net/group?id=rl-... Affected folks may also contact the PCs to discuss deadline extensions before Mar 5.
RLC 2026 Conference
Welcome to the OpenReview homepage for RLC 2026 Conference
openreview.net
Ali Khamenei is in hell. The world is a better place now!
RLC 2026 Call for Workshop is live on OpenReview! Submission deadline: Mar 12 (AoE). Full details here: rl-conference.cc/call_for_wor... @glenberseth.bsky.social @eugenevinitsky.bsky.social @twkillian.bsky.social @schaul.bsky.social @sologen.bsky.social @audurand.bsky.social @bradknox.bsky.social
RLJ | RLC Call for Workshops
rl-conference.cc
RLC2026 Call for Workshops! We’re already live: openreview.net/group?id=rl-... Here's the opportunity to help shape the conference & spotlight your own RL focus areas. Call: rl-conference.cc/call_for_wor... Deadline: Mar 12 (AoE) And don't forget the awesome banquet :) www.cirquedusoleil.com/ludo
Submit your RL papers to RLC! This is now perhaps the best venue for RL researchers.
Only ~7 days to go until the full paper deadline (Mar 5, AOE)! Wishing everyone a strong close to the submission cycle.
I am rerunning my class on robot learning this year, and I plan to push many code examples to help others get to the ugly details fast. One of these details is how BC gets off track as network sizes change. Blog and notebook below.
🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇
The compliment of the day: "What’s unusual is your willingness to follow the logic all the way through instead of stopping where it becomes socially awkward".
Has taken a long time to polish, but slowly becoming very proud of rlhfbook.com and do think it's a great resource for many people. A lot of hours (and tokens and reader feedback) going into making it right.
A significant hurdle of the empirical RL and the broader AI research is caused by the limitations of the environments in which our agents learn and build their "artificial minds". This should be compared with the richness of the real-world in which a human child flourishes.