I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026
Marcel Hussing
@marcelhussing.bsky.social
PhD student at the University of Pennsylvania. Prev, intern at MSR, and Meta FAIR. Interested in reliable and replicable reinforcement learning, robotics and knowledge discovery: https://marcelhussing.github.io/ All posts are my own.
This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13)
At this point I'm not even sure... maybe I prefer LLM reviews.
Post is a bit delayed but last week I passed my dissertation defense on algorithm output stability in RL. 🎉
This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours.
🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.
🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.
Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social
Time for round number two. Stop by #4406 to learn about stability guarantees in RL. #ICLR2026 @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social
At #ICLR2026 presenting our first poster in the morning on Relative Entropy Pathwise Policy Optimization. Stop by at #4613. 🧑🎓 @cvoelcker.bsky.social, @axelbrunnbauer.bsky.social, Michal Naumann, Pieter Abbeel, @ericeaton.bsky.social, Radu Grosu, @sologen.bsky.social @igilitschenski.bsky.social
In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)
Replicable Reinforcement Learning with Linear Function Approximation
Replication of experimental results has been a challenge faced by many scientific disciplines, including the field of machine learning. Recent work on the theory of machine learning has formalized rep...
arxiv.org
Why is the default option in the rebuttal acknowledgements at ICML to accept the paper? I'm very confused on how to use these buttons. Can someone explain them to me?
Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this!
It's this time of the year again: your baselines cannot be PPO and SAC.
Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...
How AI is changing the nature of mathematical research
What machine learning theorists learned using AI agents to generate proofs — and what comes next.
amazon.science
I'm so glad that so many research problems are finally being treated as first class citizens rather than afterthoughts. 🤔
Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making
One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability.
A new paper by @sayash.bsky.social and @randomwalker.bsky.social examines what “reliability” means in an AI context. They propose consistency, robustness, calibration, and safety, and they define these in operationally useful ways. A worthy read! www.normaltech.ai/p/new-paper-...
Excuse me? Surely telling me that didn't require much thinking.
I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous:
I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work.
🚀 Excited to share REPPO, a new on-policy RL agent! TL;DR: Replace PPO with REPPO for fewer hyperparameter headaches and more robust training. REPPO, led by @cvoelcker.bsky.social, will be presented at ICLR 2026. How does it work? 🧵👇
Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4
"Scientific reviewers should have experience publishing scientific work in related areas" is really not that hot of a take.
Clicking like on any relevant ICLR paper. Encourage people to post their work here more!
The other paper accepted to @iclr-conf.bsky.social 2026 🇧🇷. Our work on replicable RL sheds some light on how to consistently make decisions in RL. @ericeaton.bsky.social @mkearnsphilly.bsky.social @aaroth.bsky.social @sikatasengupta.bsky.social @optimistsinc.bsky.social
I think I posted about it before but never with a thread. We recently put a new preprint on arxiv. 📖 Replicable Reinforcement Learning with Linear Function Approximation 🔗 arxiv.org/abs/2509.08660 In this paper, we study formal replicability in RL with linear function approximation. The... (1/6)
Two papers accepted to @iclr-conf.bsky.social 2026! One of the is REPPO, see below! I think it deserves a lot more recognition. Let's chat about it in Rio! 🇧🇷
🤔 Want to use REPPO (cvoelcker.de/projects/rep...) but hate jax? 🤔 😮 Want to have stable on-policy RL without filling your GPU with an enormous replay buffer? 😮 🤖 Are you a roboticist and just want your RL code to run? 🤖 🎉 Fear not, we started adding new REPPO versions! 🎉 github.com/cvoelcker/rs...
Quite disheartening that there isn't a single workshop at ICLR to present my RL work but there several topics that are listed 5 or 6 times just named differently.