Marcel Hussing

@marcelhussing.bsky.social

PhD student at the University of Pennsylvania. Prev, intern at MSR, and Meta FAIR. Interested in reliable and replicable reinforcement learning, robotics and knowledge discovery: https://marcelhussing.github.io/ All posts are my own.

I've tried to be a good reviewer, read all the papers carefully, and formed my own view. I'm getting page long, clearly LLM generated rebuttals. It is not my job to prompt your LLM into convincing me. What are people's solutions to this? Should I message the AC? #NeurIPS2026

This is a robot failing to grasp a ball. Almost every robot lab produces clips like this daily… and almost all of them get thrown away. This is the most abundant but underused resource in robot learning. We’re collecting all of it now as “OopsieData”, please join us at oopsie-data.com! (1/13)

This was a very fun project. In Behavior-Consistent Deep RL, we provide a method that aligns the behavior of independently trained policies. It turns out, this works even in high dimensional spaces. Here are 6 seeds of Humanoids (all ca same return). (left) Baseline (right) Ours.

BildBild
Marcel Hussing@marcelhussing.bsky.social · 3mo ago

🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.

🚨 New Preprint Alert: Behavior-Consistent Deep Reinforcement Learning 🚨 TLDR: We introduce an approach that achieves behavioral similarity across independent algorithm executions in continuous state-action space deep RL.

Bild

Neural networks are highly non-convex, so approximate error minimizers need not look anything like each other in parameter space. But we show that nevertheless (for many model sizes) approximate error minimizers must closely agree in function/prediction space despite this!

Bild

Too many papers sound like this Hierarchical Context-Aware Diffusion-Transformer Meta-World-Model Reinforcement Learning with Causally Disentangled Preference-Aligned Self-Supervised Compositional Multi-Scale Latent Skill Priors for Long-Horizon Generalist Decision Making

One reason I work on replicable and consistent RL is because it is has always been at the top of the list of criteria for reliability.

Bild
Human Rights Data Analysis Group@hrdag.org · 5mo ago

A new paper by @sayash.bsky.social and @randomwalker.bsky.social examines what “reliability” means in an AI context. They propose consistency, robustness, calibration, and safety, and they define these in operationally useful ways. A worthy read! www.normaltech.ai/p/new-paper-...

I have seen multiple times now that a reviewer said sth like: the proofs are simple -> reject the paper. That is completely counter-productive. A theorem needs to generate new insights. If we learn something new from something simple that should be preferred. Don't believe me? Ask someone famous:

Bild

I’ve been thinking about a practical question and would love some opinions: How do your papers actually get discovered/cited? I was searching for recent work on high update ratio RL and found several very closely related papers tackling the same failure modes we study. None cited our earlier work.

Scaling Laws in Particle Physics Data! This is a result I've been itching to share and it's finally out. One of the big open questions is how much better AI-based methods at particle colliders can still become. 1/4

BildBildBild

Two papers accepted to @iclr-conf.bsky.social 2026! One of the is REPPO, see below! I think it deserves a lot more recognition. Let's chat about it in Rio! 🇧🇷

Claas Voelcker@cvoelcker.bsky.social · 7mo ago

🤔 Want to use REPPO (cvoelcker.de/projects/rep...) but hate jax? 🤔 😮 Want to have stable on-policy RL without filling your GPU with an enormous replay buffer? 😮 🤖 Are you a roboticist and just want your RL code to run? 🤖 🎉 Fear not, we started adding new REPPO versions! 🎉 github.com/cvoelcker/rs...