Conor Durkan

@conormdurkan.bsky.social

Generative modeling person https://conordurkan.com

More than anything, the R1 model and paper make me wonder how far along we'd be if everyone was still clamoring to shout their best ideas from the rooftops. I know it's naïve to think we could sustain open research of the kind we saw up to 2020 indefinitely, but still...

I like the Bayesian framing of reward-based post-training (i.e. reward-maximization with a KL penalty). (Figure from 'RL with KL penalties is better viewed as Bayesian inference', link below along with other useful references)

Bild