Gokul Swamy

@gokul.dev

PhD student at @cmurobotics.bsky.social working on efficient algorithms for interactive learning (e.g. imitation / RL / RLHF). no model is an island. prefers email. https://gokul.dev/. on the job market!

Late, but arxiv.org/abs/0804.2996 is *incredible*, so many good lines (e.g., "This comes close to being an accusation of a false claim of priority for a false discovery of an untrue fact, which would be a rare triple-negative in the history of intellectual property disputes.").

The Epic Story of Maximum Likelihood

At a superficial level, the idea of maximum likelihood must be prehistoric: early hunters and gatherers may not have used the words ``method of maximum likelihood'' to describe their choice of where a...

arxiv.org

Recent work has seemed somewhat magical: how can RL with *random* rewards make LLMs reason? We pull back the curtain on these claims and find out this unexpected behavior hinges on the inclusion of certain *heuristics* in the RL algorithm. Our blog post: tinyurl.com/heuristics-c...

Heuristics Considered Harmful: RL With Random Rewards Should Not Make LLMs Reason | Notion

Owen Oertell*, Wenhao Zhao*, Gokul Swamy, Zhiwei Steven Wu, Kiante Brantley, Jason Lee, Wen Sun

tinyurl.com

It was a dream come true to teach the course I wish existed at the start of my PhD. We built up the algorithmic foundations of modern-day RL, imitation learning, and RLHF, going deeper than the usual "grab bag of tricks". All 25 lectures + 150 pages of notes are now public!

Bild

I won't be at #ICLR2025 myself this time around but please go talk to lead authors Nico, Zhaolin, and Runzhe about their bleeding-edge algorithms for imitation learning and RLHF!

Bild

1.5 yrs ago, we set out to answer a seemingly simple question: what are we *actually* getting out of RL in fine-tuning? I'm thrilled to share a pearl we found on the deepest dive of my PhD: the value of RL in RLHF seems to come from *generation-verification gaps*. Get ready to 🤿:

Bild

Lynch was perhaps my favorite director: it was like he lifted Murakami to the silver screen. His work had enough surrealism that his message stuck, but not so much that it was hidden. To quote my favorite Twin Peaks episode: "this is the water and this is the well, drink full and descend." RIP.

David Lynch, Visionary Director of ‘Twin Peaks’ and ‘Blue Velvet,’ Dies at 78

Director David Lynch, who radicalized American film with with a dark, surrealistic artistic vision in films like 'Blue Velvet,' has died. He was 78.

variety.com

Congrats to my ever-amazing undergrad Juntao Ren (who is far too productive to be on here) for being named a Runner Up for the 2025 CRA Outstanding Undergraduate Researcher Award (cra.org/about/awards...)! He's on the PhD Job Market this year and I can't say enough good things about him!

Excited to be at NeurIPS'24, where I'll be presenting at several workshops! Looking forward to chatting about (time series & tabular) foundation models, data science agents, or ML for healthcare! Also, I'm also on the industry job market, looking forward to connect 😁!

Bild

LLM self-improvement has critical implications in synthetic data, post-training and test-time inference. To understand LLMs' true capability of self-improvement, we perform large-scale experiments with multiple families of LLMs, tasks and mechanisms. Here is what we found: (1/9)

Bild

This is a good post (the birds don't hurt either 🦜). Lately, I've been thinking a lot about how a similar principle applies in modern ML research (even in academia). I'm still not sure if this is a good thing, but it is the world we live in.

🐥pdrm🐣@pedramnavid.com · 2y ago

I’ve seen many good data people and engineers fall into the trap of thinking being right is enough, but being effective is the real unlock to career growth. Wrote some thoughts here: open.substack.com/pub/pedram/p...

@gswamy.bsky.social et al propose SPO which builds a game from a preferences, solving for the minimax winner. Handles non-Markovian, intransitive, and stochastic preferences. Nice empirical eval ranging from small demonstrative domains to huge RL domain (Mujoco). arxiv.org/abs/2401.04056 2/3.

A Minimaximalist Approach to Reinforcement Learning from Human Feedback

We present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback. Our approach is minimalist in that it does not require training a reward model nor unst...

arxiv.org