Saurav Jha

@saurav-jha.bsky.social

www.sauravjha.com.np 🇳🇵 in 🇨🇦 IVADO postdoc @mila-quebec.bsky.social Ex applied scientist #Openstream.ai; Ex Intern at #Tencent, #Sony, #Inria; PhD @unsw.edu.au 🇦🇺 ; Ex MLE #Factset

Streaming Reinforcement Learning (RL) is a huge challenge: transitions are used once and discarded immediately. This makes agents extremely sample-inefficient. But what if we could "squeeze" more information out of every single frame? Check out our latest paper!

Bild

New work, just accepted @ICLR: "The Expressive Limits of Diagonal SSMs for State-Tracking" We give a complete characterization of what diagonal SSMs can and cannot compute on state-tracking tasks and the answer is deeply connected to group theory. 🧵👇

Can LLMs play Hangman? Spoiler alert: Not yet. Check out “LLMs Can’t Play Hangman: On the Necessity of a Private Working Memory for Language Agents”, led by Davide Baldelli, Ali Parviz, AmalZouaq and Sarath Chandar.

Bild

I ran across a busy Sander at a #neurips party with a similar question - he was still patient enough to explain stuff. This talk further clarifies a good amount of my doubts. Recommend watching if you're working on diffusion / LLMs for generation!

Sander Dieleman@sedielem.bsky.social · 2y ago

I've been getting a lot of questions about autoregression vs diffusion at #NeurIPS2024 this week! I'm speaking at the adaptive foundation models workshop at 9AM tomorrow (West Hall A), about what happens when we combine modalities and modelling paradigms. adaptive-foundation-models.org