Tomorrow, Zak will talk about his new deep RL method for hard exploration problems. Join us! The talk will be hosted by Csaba.
Dylan Foster 🐢
@djfoster.bsky.social
Principal Researcher in AI/ML/RL Theory @ Microsoft Research NE/NYC. Previously @ MIT, Cornell. http://dylanfoster.net RL Theory Lecture Notes: https://arxiv.org/abs/2312.16730
Huge congratulations to Ilias Diakonikolas, Gautam Kamath, Daniel Kane, Jerry Li, Ankur Moitra, and Alistair Stewart on being awarded the Gödel prize for their breakthrough work on algorithmic robustness! www.sigact.org/prizes/g%C3%...
ACM SIGACT - Gödel Prize
sigact.org
New work in why action chunking is so important for robot control (it helps fight compounding error) arxiv.org/abs/2507.09061
Building Olmo 3 Think Foundations of Reasoning in Language Models @ NeurIPS 2025 Today 13:45 - 14:30
At #NeurIPS2025? Join us for a Social on Wednesday at 7 PM, featuring a fireside chat with Jon Kleinberg and mentoring tables. Ft. mentors @djfoster.bsky.social @surbhigoel.bsky.social @aifi.bsky.social @gautamkamath.com and more!
The coverage principle: How pre-training enables post-training New preprint where we look at the mechanisms through which next-token prediction produces models that succeed at downstream tasks. The answer involves a metric we call the "coverage profile", not cross-entropy.
The new call for Motwani postdocs application is now open! academicjobsonline.org/ajo/jobs/30865 BTW- Not quite ready for a postdoc? We updated the TCS Masters programs spreadsheet: www.cs.princeton.edu/~smattw/mast... Any career stage and in the (SF) Bay Area? Save the date for TOCA-SV on 11/7!
Stanford University, Computer Science/Theory Lab/Stanford University
Job #AJO30865, Postdoc in Theoretical Computer Science at Stanford, Computer Science/Theory Lab/Stanford University, Stanford University, Stanford, California, US
academicjobsonline.org
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking. A totally new framework based on ~backtracking~ for using process verifiers to guide inference, w/ connections to approximate counting/sampling in theoretical CS. Paper: www.arxiv.org/abs/2510.03149
MSR NYC is hiring spring and summer interns in AI/ML/RL! Apply here: jobs.careers.microsoft.com/global/en/jo...
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
microsoft.com
🚨Microsoft Research NYC is hiring🚨 We're hiring postdocs and senior researchers in AI/ML broadly, and in specific areas like test-time scaling and science of DL. Postdoc applications due Oct 22, 2025. Senior researcher applications considered on a rolling basis. Links to apply: aka.ms/msrnyc-jobs
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
aka.ms
Microsoft Research New York City (www.microsoft.com/en-us/resear...) is seeking applicants for multiple Postdoctoral Researcher positions in ML/AI! These are positions for up to 2 years, starting in July 2026. Application deadline: October 22, 2025
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
microsoft.com
Quick reminder: The deadline for our workshop on Foundations of Reasoning in Language Models (FoRLM) at NeurIPS 2025 is next Wednesday, Sept 3!
Announcing the first workshop on Foundations of Language Model Reasoning (FoRLM) at NeurIPS 2025! 📝Soliciting abstracts that advance foundational understanding of reasoning in language models, from theoretical analyses to rigorous empirical studies. 📆 Deadline: Sept 3, 2025
Announcing the first workshop on Foundations of Language Model Reasoning (FoRLM) at NeurIPS 2025! 📝Soliciting abstracts that advance foundational understanding of reasoning in language models, from theoretical analyses to rigorous empirical studies. 📆 Deadline: Sept 3, 2025
For those at ICML, Audrey will be presenting this paper at the 4:30pm poster session this afternoon! West Exhibition Hall B2-B3 W-1009
Is Best-of-N really the best we can do for language model inference? New paper (appearing at ICML) led by the amazing Audrey Huang (ahahaudrey.bsky.social) with Adam Block, Qinghua Liu, Nan Jiang, and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11
ICML's election for their board of directors has begun. I've thrown my hat in the ring. Please consider voting for Gautam Kamath. I have experience with the governance of TMLR, COLT, and ALT, and I think I've demonstrated myself as a consciencious and engaged community member.
This week's #PaperILike is "The Power of Resets in Online Reinforcement Learning" (Mhammedi et al., 2024). If you're doing RL in sim, why not use the sim to its full potential? Reset to any state! (gym.Env.reset() is not all we need.) PDF: arxiv.org/abs/2404.15417
The Power of Resets in Online Reinforcement Learning
Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that require general fun...
arxiv.org
📣Join us at COLT 2025 in Lyon for a community event! 📅When: Mon, June 30 | 16:00 CET What: Fireside chat w/ Peter Bartlett & Vitaly Feldman on communicating a research agenda, followed by mentorship roundtable to practice elevator pitches & mingle w/ COLT community! let-all.com/colt25.html
Hiring a postdoc to scale up and deploy RL-based planning onto some self-driving cars! We'll be building on arxiv.org/abs/2502.03349 and learn what the limits and challenges of RL planning are. Shoot me a message if interested and help spread the word please! Full posting to come in a bit.
Robust Autonomy Emerges from Self-Play
Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi...
arxiv.org
At the IDEAL annual meeting and saw this paper presented. Basically: reducing length of chain of thought LLM computations by deleting intermediate computations, more like classical functional programming where only function call and return values are important. arxiv.org/abs/2503.14337
PENCIL: Long Thoughts with Short Memory
While recent works (e.g. o1, DeepSeek R1) have demonstrated great promise of using long Chain-of-Thought (CoT) to improve reasoning capabilities of language models, scaling it up during test-time is c...
arxiv.org
Dhruv Rohatgi will be giving a lecture on our recent work on comp-stat tradeoffs in next-token prediction at the RL Theory virtual seminar series (rl-theory.bsky.social) tomorrow at 2pm EST! Should be a fun talk---come check it out!!
Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier arxiv.org/abs/2502.12465 New paper (another fun internship project!) with Dhruv Rohatgi, Adam Block, Audrey Huang (ahahaudrey.bsky.social), and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11
Later today, Sikata and Marcel will talk about their recent work on oracle-efficient RL with ensembles. Join us!
The abstract submission deadline for FoPt has been extended to the 21st of May (11:59pm UTC). Submission website: openreview.net/group?id=lea...
Announcing the first workshop on Foundations of Post-Training (FoPT) at COLT 2025! 📝 Soliciting abstracts/posters exploring theoretical & practical aspects of post-training and RL with language models! 🗓️ Deadline: May 19, 2025
Announcing the first workshop on Foundations of Post-Training (FoPT) at COLT 2025! 📝 Soliciting abstracts/posters exploring theoretical & practical aspects of post-training and RL with language models! 🗓️ Deadline: May 19, 2025
Announcing the first workshop on Foundations of Post-Training (FoPT) at COLT 2025! 📝 Soliciting abstracts/posters exploring theoretical & practical aspects of post-training and RL with language models! 🗓️ Deadline: May 19, 2025
Is Best-of-N really the best we can do for language model inference? New paper (appearing at ICML) led by the amazing Audrey Huang (ahahaudrey.bsky.social) with Adam Block, Qinghua Liu, Nan Jiang, and Akshay Krishnamurthy (akshaykr.bsky.social). 1/11
Last seminars before the summer break: 04/29: Max Simchowitz (CMU) 05/06: Jeongyeol Kwon (Univ. of Widsconsin-Madison) 05/20: Sikata Sengupta & Marcel Hussing (Univ. of Pennsylvania) 05/27: Dhruv Rohatgi (MIT) 06/03: David Janz (Univ. of Oxford) 06/10: Nneka Okolo (MIT)
What is the place of exploration in today's AI landscape and in which settings can exploration algorithms address current open challenges? Join us to discuss this at our exciting workshop at @icmlconf.bsky.social 2025: EXAIT! exait-workshop.github.io #ICML2025
Reinforcement learning has led to amazing breakthroughs in reasoning (e.g., R1), but can it discover truly new behaviors not already present in the base model? A new paper with Zak Mhammedi and Dhruv Rohatgi: The Computational Role of the Base Model in Exploration arxiv.org/abs/2503.07453
Join us tomorrow to attend Vlad's presentation! Related to the seminar from last week, but this time in the offline setting. Tuesday March 25, 6 PM UTC.