Tom Schaul

@schaul.bsky.social

RL researcher at DeepMind https://schaul.site44.com/ 🇱🇺

Exclusive: More than 580 Google workers including hundreds of AI researchers urge CEO Sundar Pichai to refuse classified military AI work, after senior Pentagon official tells me he is already in talks with Google. Our story on Bloomberg now, from me and Julia Love

When faced with a challenge (like debugging) it helps to think back to examples of how you've overcome challenges in the past. Same for LLMs! The method we introduce in this paper is efficient because examples are chosen for their complementarity, leading to much steeper inference-time scaling! 🧪

Aditi Mavalankar@aditimavalankar.bsky.social · last yr.

Excited to share our recent work, AuPair, an inference-time technique that builds on the premise of in-context learning to improve LLM coding performance! arxiv.org/abs/2502.18487 🧵

This year's (first-ever) RL conference was a breath of fresh air! And now that it's established, the next edition is likely to be even better: Consider sending your best and most original RL work there, and then join us in Edmonton next summer!

Reinforcement Learning Conference@rl-conference.bsky.social · 2y ago

The call for papers for RLC is now up! Abstract deadline of 2/14, submission deadline of 2/21! Please help us spread the word. rl-conference.cc/callforpaper...

Are there limits to what you can learn in a closed system? Do we need human feedback in training? Is scale all we need? Should we play language games? What even is "recursive self-improvement"? Thoughts about this and more here: arxiv.org/abs/2411.16905

Boundless Socratic Learning with Language Games

An agent trained within a closed system can master any desired capability, as long as the following three conditions hold: (a) it receives sufficiently informative and aligned feedback, (b) its covera...

arxiv.org