Ryan Truong

@heyodogo.bsky.social

ryanvt.com

We present PlayTrain, a new Reinforcement Learning (RL) framework that reimagines how we build video-game environments for RL by harnessing LLMs’ ability to write JavaScript games. PlayTrain turns single-file JS browser games into RL environments ready for efficient training.

“If something is boring after two minutes, try it for four. If still boring, then eight. Then sixteen. Then thirty-two. Eventually one discovers that it is not boring at all.” ― John Cage

Today we present a new framework for measuring human-like general intelligence in machines: studying how and how well they play and learn to play all conceivable human games compared to humans. We then propose the AI Gamestore a way to sample from popular human games to evaluate AI models.

Bild
L

Thanks Sam! Main takeaways: 1) Ground-truth vs. predictive model selection differ under noisy and scarce data—for prediction, oversimplified models may work better in avoiding overfitting. 2) When humans decide between externally provided, prefitted predictive models, they're undersensitive to 1).

Sam Gershman@gershbrain.bsky.social · 9mo ago

New work from @liushuze.bsky.social and @yangxiang.bsky.social osf.io/preprints/ps... People violate Occam's razor when selecting between predictive models. This is surprising given past research (including my own) showing a preference for simplicity.

It’s grad school application season, and I wanted to give some public advice. Caveats: -*-*-*-* 
> These are my opinions, based on my experiences, they are not secret tricks or guarantees 
> They are general guidelines, not meant to cover a host of idiosyncrasies and special cases

arxiv.org/abs/2510.11144 "Using teacher models that answer at varying levels of abstraction, from executable action sequences to high-level subgoal descriptions, we show that lifelong learning agents benefit most from answers that are abstracted and decoupled from the current state."

$How^{2}$: How to learn from procedural How-to questions

An agent facing a planning problem can use answers to how-to questions to reduce uncertainty and fill knowledge gaps, helping it solve both current and future tasks. However, their open ended nature, ...

arxiv.org