Prithviraj "Raj" Ammanabrolu

@rajammanabrolu.bsky.social

AI, RL, NLP, Games Asst Prof at UCSD Research Scientist at Nvidia Lab: http://pearls.ucsd.edu Personal: prithvirajva.com

My former MS student Chris Cui (now PhD student with @rajammanabrolu.bsky.social)motivates Text Adventure Games as testbeds for reasoning. Provides a new benchmark suite of text games. Observes that Zork still kicks LLM’s butts despite training on walkthroughs arxiv.org/abs/2504.14128

TALES: Text Adventure Learning Environment Suite

Reasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse reasoning capabiliti...

arxiv.org

🔥Excited to share our new work: "A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning"! We study what actually works for agentic multi-turn RL with varying 🌎Environment, 🤖Policy, and ⭐Reward. We conduct various ablations and empirical analysis on 🧩TextWorld, 🧙ALFWorld, and 🧑‍💻SWE-Gym.

I recently left Mosaic/Databricks Research. It's been a ride building out the RL team from <4 ppl to 20+ across two companies & acquisition +figuring out RL as a Service in prod. Mosaic had insane talent density Some "relaxation" while I put out Prof fires for a smol bit then new adventures!

BildBildBild

The thing that feels so off about the core tech world is that every convo is very transactional. Maybe true elsewhere too. "Oh you're an expert in RL, can you answer questions about my new startup?" Every single (Bay) party. No I do not want to consult. I just wanna hang out.

Bild

I've heard this personally from multiple PMs at AI companies. Students are one of the biggest demographics and they need to "break in" and have even more usage to improve their metrics. Classic corporate economic incentives

Mark Riedl@markriedl.bsky.social · last yr.

AI companies in the US gave access to their systems to students for free during college exams China disabled access to AI systems during nationwide college exams www.theverge.com/news/682737/... Feel free to draw your own conclusions

"Foundation" models for embodied agents are all the rage but how to actually do complex looong context reasoning? Can we scale Beyond Needle(s) in the (Embodied) Haystack? ∞-THOR is an infinite len sim framework + guide on (new) architectures/training methods for VLA models

What's with these arguments over whether X or Y or whatever was the first LLM RL library? These all came in the last 3 months We wrote multiturn RL4LMs like 3+ years ago github.com/allenai/RL4LMs There were other simple versions even before. ML ppl approaching goldfish memory

GitHub - allenai/RL4LMs: A modular RL library to fine-tune language models to human preferences

A modular RL library to fine-tune language models to human preferences - allenai/RL4LMs

github.com

Yeah this is actual unironic advice for any 1st years: you seriously need to go get beers or otherwise hangout socially with more senior grad students (no one else will really know) and hopefully get ensconced enough that someone will tell you “hey just fyi Dr. GoodPubs is low-key a total sociopath”

BBrian Chang@brianchang.bsky.social · last yr.

It’s another way nepotism manifests? You really can’t know who would be a good or bad advisor other than rumors. If you have parents in academia, they can help you sift through the rumors