Sai Prasanna

@saiprasanna.in

See(k)ing the surreal Causal World Models for Curious Robots @ University of Tübingen/Max Planck Institute for Intelligent Systems 🇩🇪 #reinforcementlearning #robotics #causality #meditation #vegan

If open-endedness has to be fundamentally subjectively measured, what are the factors of the agent makes it so if we fix humans as the final arbiter or evaluator. Does embodiment/action space etc of the agent matter for a human evaluator of open-endedness?

I realized how I background process tonnes of information, from work/research and emotional stuff. And it works well, leads to good research ideas, wise processing of tough situations! But It's so hard to learn to trust this as conscious thinking for solving problems feels more under my "control"

TIL: "Clever Hans cheat" for next-token prediction. A subtle but interesting issue with next-token prediction. In the purely forward next token prediction objective, teacher forcing can lead to learning dynamics where the models don't even generalize "in-distribution"!! arxiv.org/abs/2403.06963

The pitfalls of next-token prediction

Can a mere next-token predictor faithfully model human intelligence? We crystallize this emerging concern and correct popular misconceptions surrounding it, and advocate a simple multi-token objective...

arxiv.org

If I have a really good photo that could be potentially used in many contexts, what's the best place to make money with it? My friend has a really good eye for photos and we want to try a side venture selling some of her stuff

Looks like a cool study. Lots to learn from ants about large scale coordination www.pnas.org/doi/10.1073/... "Our results exemplify how simple minds can easily enjoy scalability while complex brains require extensive communication to cooperate efficiently." h/t @petersuber.bsky.social

Comparing cooperative geometric puzzle solving in ants versus humans | PNAS

Biological ensembles use collective intelligence to tackle challenges together, but suboptimal coordination can undermine the effectiveness of grou...

pnas.org

Does augmenting ourselves with V/LLMs to cognitive gaps make self actualization even more difficult on average? Stands stark in contrast with (more difficult/slower to show positive outcomr) augmentation strategies like meditation or psychedelics

One thing coding with LLMs has helped me a lot during the past months is for visualisations. I'm churning out code to visualize many aspects of agent behavior which I wouldn't have done before due to my mental friction in writing such code. Such code to do visualisations is also easy to verify.

When predicting discrete joint distribution of two variables with a neural network, what loss is the best to use? KL on the joint and two marginals? Or is there anything better?

In an effort to play a small part in creating additional value on this site, I'm going to post one-per-day a paper we wrote that was published in 2024. Together with memes. Skipping holidays/weekends. In random order. Would love your thoughts on them. I'll keep them threaded for easy finding! >

As an analogy, I also think we should prioritize decarbonizing electricity over making it renewable (e.g. I think we should use nuclear power for the foreseeable future). And the reason is simple: resources are finite, and climate change is the more pressing problem.

1/ You may have seen us talk about formaldehyde — a chemical that causes an inescapable cancer risk for everyone in America. It’s in the air we breathe. And it’s in our homes: our couches, our clothes, even babies’ cribs. So what can you do to reduce your exposure? THREAD 🧵

How much of the benefit of learning purely in simulations like Dreamer etc, comes from the fact that the shitty model at the start of the training somehow allows the agent to explore large parts of the state space efficiently?

An updated intro to reinforcement learning by Kevin Murphy: arxiv.org/abs/2412.05265! Like their books, it covers a lot and is quite up to date with modern approaches. It also is pretty unique in coverage, I don't think a lot of this is synthesized anywhere else yet

Reinforcement Learning: An Overview

This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based RL, policy-gradient methods, model-based met...

arxiv.org

If you're in NeurIPS, you gotta meet Lennart! He's one of the friendliest folks who can describe a new concept to you in a must clear, & intuitive way, while also listening to your ideas with a similar intensity! Maybe your next nobel collaborator, you never know 😉

Lennart Purucker@lennartpurucker.bsky.social · 2y ago

Excited to be at #NeurIPS2024 tomorrow! 🎉 Let’s connect if you are interested in tabular data and: 🤖 AutoML (e.g., AutoGluon) 📊 Data Science (e.g., LLMs for Feature Engineering) 🏛️ Foundation Models (e.g., TabPFN) Looking forward to insightful discussions—feel free to reach out!