Mengye Ren

@mengyer.bsky.social

Assistant Professor of CS & DS at NYU. Machine Learning, Human-like AI, Continual Learning | Head of @agentic-ai-lab.bsky.social mengyeren.com

CDS Assistant Professor Mengye Ren (@mengyer.bsky.social) argues in a new paper that AI selfhood requires continual learning. Today’s LLMs wake from amnesia each session and read a diary of a past self — never extending themselves through new experience. nyudatascience.medium.com/to-have-a-se...

To Have a Self, an AI Must Live a Life

Every copy of a large language model begins each conversation identical to every other copy. They share the same weights and the same…

nyudatascience.medium.com

What does it mean to create a new concept rather than retrieve a familiar one? I propose that creativity is what's unfamiliar at first but quickly learnable by an adaptive observer, and show that meta-learning through a frozen Diffusion model produces stylistic & conceptual creations.

BildBildBild

AI agents often struggle to plan movements because their internal representations of the physical world can be overly tangled. CDS PhD student Ying Wang shows how straightening these pathways improves AI navigation. Accepted to ICML 2026. nyudatascience.medium.com/improving-wo... 1/2

Improving World Models: A Neuroscience-Inspired Approach to Latent Planning

Humans instinctively map out the physical consequences of their actions before taking them, seamlessly predicting that a dropped glass will…

nyudatascience.medium.com

I have updated my tutorial on making Vision Language Action models. This tutorial starts with a basic Transformer and walks people through the steps to transform it into a full VLA that uses PaliGemma as the pretrained VLM. Links below.

Bild

Corporate PRs are becoming a disservice to science. We see amazing things with no idea how they were done. It's just a way to grab smart people and pump equity, while discouraging junior students by making them think there's nothing left to be done in research.

Check out our Midway Network paper on learning hierarchical latent motion tokens from watching videos. Recently got accepted to ICLR 2026!

Chris Hoang@choang.bsky.social · 6mo ago

Animals learn to recognize objects and how they move from observation. SSL on videos emulates “learning by observing” but only for recognition or motion, not both! Our #ICLR2026 work Midway Network is the first to learn both recognition and motion understanding from videos via latent dynamics 🧵

Verifiers are increasingly being used today in RL to provide rewards. We did a systematic study on when it is the best to use LLMs to verify solutions.

NYU Center for Data Science@nyudatascience.bsky.social · 6mo ago

Do stronger LLMs make better verifiers? Not necessarily when grading themselves. New work led by Courant PhD student @jacklu-me.bsky.social and CDS Asst Prof @mengyer.bsky.social shows that cross-family verification outperforms self-verification. nyudatascience.medium.com/study-reveal...

Our latest research Midway Networks learn recognition and motion representations from scratch by letting the network learn from watching videos. The latent motion vectors are refined in a top-down hierarchy. Interesting tracking results using our forward perturbation viz.

BildBild
NYU Center for Data Science@nyudatascience.bsky.social · 8mo ago

Research from CDS Asst Prof @mengyer.bsky.social and Courant PhD student Christopher Hoang shows how the Midway Network learns object recognition and motion jointly from raw video, using motion latents and a gating unit to model real dynamics. nyudatascience.medium.com/watching-the...

We are so comfortable with the concept of pretraining in foundation models today that we assume an AI is supposed to have seen everything humanity has created. 1/2

The proposed 5% remittance tax on non-citizens is another blatant attack on foreign workers, who've contributed tremendously to the U.S. economy, have little path to citizenship, no birthright for their kids, and will face double taxation on their hard-earned dollars.

How can we leverage naturalistic videos for visual SSL? Naturalistic, i.e. uncurated, videos are abundant and can emulate the egocentric perspective. Our paper at ICLR 2025, PooDLe🐩, proposes a new SSL method to address the challenges of learning from naturalistic videos. 🧵

Bild