In new post, I write about how we misunderstand the famous lottery ticket hypothesis. We tend to think that a network needs to be exponentially large for there to be a subnetwork to win the lottery. This is incorrect---a small network suffices! I use a dart throwing analogy to make sense of this.
Vaishnavh Nagarajan
@vaishnavh.bsky.social
Foundations of AI. I like simple and minimal examples and creative ideas. I also like thinking about the next token 🧮🧸 Google | PhD, CMU | https://arxiv.org/abs/2504.15266 | https://arxiv.org/abs/2403.06963 vaishnavh.github.io
In my next blogpost, I write about how I view technical communication: it's like trying to communicate an escape route to someone without a map but with a catch: you're not with them. You only have a walkie-talkie. Also, they're in panic.
Consolidated my armchair thoughts about "how may an LLM (not) differ from a human who thinks in images/text?". I split this as 2 qns: - is text sufficient to be correct about say, a circle? - does correctness imply sharing the same "understanding" as humans? vaishnavh.github.io/blog/what-ll...
In practice, scientists do not seem to behave like "neutral, rational agents" but rather behave like "zealous advocates" for an idea that has "hired them". I wrote about how I think this "courtroom" view of science works and what I learned from it! vaishnavh.github.io/blog/emotion...
Also, what's the catch with punishing bad reviews by preventing future submissions? Say: if your reviews are egregiously bad as flagged by multiple ACs across at least two conferences, you won't be able to submit papers to the next N conferences. (Possible that I'm missing something here.)
Curious why conferences don't have a system where the authors of every paper together guarantee N reviews per paper (and they can distribute the load amongst themselves). This way wouldn't we tax authors in proportion to the number of papers they burden the system with?
Curious why conferences don't have a system where the authors of every paper together guarantee N reviews per paper (and they can distribute the load amongst themselves). This way wouldn't we tax authors in proportion to the number of papers they burden the system with?
A recent paper (arxiv.org/abs/2602.18671) made me question something basic: do the logits of a language model model the next-token or the full sequence distribution? It really messed with my brain (in a fun way!). I wrote about the paper to clarify my thinking. vaishnavh.github.io/blog/joint-o...
What does a language model model? - Vaishnavh Nagarajan
TL;DR: Does the next-token logit track the conditional or the joint probability of the whole sequence?I had an invisi...
vaishnavh.github.io
Really liked this paper which ties up two observations that are equally mindboggling (low-rank logits & subliminal/weird generalization effects) and presents one other such observation arxiv.org/abs/2602.04863
Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understa...
arxiv.org
The visual world is composed of objects, and those objects are composed of features. But do VLMs exploit this compositional structure when processing multi-object scenes? In our 🆒🆕 #ICLR2026 paper, we find they do – via emergent symbolic mechanisms for visual binding. 🧵👇
Currently reading "a mathematician's apology" by GH Hardy. This is excerpt the foreword by CP Snow describing Hardy's personality and his work:
in associative memory, the latent space doesn't really encode any interesting distance. imagine you're trying to store which countries share borders. you could simply write down a list of adjacent countries OR you could visualize the world map in your head. this is "associative" vs "geometric".
Rare to see such long term efforts these days 🫡
This was a colossal multi-year effort driven by an incredible team that gave this everything: Marc Finzi, Shikai Qiu, Yiding Jiang, Pavel Izmailov, Zico Kolter. Much more in the paper! arxiv.org/abs/2601.03220 7/7
We introduce epiplexity, a new measure of information that provides a foundation for how to select, generate, or transform data for learning systems. We have been working on this for almost 2 years, and I cannot contain my excitement! arxiv.org/abs/2601.03220 1/7
Please welcome Google's Open Source efforts to Blue Sky at @opensource.google!
1/ We found that deep sequence models memorize atomic facts "geometrically" -- not as an associative lookup table as often imagined. This opens up practical questions on reasoning/memory/discovery, and also poses a theoretical "memorization puzzle."
If X, Y, Z are iid high-dim Gaussian N(0, I), what's the angle between X-Y and Z-Y? A. Concentrates at 90 deg B. Concentrates, NOT at 90° C. Doesn't concentrate anywhere. My (and most people's) instincts got this wrong! vaishnavh.github.io/blog/high-di...
Angles between high-dimensional vectors - Vaishnavh Nagarajan
Switch off your brain and answer this:Given three points $\mathbf{X}, \mathbf{Y}, \mathbf{Z}$ sampled from a high-dim...
vaishnavh.github.io
Congratulations to CSD faculty Aditi Raghunathan and her research collaborators on receiving an ICML Outstanding Paper award for Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (icml.cc/virtual/2025...). Paper: arxiv.org/abs/2504.15266
ICML 2025 AwardsICML 2025
icml.cc
Reading the dedications of a PhD thesis is often a cure for a bad day. There’s so much affection in them
As NeurIPS review deadline is around the corner, please remember that you cannot use any non-local LLM like chatgpt/gemini for understanding the paper and drafting/revising your review as that breaks the confidentiality agreement. NeurIPS 2025 Official LLM Policy: neurips.cc/Conferences/...
LLM Policy
neurips.cc
I really enjoyed "When We Cease to Understand the World", although it's more fiction than history of science
“Science in history” by Bernal is my first recommendation. The work of Ian Hacking is a good recommendation for Probability
How do task dynamics impact learning in networks with internal dynamics? Excited to share our ICML Oral paper on learning dynamics in linear RNNs! with @clementinedomine.bsky.social @mpshanahan.bsky.social and Pedro Mediano openreview.net/forum?id=KGO...
Learning dynamics in linear recurrent neural networks
Recurrent neural networks (RNNs) are powerful models used widely in both machine learning and neuroscience to learn tasks with temporal dependencies and to model neural dynamics. However, despite...
openreview.net
When we are doing science, we are unknowingly executing our mythology, taken from movies and friends and textbooks, of what science is. History of science helps us ground that myth in reality
I finally wrote a full-fledged blog about this: reading the history of science is an **amazing** yet under-recognized way to develop (emotional) maturity as a researcher. If you have thoughts/recommendations, please share! vaishnavh.github.io/2025/04/29/h...
I finally wrote a full-fledged blog about this: reading the history of science is an **amazing** yet under-recognized way to develop (emotional) maturity as a researcher. If you have thoughts/recommendations, please share! vaishnavh.github.io/2025/04/29/h...
Why PhD students should read the history of science - Vaishnavh Nagarajan
Here’s a secret that I accidentally discovered during my PhD: consuming history-of-science content is an efficient wa...
vaishnavh.github.io
I weirdly have discovered a similar trick! Something about how fumbling science is across history makes the present sense of fumbling more palatable
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf.bsky.social: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @aaditya6284.bsky.social and Peter Latham
This paper is quite nice. It mixes some useful toy models of creativity with insights about how to induce more creativity in LLMs that are better than greedy sampling
📢 New #paper on creativity & multi-token prediction! We design minimal open-ended tasks to argue: → LLMs are limited in creativity as they learn to predict the next token → creativity can be improved via multi-token learning & injecting noise ("seed-conditioning" 🌱) 1/ #MLSky #AI #arxiv 🧵👇🏽
This isn't fake news. One of the craziest AI research papers I've been on in a while. Weird ablations on RLVR shows that the Qwen 2.5 models can learn with literally random rewards, likely due to some funkiness in mid-training and the GRPO setup.