How far can we compress billion-parameter LLMs? We introduce requential coding, which achieves < 1-bit per param compression, and explains why scaling doesn't hit a generalization wall! arxiv.org/pdf/2607.11883 w/ Shikai Qiu, Marc Finzi, Yujia Zheng, Kun Zhang 1/🧵
Andrew Gordon Wilson
@andrewgwils.bsky.social
Machine Learning Professor https://cims.nyu.edu/~andrewgw
In cooking, execution is more important than the dish itself, even for simple dishes. Hummus can be great or terrible. The same is true of scientific ideas. Almost nothing works as we wish at first. Persistence, high standards, and attention to detail make all the difference.
📍 In person at COLM 2026, San Francisco 🗓️ Submission deadline: June 23, 2026, 11:59 PM AoE 🌐 science-ai-2026.github.io 🧑🏫 Speakers: @suryaganguli.bsky.social, Jikai Jin, Zhiyuan Li, @hectorliu.bsky.social, @valentinapy.bsky.social , Ludwig Schmidt, MohammadShoeybi, @andrewgwils.bsky.social
Scientific Understanding of Foundation Models | COLM 2026
A workshop on building rigorous scientific understanding of foundation models — from scaling laws and emergent capabilities to principled evaluation and mechanistic explanation.
science-ai-2026.github.io
Anyone want to submit a workshop proposal all about deep learning? I think this area really might take off.
Perhaps I'm an outlier, but generally the value I derive from art is not from its backstory. I love a Bach fugue not because he was suffering, content, had many children, or whatever else, but because it's an extraordinary composition. I'd feel the same about AI generated art.
How much does a language model forget when finetuned on new tasks? We show both model size and optimization matter and forgetting can be nearly eliminated with self-generated replay! arxiv.org/abs/2605.26097 w/Martin Marek, Dongkyu Cho, Shikai Qiu, Rumi Chunara, and Pavel Izmailov. 1/8
May all of your NeurIPS submissions be high epiplexity.
"Does it still make sense to get a CS degree?" A CS degree has never been primarily about software engineering. It's about core skills, about learning how to think. That never goes obsolete. But really you should get a physics degree.
Never be embarrassed about explaining something basic. The best work has no pretense, no ego.
Me in every meeting: "have you considered epiplexity?"
Using advanced AI optimizers like Muon doesn’t have to rely on guesswork. Courant PhD students Shikai Qiu and Zixi (Charlie) Chen, CDS PhD Student Hoang Phan, CDS Asst. Prof. Qi Lei, and CDS Prof. @andrewgwils.bsky.social bridge theory and practice. nyudatascience.medium.com/building-the...
Building the Science of Scaling: Improving the Efficiency of Deep Learning Optimizers
A profound regime change in the field of optimization may be around the corner. For a decade, the Adam optimizer has overwhelmingly…
nyudatascience.medium.com
For the most part, people see what they want to see. If they want to find fault, they will. If they want to be supportive, they will. Smart people can convincingly rationalize virtually any position. But underneath it all often lies something fundamentally irrational, and far from objective.
Alec Radford (and others behind GPT, let's not forget there were other authors) deserve credit. Conventional wisdom said it shouldn't work well. It didn't work well. They got brutal feedback: stop wasting time building a glorified autocomplete. But they persisted and the results were mindblowing.
There's a new generation of empirical deep learning researchers, hacking away at whatever seems trendy, blowing with the wind... no accumulation of real understanding, or foundations. No real passion or depth, just light amusement and career advancement. I'm hoping it's a phase.
I don't like how the world is becoming increasingly isolating and impersonal. I don't want to scan a QR code with my phone to order at a restaurant. I'd like to talk with a person. Expediency isn't all that matters. Am I alone in this?
What if Watson & Crick discovered the double helix structure of DNA at Nando's instead of The Eagle pub? Would they have a commemorative perinaise, or stick with the plaque?
We introduce epiplexity, a new measure of information that provides a foundation for how to select, generate, or transform data for learning systems. We have been working on this for almost 2 years, and I cannot contain my excitement! arxiv.org/abs/2601.03220 1/7
One of the underrated papers this year: "Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful" (arxiv.org/abs/2507.07101) (I can confirm this holds for RLVR, too! I have some experiments to share soon.)
Excited about our new paper that unifies discrete, Gaussian, and simplicial diffusion, enabling model comparison, likelihood evaluation, stable training, and more, including a DNA design application! Amazing work from @alannawzadamin.bsky.social, Alina, Lily, and team! arxiv.org/abs/2512.15923
A Unification of Discrete, Gaussian, and Simplicial Diffusion
To model discrete sequences such as DNA, proteins, and language using diffusion, practitioners must choose between three major methods: diffusion in discrete space, Gaussian diffusion in Euclidean spa...
arxiv.org
Thrilled to start 2026 as faculty in Psych & CS @ualberta.bsky.social + Amii.ca Fellow! 🥳 Recruiting students to develop theories of cognition in natural & artificial systems 🤖💭🧠. Find me at #NeurIPS2025 workshops (speaking coginterp.github.io/neurips2025 & organising @dataonbrainmind.bsky.social)
Excited to be speaking at the SPIGM workshop at NeurIPS tomorrow, 10:30-11 am, Room 20C. My talk will be "Probabilistic Inference is the Future of Foundation Models". See you there! spigmworkshopv3.github.io/schedule/
A nice list. But, it doesn't actually go much beyond electronics. In terms of quality of life, I think some of these "conveniences" are a downgrade in practice. I miss blockbuster. I miss watching my favourite shows when they aired on a TV schedule. I miss 90s gaming. I miss being able to unplug.
Lot of people should read this exhaustive list of life improvements from the 90s: gwern.net/improvement
My full interview with MLStreetTalk has just been posted. I really enjoyed this conversation! We talk about the bitter lesson, scientific discovery, Bayesian inference, mysterious phenomena, and key principles for building intelligent systems. www.youtube.com/watch?v=M-jT...
The Real Reason Huge AI Models Actually Work
YouTube video by Machine Learning Street Talk
youtube.com
I'm excited to be giving a keynote talk at the AutoML conference 9-10 am at Cornell Tech tomorrow! I'm presenting "Prescriptions for Universal Learning". I'll talk about how we can enable automation, which I'll argue is the defining feature of ML. 2025.automl.cc/program/
Research doesn't go in circles, but in spirals. We return to the same ideas, but in a different and augmented form.
CDS/Courant Professor Andrew Gordon Wilson (@andrewgwils.bsky.social) argues mysterious behavior in deep learning can be explained by decades-old theory, not new paradigms: PAC-Bayes bounds, soft biases, and large models with a soft simplicity bias. nyudatascience.medium.com/deep-learnin...
Deep Learning’s Most Puzzling Phenomena Can Be Explained by Decades-Old Theory
Andrew Gordon Wilson argues that many generalization phenomena in deep learning can be explained using decades-old theoretical tools.
nyudatascience.medium.com
Regardless of whether you plan to use them in applications, everyone should learn about Gaussian processes, and Bayesian methods. They provide a foundation for reasoning about model construction and all sorts of deep learning behaviour that would otherwise appear mysterious.
A common takeaway from "the bitter lesson" is we don't need to put effort into encoding inductive biases, we just need compute. Nothing could be further from the truth! Better inductive biases mean better scaling exponents, which means exponential improvements with computation.
Gould mostly recorded baroque and early classical. He only recorded a single Chopin piece, as a one-off broadcast. But like many of his efforts, it's profoundly thought provoking, the end product as much Gould as it is Chopin. I love the last mvt (20:55+). www.youtube.com/watch?v=NAHE...
Glenn Gould plays Chopin Piano Sonata No. 3 in B minor Op.58
YouTube video by The Piano Experience
youtube.com