Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n
Andrew Saxe
@saxelab.bsky.social
Professor at the Gatsby Unit and Sainsbury Wellcome Centre, UCL, trying to figure out how we learn
Confused about all this talk of compositionality in neuroscience? Read our new perspective www.nature.com/articles/s41... with authors Reidar Riveland and Alex Pouget.
The compositionality continuum as a principle for studying the neural basis of intelligence - Nature Neuroscience
Compositionality exists on a continuum of increasing complexity rather than as a binary trait. Studying its implementation in simpler biological and artificial systems is the most tractable path towar...
nature.com
If an action results in error, each neuron requires an individualized teaching signal that guides change in its output. This is the credit assignment problem of learning. Are there neurons in the brain that can compute such a sophisticated teaching signal? Yes. www.biorxiv.org/content/10.6...
Climbing fibers encode the gradient of a loss function for the cerebellum
Neurons in the brain are often many synapses away from motoneurons, yet if a movement results in error, each distant neuron needs a teacher that considers its specific contribution to production of th...
biorxiv.org
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning. In our new @icmlconf.bsky.social paper, we address this gap! 📅 July 9th, Poster #4502 Session 8! 🧵 arxiv.org/pdf/2602.20062
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization. Led by Sara Dragutinović and advised by Rajesh Ranganath arxiv.org/abs/2603.00742
Ever wondered how the hippocampal cognitive map is read-out? @changmin-yu.bsky.social @zilong-ji.bsky.social with Jake & John, show that, during navigation, theta sweeps indicate remembered goal-directions (cf. current/next movements or perceptual targets) 1/2 www.nature.com/articles/s41...
1/7 Excited to share my last PhD article, just accepted to ICML 2026! In it, we (me, Alexandre Payeur, Guillaume Lajoie) used dynamical systems theory to study "local" learning in linear recurrent neural networks. See link for the paper, and thread for a brief summary. arxiv.org/abs/2606.00243
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
Biological and neuromorphic recurrent neural networks (RNNs) are subject to spatial and temporal locality constraints on the information that can plausibly be used during learning. A common strategy t...
arxiv.org
Come chat about this @iclr-conf.bsky.social! Friday 3:15 PM, Pavilion 4, Poster #4216
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
We’ve got an exciting new thing to share! We have causal evidence (using TMR) that memory reactivation during sleep promotes abstract understanding of underlying structure, allowing transfer learning in a new domain with zero superficial feature overlap with the learned one.
Super excited to share this preprint! How do we disentangle underlying structure from the particular features of a learning episode to benefit future learning? We find that memory reactivation during sleep promotes this structure abstraction process. www.biorxiv.org/content/10.6...
New preprint! 🧠 How do RNNs learn abstract rules from sequences, independent of specific stimuli? By Vezha Boboeva, with Alberto Pezzotta & George Dimitriadis "From sequences to schemas: low-rank recurrent dynamics underlie abstract relational representations" www.biorxiv.org/content/10.6...
Two Analytical Connectionism-related updates: 1. ⏰ 1 week left to apply! Interested in language + AI & cognition? Don’t miss it: www.analytical-connectionism.net/school/2026/ 2. 📜 Lecture notes from the first two editions are finally out: proceedings.mlr.press/v320/
2026 School on Analytical Connectionism
A 2-week summer course hosted at Chalmers University of Technology on analytical approaches to language acquisition and higher-level cognition.
analytical-connectionism.net
📢 We’re now accepting applications for the 2026 School on Analytical Connectionism dedicated this year to Language Acquisition. 📍 Gothenburg, Sweden 🗓️ August 17–28, 2026 ☠️ Apply by April 17! 🔗 analytical-connectionism.net/school/2026/ 👇 Meet the experts joining us this summer!
We’re hiring a Group Leader! Join us to lead a transformative initiative in human systems neuroscience. Find out more and apply ⤵️ www.sainsburywellcome.org/content/curr...
Postdoc opening! Come work with us on deep learning theory relevant to AI safety Deadline: 7 Apr 2026 Details and application: www.ucl.ac.uk/work-at-ucl/...
UCL – University College London
UCL is consistently ranked as one of the top ten universities in the world (QS World University Rankings 2010-2022) and is No.2 in the UK for research power (Research Excellence Framework 2021).
ucl.ac.uk
Very excited by this year's Analytical Connectionism Summer School! A dream lineup of speakers on the topic of language acquisition in minds and machines Bursaries available to cover costs Aug 17 – Aug 28, 2026 Gothenburg Details: www.analytical-connectionism.net//school/2026/
A great entry into the proposals available for physiologically plausible gradient descent! I think the way they use dendrite targeting inhibition in this model is particularly elegant. Time to start testing these ideas folks!!! #neuroscience 🧪 #NeuroAI
Our latest publication grapples with how the brain could implement gradient descent by sending learning targets top-down, gating plasticity with dendritic inhibition, and updating synaptic weights with biologically observed learning rules like BTSP. www.cell.com/cell-reports...
The First 1,000 Days (1kD) Project - Collecting and Analyzing an Ultra-Dense Naturalistic Dataset of Human Baby Development https://www.biorxiv.org/content/10.64898/2026.03.19.712982v1
Looking for alternatives to quadratic functions for closed-form analysis in optimization? This post explores matrix Riccati dynamics and their applications to neural networks. francisbach.com/closed-form-...
Here's a lovely #blueprint on a new study from our lab led by @royeyono.bsky.social. tl;dr: it implies that there may be interneurons whose role is to normalize credit assignment signals during learning. #neuroscience 🧪
How do neural circuits in the brain implement normalization? 🧠 In our new paper, we show that just normalizing sensory input isn't enough. Crucially, we must also normalize the error signals! 🧵👇 Paper: arxiv.org/abs/2603.17676
A new Department of Cognitive Science is being created at Bocconi University in Milan, Italy. Here is the call for a cluster hire (for around 10 faculty) in all areas of cognitive science, at both junior and senior levels: www.unibocconi.it/en/faculty-a... Deadline: May 4th, 2026
Open Rank Faculty Cluster Hire Search for the New Department of Cognitive Science at Bocconi - Bocconi University
unibocconi.it
Poster tonight at #cosyne26 (1-079)! @wanqingjiang.bsky.social & @noehamou.bsky.social show that mice learn hidden community structure in a 15-odour graph even when transition statistics are flat. Fun collaboration with @saxelab.bsky.social that started with East London coffees ☕!
Really neat work by Fountas and colleagues at UCL: arxiv.org/abs/2603.04688 They propose that consolidation reflects a form of "predictive forgetting" that aids generalization.
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
Standard accounts of memory consolidation emphasise the stabilisation of stored representations, but struggle to explain representational drift, semanticisation, or the necessity of offline replay. He...
arxiv.org
Thanks @natmesanash.bsky.social for covering our new work, in @thetransmitter.bsky.social!
Whether the hippocampus is involved in unrewarded learning has been a controversial question. A new preprint finds that it may be critical for passively learning information. By @natmesanash.bsky.social #neuroskyence www.thetransmitter.org/memory/hippo...
📢📢 Announcing this year's conference on the Mathematics of Neuroscience & AI (Rome, 9-12th June). We’ve got a stellar line-up and venue, and invite everyone to join: www.neuromonster.org
📢 Job alert - Deep Learning Theory & AI Safety Applications open for a postdoc fellow (@saxelab.bsky.social lab) to study artificial deep networks using techniques from applied maths & stat physics. ⏰ Deadline: 26 Mar 2026 🤝 In collaboration with @stefsm.bsky.social ℹ️ www.ucl.ac.uk/life-science...
Excited to be co-organising a #cosyne2026 workshop with Alison Comrie on 'algorithms for learning from scratch'! With a great line-up of speakers, we'll be tackling the question of what processes enable naive biological & artificial agents to adapt to new situations. Info here: tinyurl.com/4u8enf7k
learningfromscratch
march 16th, workshop day 1 @ cosyne 2026
sites.google.com
📢 We’re now accepting applications for the 2026 School on Analytical Connectionism dedicated this year to Language Acquisition. 📍 Gothenburg, Sweden 🗓️ August 17–28, 2026 ☠️ Apply by April 17! 🔗 analytical-connectionism.net/school/2026/ 👇 Meet the experts joining us this summer!
Thrilled to finally share this work! 🧠🔊 Using a new reinforcement-free task we show mice (like humans) extract abstract structure from sound (unsupervised) & dCA1 is causally required by building factorised, orthogonal subspaces of abstract rules. Led by Dammy Onih! www.biorxiv.org/content/10.6...
biorxiv.org
Excited to launch Principia, a nonprofit research organisation at the intersection of deep learning theory and AI safety. Our goal is to develop theory for modern machine learning systems that can help us understand complex network behaviors, including those critical for AI safety and alignment. 1
Our paper is out in @natneuro.nature.com! www.nature.com/articles/s41... We develop a geometric theory of how neural populations support generalization across many tasks. @zuckermanbrain.bsky.social @flatironinstitute.org @kempnerinstitute.bsky.social 1/14