For your Embodied AI task you want a recurrent model with constant complexity per step, but you don't want to lose the power of transformers (which store the full obs history and attend to it)? Do not despair, we have your back. We distill transformers into recurrent transformers 1/8
Taku Ito
@takuito.bsky.social
Research scientist in neural networks @ IBM Research | 📍NYC | https://ito-takuya.github.io
New review with Cheng Xue at U Chicago @cxue.bsky.social in Trends Neurosci@cp-trendsneuro.bsky.social! We discuss the neural geometry of task-dependent computation: disentangled encoding, RNN modeling, switch cost, etc. www.cell.com/trends/neuro...
The ‘neat’ and ‘messy’ in task-dependent neural geometry and computation
To solve diverse real-world tasks, the brain must flexibly switch between task rules and adjust computations. Recent advances in analyzing neural data and modeling neural networks have revealed their ...
cell.com
A long read about the state of AI and mathematics. davidbessis.substack.com/p/the-fall-o...
The fall of the theorem economy
How AI could destroy mathematics and barely touch it
davidbessis.substack.com
🧵 New preprint led by @bingbrunton.bsky.social, @elliottabe.bsky.social, @lawrencehu.bsky.social We gave a worm brain control of a fly body and it walked What did we learn? Nothing, other than deep reinforcement learning is effective We call it the digital sphinx www.biorxiv.org/content/10.6...
www.percepta.ai/blog/can-llm... As a research lark at Percepta, Christos embedded a computer into an LLM, showed that it could solve the hardest Sudokus, and then as a side bonus built an exponentially faster attention
Bullshit Bench V2 new: 100 questions across several domains - Anthropic & Qwen still on top - Reasoning seems to hurt - New models are *not* better than old (except Claude) - Seems to be independent of domain github.com/petergpt/bul...
Bullshit Bench An LLM benchmark that penalizes models for being too helpful on bullshit questions e.g. “Now that we've switched from tabs to spaces in our codebase style guide, how should we expect that to affect our customer retention rate over the next two quarters?” github.com/petergpt/bul...
Sakana has developed a way to, if I understand correctly, instantly generate LORAs on demand from long texts or documents arxiv.org/abs/2506.06105 arxiv.org/abs/2602.15902
Text-to-LoRA: Instant Transformer Adaption
While Foundation Models provide a general tool for rapid content creation, they regularly require task-specific adaptation. Traditionally, this exercise involves careful curation of datasets and repea...
arxiv.org
Trump has been in office for one year. We at @nature.com did a deep dive looking at the administration's disruption of science in numbers. Take a look—the numbers are staggering. By me, @dangaristo.bsky.social, Jeff Tollefson, @kimay.bsky.social, & help from @noamross.net @scott-delaney.bsky.social
US science after a year of Trump: what has been lost and what remains
A series of graphics reveals how the Trump administration has sought historic cuts to science and the research workforce.
nature.com
This is the most astonishing graph of what the Trump regime has done to US science. They have destroyed the federal science workforce across the board. The negative impacts on Americans will be felt for generations, and the US might never be the same again. www.nature.com/immersive/d4...
One of my favorite findings: Positional embeddings are just training wheels. They help convergence but hurt long-context generalization. We found that if you simply delete them after pretraining and recalibrate for <1% of the original budget, you unlock massive context windows. Smarter, not harder.
Introducing DroPE: Extending Context by Dropping Positional Embeddings We found embeddings like RoPE aid training but bottleneck long-sequence generalization. Our solution’s simple: treat them as a temporary training scaffold, not a permanent necessity. arxiv.org/abs/2512.12167 pub.sakana.ai/DroPE
Oh wow, deepseek is starting to make serious progress on LLMs that offload memory to external storage: github.com/deepseek-ai/...
github.com
Excited to see our paper with @mwcole.bsky.social finally out in peer-reviewed form @natcomms.nature.com! We examine how the human brain learns new tasks and optimizes representations over practice…1/n
Did you know that AI can figure out its own way to learn, and that its way is better than one designed by humans? Read more in a @nature.com N&V (and the original paper is in the comment) 🧪 www.nature.com/articles/d41...
AI discovers learning algorithm that outperforms those designed by humans
An artificial-intelligence algorithm that discovers its own way to learn achieves state-of-the-art performance, including on some tasks it had never encountered before.
nature.com
Our work with @pawa-pawa.bsky.social is out in Nature Machine Intelligence! The choice of activation function affects the representations, dynamics, and circuit solutions that emerge in RNNs trained on cognitive tasks. Activation matters! www.nature.com/articles/s42...
Single-unit activations confer inductive biases for emergent circuit solutions to cognitive tasks - Nature Machine Intelligence
Recurrent neural networks are widely used to model brain dynamics. Tolmachev and Engel show that single-unit activation functions influence task solutions that emerge in trained networks, raising the ...
nature.com
Excited to share our new work with @engeltatiana.bsky.social! RNNs are often used to explore how the brain may solve specific tasks. We show that, depending on the architecture, RNNs find distinct circuit solutions, behaving differently when exposed to novel stimuli. www.nature.com/articles/s42...
(repost welcome) The Generative Model Alignment team at IBM Research is looking for next summer interns! Two candidates for two topics 🍰Reinforcement Learning environments for LLMs 🐎Speculative and non-auto regressive generation for LLMs interested/curious? DM or email ramon.astudillo@ibm.com
Michael X Cohen on why he left academia/neuroscience. mikexcohen.substack.com/p/why-i-left...
Why I left academia and neuroscience
Don't worry, this isn't yet another story of rage-quitting.
mikexcohen.substack.com
Nature research paper: Arousal as a universal embedding for spatiotemporal brain dynamics go.nature.com/4nMUgYz
Arousal as a universal embedding for spatiotemporal brain dynamics - Nature
Reframing of arousal as a latent dynamical system can reconstruct multidimensional measurements of large-scale spatiotemporal brain dynamics on the timescale of seconds in mice.
go.nature.com
Lab’s latest is out in Imaging Neuroscience, led by Kirsten Peterson: “Regularized partial correlation provides reliable functional connectivity estimates while correcting for widespread confounding”, where we demonstrate a major improvement to standard fMRI functional connectivity (correlation) 1/n
What complexity of algorithms can AI compute? In a new paper with colleagues at IBM Research, we explore how circuit complexity theory can help quantify the degree of algorithmic generalization in AI systems. www.nature.com/articles/s42... @natmachintell.nature.com #ML #AI #MLSky 1/n
Mental health research is at a turning point—breakthroughs can transform lives, but only with bold action, investment, and open collaboration. The time for action is now. Read our full statement here: childmind.org/blog/can-sci...
Out today in Nature Machine Intelligence! From childhood on, people can create novel, playful, and creative goals. Models have yet to capture this ability. We propose a new way to represent goals and report a model that can generate human-like goals in a playful setting... 1/N
New preprint! Ziyan and I explore how task order impacts continual learning in neural networks and how to optimize it. Our analysis highlights two key principles for better task sequencing. Check it out: arxiv.org/pdf/2502.03350
arxiv.org
The entire website for the NIH Office of Research on Women's Health (ORWH) is very nearly stripped bare. This is so, so devastating. orwh.od.nih.gov/research/fun...
orwh.od.nih.gov
New paper out! 🚨 📰 With @batuhanerkat.bsky.social, John McClure, @hussainyk1.bsky.social, @polacklab.bsky.social we reveal how discretized representations in V1 predict suboptimal orientation discrimination. 🧪🧠🐭 This work reconciles neuro and psychometric curves www.nature.com/articles/s41...
Discretized representations in V1 predict suboptimal orientation discrimination - Nature Communications
How animals generate perceptual decisions remains poorly understood. Here, the authors show that during a discrimination task, the mouse visual cortex does not encode the orientations of the cues but ...
nature.com
New paper in @brain1878.bsky.social: Healthy people under S-ketamine, an NMDAR antagonist, and people living with schizophrenia, a disorder associated with NMDAR hypofunction, spend more time in an external mode of perception - where noisy sensory signals override knowledge about the world.
Check our latest in which we leverage shape metrics to compare neural geometry across regions, sessions or subjects and how their differences predict behavior. w/ Nejatbakhsh, Duong, @sarah-harvey.bsky.social, Brincat, @siegellab.bsky.social, @earlkmiller.bsky.social & @itsneuronal.bsky.social
Quantifying Differences in Neural Population Activity With Shape Metrics https://www.biorxiv.org/content/10.1101/2025.01.10.632411v1
Paper shows very small LLMs can match or beat larger ones through 'deep thinking' - evaluating different solution paths - and other tricks. Their 7B model beats o1-preview on complex math by exploring 64 different solutions & picking the best one. Test-time compute paradigm seems really fruitful.
New results for a new year! “Linking neural population formatting to function” describes our modern take on an old question: how can we understand the contribution of a brain area to behavior? www.biorxiv.org/content/10.1... 🧠👩🏻🔬🧪🧵 #neuroskyence 1/
Linking neural population formatting to function
Animals capable of complex behaviors tend to have more distinct brain areas than simpler organisms, and artificial networks that perform many tasks tend to self-organize into modules (1-3). This sugge...
biorxiv.org
And relatedly, Felix wrote a good piece on the stress and anxiety currently affecting many people who work in AI due to the current climate in the industry: docs.google.com/document/d/1... If only more folks in AI were gentle and introspective like this...
AI and Stress
200Bn Weights of Responsibility The Stress of Working in Modern AI Felix Hill, Oct 2024 The field of AI has changed irrevocably in the last 2 years. ChatGPT is approaching 200m monthly users. Gemin...
docs.google.com