Codex and Claude Code have a neat auto approve feature where 1) it doesn't ask for permissions, but 2) you get to feel safe. Well --- do you? The premise is it is asking an agent for permission. But if you don't trust the driver agent, why should you trust the review agent?
Aaron Roth
@aaroth.bsky.social
Professor at Penn, Amazon Scholar at AWS. Interested in machine learning, uncertainty quantification, game theory, privacy, fairness, and most of the intersections therein
A world without open problems Here are some that fell today: K-server: arxiv.org/abs/2609.15979 Matroid Secretary: arxiv.org/abs/2609.145... Matrix Spencer: arxiv.org/abs/2609.15025 (Well Matrix Spencer was maybe also a few weeks ago, but who's counting? arxiv.org/abs/2608.28816 )
The $k$-server conjecture is true
The $k$-server conjecture states that a deterministic online algorithm can achieve competitive ratio $k$ on every metric space. We prove the conjecture. Specifically, we show that the work function al...
arxiv.org
So we are clearly moving to a world that will be without open problems, but this doesn't mean math is going away. Interesting work will correspond to discovering new questions and coherent theoretical frameworks. Well defined agreed-to-be-interesting problems won't last long.
I sometimes see people trying to use complexity theory to argue that building "true" AI is impossible. I find this unreasonably annoying. It requires ignoring what is in front of your face and it ignores that worst-case complexity has been an awful guide in machine learning.
Years of iterating against the same benchmarks should, by textbook logic, produce overfitting. It largely doesn't. New research explains why: strategies that generalize can be expressed in too compact a form to allow memorization, while the ones that overfit don't survive a compression.
Why don’t machine learning research agents overfit?
New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.
amazon.science
A blog post on some neat work with @zstevenwu.bsky.social and Martin Bertran: www.amazon.science/blog/why-don...
Why don’t machine learning research agents overfit?
New research indicates that AI agents learn compressible models of data, which don’t have enough space to enable memorization.
amazon.science
Well, this Navier-Stokes affair did blow up in finite time
A new semester, and the first lecture is in the books for my class on the "Mathematical Foundations of AI Alignment". aaroth.github.io/cis-7000-ai-... What does that mean? Good question. We have about a semester in which to figure it out.
CIS 7000 — Mathematical Foundations of AI Alignment
Course topics and reading list.
aaroth.github.io
Recordings from FORC 2026 are now available! Please check them out. Also, subscribe to FORC's new YouTube channel while you're at it! www.youtube.com/playlist?lis...
FORC 2026 - YouTube
Talk recordings from FORC 2026
youtube.com
We have a new online boosting algorithm which is very efficient and effective. Unlike prior algorithms which maintain many weak learners and ensemble them, we operationalize the "dual view" of boosting. We don't maintain an ensemble. We try to construct an online hard core distribution.
I'm pleased to share our #ICML2026 tutorial on machine unlearning! Presented by Vinith Suriyakumar and myself, it includes a full set of videos recorded and posted to YouTube! Please check it out: unlearning-tutorial.github.io
People used to be able to impress and intimidate reviewers with complicated proofs. This will change. In the age of AI inscrutable proofs are cheap. It is understandable proofs that are valuable. Opaque complexity is now it is a sign of laziness or lack of insight.
"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/
Right now we have "problem overhang" - lots of problems we as a community are interested in because smart and charismatic people thought about them and convinced us that these problems are important. So we are happy/interested to see them solved by AI.
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
I've been referring people to this perspective every day for the last couple weeks. As we work together to determine new norms, we should not tolerate pure AI or otherwise bad writing. Tell your friends if they fall into this trap.
AI is getting good at math. What are our jobs as researchers now that we have proof machines? The raw proofs that come from LLMs are difficult to understand, even if correct. So its now easy to quickly write many badly written papers that nevertheless contain correct proofs of interesting theorems.
Reminder, this is in a few hours!! j o i n u s
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Looking forward to being at #ICML2026 in Seoul next week! Unfortunately I'll be there for only 60 hours, but 2.5 of those will be at our machine unlearning tutorial (co-presented with Vinith Suriyakumar)! Monday July 6 at 9 AM -- don't miss out!
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Calibration, Decisions, and Collaboration in Learning | ICML 2026
An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration.
calibration-tutorial.github.io
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Calibration, Decisions, and Collaboration in Learning | ICML 2026
An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration.
calibration-tutorial.github.io
AI is getting good at math. What are our jobs as researchers now that we have proof machines? The raw proofs that come from LLMs are difficult to understand, even if correct. So its now easy to quickly write many badly written papers that nevertheless contain correct proofs of interesting theorems.
Interested to see how this goes. Now that the cost of generating a paper-like-object has dropped so low, publication venues are going to start having to impose costs on submissions of various sorts. We'll need experiments to figure out the best way to do this without disrupting science.
TMLR has been facing an significant uptick in the number of submissions since the start of 2026. This is placing an extreme burden on our amazing team of reviewers & action editors. To ease this burden, TMLR will be implementing submission quotas, effective July 1. 1/n medium.com/@TmlrOrg/ann...
For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557
For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557
Modern LLMs are incredibly good compression algorithms, which can shed light on why autonomous data science agents don't overfit as much as you might think. arxiv.org/abs/2606.11045
In the last 48h: - Jr researcher asked me wheter to use AI in making talks - Saw two talks, with AI {slop, enhanced} slides Collected my thoughts and wrote a post. Tl;dr: don't steal your own thinking, don't remove *you* from your talks. Also, give a &#@% about your talks.
A clearly hallucinated citation! NeurIPS 2026 decisions aren't out yet. But wait --- the hallucination is also present in the bibtex entries from openreview openreview.net/forum?id=fAj... and Google Scholar scholar.googleusercontent.com/scholar.bib?...
Recently we showed that the minimax optimal rate for multicalibration is T^{2/3}. But that doesn't mean you have to do that badly on all instances. We give an algorithm that can adapt to easy instances and get better rates while still being minimax optimal in the worst case. arxiv.org/abs/2605.09273
I'm giving this talk at the MIT CS theory seminar tomorrow. Stop by if you are around!
I've recently been getting invitations to talk about how to use AI tools to assist with TCS research. Its something I've been doing a lot, but don't have structured thoughts about how to explain process. But I'm going to try -- first such talk is tomorrow: t.co/wlHPBzXzDm