"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/
Aaron Roth
@aaroth.bsky.social
Professor at Penn, Amazon Scholar at AWS. Interested in machine learning, uncertainty quantification, game theory, privacy, fairness, and most of the intersections therein
Right now we have "problem overhang" - lots of problems we as a community are interested in because smart and charismatic people thought about them and convinced us that these problems are important. So we are happy/interested to see them solved by AI.
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
It was fun giving this tutorial. If you missed it, check out calibration-tutorial.github.io where all of our materials are available, including slides, hundreds of pages of lecture notes, an interactive demo, and an annotated reading list.
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
I've been referring people to this perspective every day for the last couple weeks. As we work together to determine new norms, we should not tolerate pure AI or otherwise bad writing. Tell your friends if they fall into this trap.
AI is getting good at math. What are our jobs as researchers now that we have proof machines? The raw proofs that come from LLMs are difficult to understand, even if correct. So its now easy to quickly write many badly written papers that nevertheless contain correct proofs of interesting theorems.
Reminder, this is in a few hours!! j o i n u s
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Looking forward to being at #ICML2026 in Seoul next week! Unfortunately I'll be there for only 60 hours, but 2.5 of those will be at our machine unlearning tutorial (co-presented with Vinith Suriyakumar)! Monday July 6 at 9 AM -- don't miss out!
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Calibration, Decisions, and Collaboration in Learning | ICML 2026
An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration.
calibration-tutorial.github.io
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.
Calibration, Decisions, and Collaboration in Learning | ICML 2026
An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration.
calibration-tutorial.github.io
AI is getting good at math. What are our jobs as researchers now that we have proof machines? The raw proofs that come from LLMs are difficult to understand, even if correct. So its now easy to quickly write many badly written papers that nevertheless contain correct proofs of interesting theorems.
Interested to see how this goes. Now that the cost of generating a paper-like-object has dropped so low, publication venues are going to start having to impose costs on submissions of various sorts. We'll need experiments to figure out the best way to do this without disrupting science.
TMLR has been facing an significant uptick in the number of submissions since the start of 2026. This is placing an extreme burden on our amazing team of reviewers & action editors. To ease this burden, TMLR will be implementing submission quotas, effective July 1. 1/n medium.com/@TmlrOrg/ann...
For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557
For a long time we didn't know if test-time randomization was needed for sample-optimal multicalibration and omniprediction. It's not. arxiv.org/abs/2606.20557
Modern LLMs are incredibly good compression algorithms, which can shed light on why autonomous data science agents don't overfit as much as you might think. arxiv.org/abs/2606.11045
In the last 48h: - Jr researcher asked me wheter to use AI in making talks - Saw two talks, with AI {slop, enhanced} slides Collected my thoughts and wrote a post. Tl;dr: don't steal your own thinking, don't remove *you* from your talks. Also, give a &#@% about your talks.
A clearly hallucinated citation! NeurIPS 2026 decisions aren't out yet. But wait --- the hallucination is also present in the bibtex entries from openreview openreview.net/forum?id=fAj... and Google Scholar scholar.googleusercontent.com/scholar.bib?...
Recently we showed that the minimax optimal rate for multicalibration is T^{2/3}. But that doesn't mean you have to do that badly on all instances. We give an algorithm that can adapt to easy instances and get better rates while still being minimax optimal in the worst case. arxiv.org/abs/2605.09273
I'm giving this talk at the MIT CS theory seminar tomorrow. Stop by if you are around!
I've recently been getting invitations to talk about how to use AI tools to assist with TCS research. Its something I've been doing a lot, but don't have structured thoughts about how to explain process. But I'm going to try -- first such talk is tomorrow: t.co/wlHPBzXzDm
We updated our paper --- and solved the open problem highlighted in the old version. Now our lower bound construction has only polylog(1/eps) many groups instead of poly(1/eps) many groups. The construction is also simplified.
Excited about a new paper! Multicalibration turns out to be strictly harder than marginal calibration. We prove tight Omega(T^{2/3}) lower bounds for online multicalibration, separating it from online marginal calibration for which better rates were recently discovered.
Why do all LLMs predict 27 as their favorite number? There may be a principled explanation. Learn more at Agents in the Wild at #ICLR2026. @ericeaton.bsky.social, me, @surbhigoel.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, @optimistsinc.bsky.social
How many samples do you need from an unknown distribution in order to train a model with multicalibration error at most epsilon? Answer: 1/epsilon^3 samples is both necessary and sufficient.
April is #AIMonthAtPenn! On 4/24, Warren Center faculty affiliate Aaron Roth will give the George H. Heilmeier Faculty Award Lecture in Amy Gutmann Hall. More information and registration here: ai.upenn.edu/heilmei...
Say hi to @marcelhussing.bsky.social at ICLR
In “Replicable Reinforcement Learning with Linear Function Approximation,” @optimistsinc.bsky.social, @marcelhussing.bsky.social, @mkearnsphilly.bsky.social, @aaroth.bsky.social, @sikatasengupta.bsky.social, & more develop replicable methods for linear function approximation in RL: (5/12)
I've recently been getting invitations to talk about how to use AI tools to assist with TCS research. Its something I've been doing a lot, but don't have structured thoughts about how to explain process. But I'm going to try -- first such talk is tomorrow: t.co/wlHPBzXzDm
https://www.cics.umass.edu/events/research-ai-era-seminar-aaron-roth
t.co
AI Agents like Codex are very good at figuring out taxes, including obscure local ones that Intuit doesn't bother with (looking at you, Philadelphia local taxes). Businesses that provide financial/legal services that involve reasoning through dense but public documentation are in trouble.
Announcing the #ICML2026 tutorials! All ten tutorials will be presented the first day of the conference, Monday July 6. Read the blog post for more details on the selection process! blog.icml.cc/2026/04/02/a...
So many interesting things here. (N.b. I get to think ab interfaces for this this all day long at work :)) One thing I find interesting here is how similar the real work of science is to that of the humanities, both of which are centered around human judgement ab what is relevant and interesting.
Michael @mkearnsphilly.bsky.social ) and I wrote a blog post about our experiences using AI for research, and our thoughts on what these developments will mean for research, publication, and education: www.amazon.science/blog/how-ai-...
Very cool work. Empirical science has many researcher-degrees-of-freedom which makes it hard to interpret specific studies --- these are only a single trajectory through the data analysis multiverse. Human researchers are opaque. But with agents you can explore the whole space!
There's growing evidence that LLMs can p-hack. But p-hacking also points to something bigger: a data science multiverse of defensible analytical choices. We wrote a paper (arxiv.org/abs/2602.18710) on using LLM agents to map this multiverse systematically. 🧵