1/ CS majors are drilled to think about "worst-case" performance of algorithms. By contrast, much of the discourse on AI evals focuses on average-case or best-case (e.g. LLM X can solve IMO problems). Maybe one key to "reliability" is certifying the 1st quantile of outputs too, not just the mean.
Shuvom Sadhuka
@shuvoms.bsky.social
phd-ing at mit csail, https://shuvom-s.github.io/, hertz fellow
How can you evaluate agent trajectories with only black-box access to a verifier and the agent? Introducing E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing ❌no finetuning ❌no additional GPU compute ⬛black-box access ✔️controllable error rates w/ guarantees
I’m excited to share our new paper A Bayesian Model for Multi-stage Censoring, which I will present at #ML4H2025 in San Diego! 🧵 below:
What can you do when you need to evaluate a set of models but don't have too much labeled data? Solution: add unlabeled data. It was great to co-lead this project with @dmshanmugam.bsky.social and come talk to us at NeurIPS!
New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.
Slightly interesting observation: I asked chatgpt "generate an image that represents what you know about me. don't ask questions" and it drew me as a woman. I've never revealed my gender to chatgpt, so I asked it why it drew me as a woman. Here's what it said:
I wrote up some thoughts on what it means to measure the entropy of natural languages and connections to LLMs, loosely inspired by an awesome paper from Shannon in 1951(!) Check it out: shuvom-s.github.io/blog/2025/me... [1/3]
Measuring Entropy | Shuvom Sadhuka
How would you measure the entropy of natural language?
shuvom-s.github.io
I'm in Vancouver and will be giving a spotlight talk at #ML4H tomorrow, Dec. 15, at 4:30pm on some ongoing work on modeling multi-stage selection problems in clinical settings. Work done with (high school senior!) Sophia Lin, Bonnie Berger, and @emmapierson.bsky.social. I hope to see you there!
I wrote a "blog post" a year ago on advice for applying to fellowships but never ended up posting it anywhere. I thought I might as well share it, especially since fellowship applications are around the corner! Feel free to reply w/ edits, questions, etc. https://shorturl.at/ZOGVH
The @MITEECS Graduate Application Assistance Program is open for applications! Our team of PhD students offers 1-on-1 mentorship and office hours to EE/CS graduate school applicants (to any school, not just MIT) from around the world. Deadline: 10/15/24. https://eecs-gaap.mit.edu/
I'll be at the ICLR DMLR and PMLDS workshops presenting ongoing work with @dmshanmugam (+ a wonderful team including @manish_raghavan, @mit_caml, @lab_berger, @2plus2make5)! Details in thread below 🧵 [1/4]
If you’re considering applying to grad school, especially in CS, please consider applying to the @MITEECS graduate application assistance program (
I'll be organizing a new algorithmic fairness reading group at MIT this semester. Website (w/ signup link) is here: We have funding for food, and all are welcome/encouraged to come :).
The new @MIT_SCC had an essay competition on the future of computing, and I wrote a (non-technical) essay on genomic privacy. Here's a rough summary: [1/n]