Shuvom Sadhuka

@shuvoms.bsky.social

phd-ing at mit csail, https://shuvom-s.github.io/, hertz fellow

1/ CS majors are drilled to think about "worst-case" performance of algorithms. By contrast, much of the discourse on AI evals focuses on average-case or best-case (e.g. LLM X can solve IMO problems). Maybe one key to "reliability" is certifying the 1st quantile of outputs too, not just the mean.

Bild

How can you evaluate agent trajectories with only black-box access to a verifier and the agent? Introducing E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing ❌no finetuning ❌no additional GPU compute ⬛black-box access ✔️controllable error rates w/ guarantees

BildBild

What can you do when you need to evaluate a set of models but don't have too much labeled data? Solution: add unlabeled data. It was great to co-lead this project with @dmshanmugam.bsky.social and come talk to us at NeurIPS!

Divya Shanmugam@dmshanmugam.bsky.social · 10mo ago

New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.

Slightly interesting observation: I asked chatgpt "generate an image that represents what you know about me. don't ask questions" and it drew me as a woman. I've never revealed my gender to chatgpt, so I asked it why it drew me as a woman. Here's what it said:

Bild

I'll be at the ICLR DMLR and PMLDS workshops presenting ongoing work with @dmshanmugam (+ a wonderful team including @manish_raghavan, @mit_caml, @lab_berger, @2plus2make5)! Details in thread below 🧵 [1/4]

Bild

I'll be organizing a new algorithmic fairness reading group at MIT this semester. Website (w/ signup link) is here: We have funding for food, and all are welcome/encouraged to come :).