Jessica Hullman

@jessicahullman.bsky.social

Ginni Rometty Prof @NorthwesternCS | Fellow @NU_IPR | AI, people, uncertainty, beliefs, decisions, metascience | Blog @statmodeling

About a year ago, I wrote skeptically about LLMs in peer review -- not because of skepticism about their inherent capabilities, but because I don't want the research community to optimize for the taste of any one person/system. What's changed since then?

What’s next for machine learning peer review?

A bit over a year ago, I wrote about the dangers of using LLMs for peer review. The most serious concern I had was algorithmic monoculture: the research community would collectively end up optimizing ...

bryanwilder.substack.com

Multiverse Analyses “[Only 6/152 (3.9%)] studies discussed whether their competing specifications were defensible or principled (distinguishing between equivalent, non-equivalent, or uncertain specifications) in the sense of Del Giudice & Gangestad (2021) [screenshot below]” doi.org/10.1177/2515...

Bild
Alejandro Sandoval-Lentisco@asandovall.bsky.social · 2w ago

How often are multiverse analyses implemented, not just cited/discussed? How are they implemented? 🧵 New preprint on the uptake and implementation of multiverse-style analyses 👇 www.biorxiv.org/content/10.6...

Newly proposed rules from the federal government will irreparably damage US science: please express your dissent and post a comment before the OMB public comment period closes this Monday, July 13. Read about the endgame here: www.theguardian.com/commentisfre...

Is the US trying to make scientists’ work so difficult that they simply give up? | Daniel Malinsky

New Trump administration rules would undermine longstanding research practices. It’s death by a thousand cuts

theguardian.com

It took a while but the effects of unstable government funding on research are showing up. Smaller PhD cohorts. Faculty spending more time in industry or scaling down their plans. Faculty leaving to other countries. Another unforced disaster.

I recall my dad saying to me as a kid "Isn't it terrible that we have to sell our time to survive." He was rarely serious or philosophical so it stuck with me. Also his drug talk when I was 12: "Marijuana, LSD, won't hurt you. Stay away from coke & heroin." I rarely listened but took that to heart

noam@noamchompers.bsky.social · last mo.

Do you guys also have ‘thing that was said to me when I was young that echos in my mind forever’? Can’t tell if that’s a fun prompt or a very dark one lol

Sharing our latest endeavour here. How reproducible is your paper? @philipjakobbln.bsky.social and I built rigor.me to ease the burden of computational reproducibility. If you provide a paper, data and code, we execute it and tell you what works (and what fails). Beta is available now: rigor.me

Rigor

rigor.me

RigorLabs@rigorlabs.bsky.social · last mo.

AUTOMATING COMPUTATIONAL REPRODUCIBILITY My colleague @philipjakobbln.bsky.social and I are currently engaged in a research project where we reproduce scientific results en masse. To that end, we built rigor.me, a platform for automatically reproducing papers using agents 🧵

Lol, a paper that describes my life. My natural reaction to feeling like I'm finally getting the hang of something has always been to drop it and try something else. And somehow I still often feel like I'm not getting out fast enough.

Bild

"The more phil of science I read, and the more familiar I became with different pockets of the metascience community, the more I came to realize how little consensus there is about how to evaluate science. But I don’t blame those who default to thinking the metascientists have figured things out.

Samuel Moore@samuelmoore.org · last mo.

"Whether a paper that fails a robustness check or skips preregistration is worse science depends on assumptions about what those signals are proxies for — assumptions reformers have been debating for years without resolution." jessicahullman.substack.com/p/ai-review-...

Check out our #FAccT2026 tutorial tomorrow on Bridging Predictions and Interventions in Social Systems! We're going to be building a predictions - interventions index to track ADS systems, evaluations, and gaps therein. Bring a device; I'll try to bring worksheets too :)

Friday, June 26, 3:30p-4:30p Musset Level A

Using AI to support peer review seems unavoidable, but what quality checks should AI implement? We can take some lessons from metascience on the hard reward design problem that is AI review. I wrote a paper synthesizing a few points the emerging lit seems at risk of confusing. 1/

Bild

Excellent essay by Leif, Tyler, and Ben. Interestingly, reification of “internal states” is one of the pitfalls of the intentional stance of Dennett, something we are seeing a lot of AI interoperability research succumbing to.

Ben Recht@beenwrekt.bsky.social · 2mo ago

In _The Ideas Letter_, Leif Weatherby, Tyler Shoemaker, and I wrote about the creation of reality through pseudoscience, religion, and psychosis. Of course, the piece is mostly about AI.

After two days with Claude Fable 5 the best way I can describe it is "relentlessly proactive" - here's an example where I dropped in a screenshot of a bug and it span up custom CORS Python servers and used pyobjc-framework-Quartz to capture screenshots simonwillison.net/2026/Jun/11/...

Claude Fable is relentlessly proactive

After two days of experience with Claude Fable 5 I think the best way to describe it is relentlessly proactive. It knows a whole lot of tricks and it will …

simonwillison.net

It's interesting how quick we are to assume science is aligned by default. There's an idea that good science doesn't require human involvement, it's about the "Truth" which can be discovered & verified independent of our judgment. In doing so we ignore most of the history of science.

Alexander Kustov@akoustov.bsky.social · 2mo ago

At the same time, if you are a scientist and you want to solve a problem or simply get things done, you are going to be perfectly content if the solution is machine-generated, whether it is about analysis, the write-up or both.

Lots of people with strong opinions about use of AI detectors to filter what papers we review (or what stories we consider for awards). I wrote up a toy model of using AI detection to infer author type to get at some implict assumptions we make when we argue that AI detection is or is not useful

AAndrew Gelman et al.@statmodeling.bsky.social · 2mo ago

When is detecting AI-generated text worthwhile? statmodeling.stat.columbia.edu/2026/06/06/w...