Sean Trott

@seantrott.bsky.social

Most of the statistical tests you learned in stats class — t-test, ANOVA, correlation — are actually special cases of a single thing: linear regression. Ch 7 of Experimentology argues that thinking in models, not tests, is more flexible and a better foundation for theory. 🧵 experimentology.io

Bild

In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵

Bild

If you're a faculty, research scientist, postdoc, or senior PhD in any area of science, select your top 6 virtues at this link: shorturl.at/bV2lF You can also tell us if you think we missed any. We want broad participation, so pls RT! 🙏 In collab w/@devezer.bsky.social @statmodeling.bsky.social

Epistemic Virtues for Scientific Inquiry — Community Survey

In Six Memos for the Next Millenium, Italo Calvino proposed six literary virtues he felt should be preserved and cultivated in literature regardless of how the world changed: lightness, quickness, exa...

shorturl.at

Does vision training change how language is represented and used in meaningful ways?🤔The answer is a nuanced yes! Comparing VLM-LM minimal pairs, we find that while the taxonomic organization of the lexicon is similar, VLMs are better at _deploying_ this knowledge. [1/9]

Bild

I think understanding which factors lead to convergence and divergence (both in behavior and internal mechanisms) across networks is crucial to understanding what kinds of systems we're studying and what kinds of claims we can generalize across model instances. Very cool work!

Ann Huang@annhuang42.bsky.social · 8mo ago

📍Excited to share that our paper was selected as a Spotlight at #NeurIPS2025! arxiv.org/pdf/2410.03972 It started from a question I kept running into: When do RNNs trained on the same task converge/diverge in their solutions? 🧵⬇️

A confounding thing for the linguistics of LMs: the best way to assess their grammatical ability is string probability. Yet string probability and grammaticality are famously not the same! Really excited to have this out, where we give a formal account, w/ experiments, of how to make sense of that!

Jennifer Hu@jennhu.bsky.social · 9mo ago

New work to appear @ TACL! Language models (LMs) are remarkably good at generating novel well-formed sentences, leading to claims that they have mastered grammar. Yet they often assign higher probability to ungrammatical strings than to grammatical strings. How can both things be true? 🧵👇

Screenshot of a figure with two panels, labeled (a) and (b). The caption reads: "Figure 1: (a) Illustration of messages (left) and strings (right) in toy domain. Blue = grammatical strings. Red = ungrammatical strings. (b) Surprisal (negative log probability) assigned to toy strings by GPT-2."