Alicia Curth

@aliciacurth.bsky.social

Machine Learner by day, 🦮 Statistician at ❤️ In search of statistical intuition for modern ML & simple explanations for complex things👀 Interested in the mysteries of modern ML, causality & all of stats. Opinions my own. https://aliciacurth.github.io

Oh friends who are complaining about not enough Real Math^tm in their feed, I am here to help. Well, Alicia is here to help, at least!

Alicia Curth@aliciacurth.bsky.social · 2y ago

Part 2: Why do boosted trees outperform deep learning on tabular data?? @alanjeffares.bsky.social & I suspected that answers to this are obfuscated by the 2 being considered very different algs🤔 Instead we show they are more similar than you’d think — making their diffs smaller but predictive!🧵1/n

To emphasise just how accurately that reflects Alan’s approach to research (which I 100% subscribe to btw), I feel compelled to share that this is the actual slide I use whenever I present the U-turn paper in Alan’s absence 😂 (not a joke)

Bild
Alan Jeffares@alanjeffares.bsky.social · 2y ago

If people knew how much of my PhD has consisted of reading about something new, referencing back to Elements of Statistical Learning, and simply writing down what I learned… It feels like a cheat code!

btw this is why friends dont let friends skip the “boring classical ML” chapters in Elements of Statistical Learning‼️ (True story: the origin of this case study is that @alanjeffares.bsky.social[big EoSL nerd] looked at the neural net eq&said “kinda looks like GBTs in EoSL Ch10”&we went from there)

Alicia Curth@aliciacurth.bsky.social · 2y ago

but WAIT A MINUTE — isn’t that literally the same formula as the kernel representation of the telescoping model of a trained neural network I showed you before?? Just with a different kernel?? Surely this diff in kernel must account for at least some of the observed performance differences… 🤔7/n

Part 2: Why do boosted trees outperform deep learning on tabular data?? @alanjeffares.bsky.social & I suspected that answers to this are obfuscated by the 2 being considered very different algs🤔 Instead we show they are more similar than you’d think — making their diffs smaller but predictive!🧵1/n

Bild

aren’t smoothers just THE BEST?? understanding double descent, random forests, neural network complexity, and now causal inference — smoothers are just at your service when you need them

Michael Knaus@mcknaus.bsky.social · 2y ago

How concretely? Smoothers, we need smoothers!!! Outcome nuisance parameters have to be estimated using methods like (post-selection) OLS, (kernel) (ridge) or series regressions, tree-based methods, … Check out @aliciacurth.bsky.social for nice references and cool insights using smoothers in ML.

New WP 🚨 1. Recipe to write estimators as weighted outcomes 2. Double ML and causal forests as weighting estimators 3. Plug&play classic covariate balancing checks 4. Explains why Causal ML fails to find an effect of 1 with noiseless outcome Y = 1 + D 5. More fun facts arxiv.org/abs/2411.11559

BildBildBildBild

From double descent to grokking, deep learning sometimes works in unpredictable ways.. or does it? For NeurIPS(my final PhD paper!), @alanjeffares.bsky.social & I explored if&how smart linearisation can help us better understand&predict numerous odd deep learning phenomena — and learned a lot..🧵1/n

Bild

I finally decided to double up with an account here hoping to find more scientific discourse :) So: Hi, I’m Alicia, Machine Learning Researcher at MSR (since last month)! Prev I was a PhD student in Cambridge trying to make sense of the mysteries of modern Machine Learning (— to be continued!!) :)