Cory McCartan

@corymccartan.com

Asst. Prof. of Statistics & Political Science at Penn State. I study stats methods, gerrymandering, & elections. Founder of UGSDW and proud alum of HGSU-UAW L. 5118. 🏳️‍🌈 corymccartan.com

Don't posit a theory that is only interesting because it implies X causes variation in Y, test for a statistical relationship between X and Y, but then say you are not interpreting causally. Of course you are interpreting it causally. That was the whole point of what you just did.

Mechanically, what this means is that it’s assuming fairly smooth interactions between covariates. So if you want flexible models of interaction you have to build them in yourself. Which is, after all, what you’d expect with BART.

This is definitely interesting. It’s probably worth noting in high dimensions that mixed sobolev spaces are much smoother than usual sobolev spaces. So think of this kernel as closer to a squared exponential than a Matern-1. But still rough enough that the paths aren’t analytic.

Noah Greifer@noahgreifer.bsky.social · 2d ago

Fascinating and clear paper by @corymccartan.com and @melodyyhuang.bsky.social, greatly enhancing our understanding of how Bayesian Additive Regression Trees (BART) works and why it is so effective. A must-read for my fellow BART enthusiasts. #statssky #causalinference

New WP w/@melodyyhuang.bsky.social studying the success of BART models, which regularly win causal inference competitions! We argue that BART should be thought of as a random features approximation to a limiting GP. This view helps understand BART & apply it in more places arxiv.org/abs/2607.28844

Seeing the Forest for the Trees: The Gaussian Process Limit of BART

Cory McCartan & Melody Huang

Abstract:
Bayesian Additive Regression Trees (BART) have shown state-of-the-art performance in both prediction and causal inference problems. Previous theoretical work has attempted to explain BART’s superior performance by establishing posterior contraction rates for standard BART models, but these rates depend strongly on the number of covariates. Here, we take a different approach and study the behavior of BART as the number of trees grows towards infinity. We show that in this regime, BART converges to a Gaussian process (GP) with a particular kernel. The kernel and its corresponding reproducing kernel Hilbert space (RKHS) have favorable inferential properties that help explain BART’s excellent performance. We introduce random tree features as an approximation to this limiting GP, and establish minimax-optimal learning rates for ridge
regression on these random features that depend only logarithmically on dimension. In addition to providing insight into the empirical success of BART, random tree features offer a computational benefit over traditional MCMC estimation. The random-features approximation also allows
practitioners to easily incorporate BART into any model which has a linear predictor, expanding the applicability and flexibility of BART.

Out today in the APSR: our paper studying which gerrymandering reforms work! Long story short, do what Michigan does! Slightly longer story: we project these different complex reforms onto a 1-dimensional "partisan leeway" axis, and do continuous-trt DiDiD there. doi.org/10.1017/S000...

First page of paper

If there is a practical difference then people should just talk in simple terms about what is to be done. The dictionary definition stuff is largely exhausting, pedantic navel gazing

Thrilled to share that I received POLMETH's Statistical Software Award for 2026 for the #rstats package geomander. It's a performant toolkit for working with spatial data, especially election and demographic data. Learn more about the package at christophertkenny.com/geomander/

Geographic Tools for Studying Gerrymandering

A compilation of tools to complete common tasks for studying gerrymandering. This focuses on the geographic tool side of common problems, such as linking different levels of spatial units or estimatin...

christophertkenny.com

Peacetime politics involves less physical danger than war, but it has some dangers in common: both require what Clausewitz calls “moral courage” by which he means the willingness to accept risks and moral responsibility. This is what I saw lacking on that panel. 1/🧵

Liberal Currents@liberalcurrents.com · 2w ago

"I had only one reaction to the panelists and their cheerleaders: these gutless losers will never win. If they have their way, we will, at best, remain in this crisis for a generation longer, and at worst we will simply lose the war." www.liberalcurrents.com/tit-for-tat-...

Here's something I post from time to time. My answer to a reader who asked me: what could journalists do NOW to break with some of their more corrosive habits.

Bild

Interesting read. I think both can be true: (a) the public reads more than ever, and (b) standardized tests, phones (TM), & pandemic have hindered college students' ability to engage with longer works. Academics see (b) and are receptive to hearing that reading has cratered everywhere

Post nicht verfügbar.

I have just resigned from the board of "Statistics and Computing", along with 17 other Associate Editors. This was a difficult decision: this great journal has published many tremendous articles under the leadership of EiC Ajay Jasra, and before him David Hand, Gilles Celeux, and Mark Girolami.

If there’s an issue Democrats could run on and win in a landslide, it’s anti-corruption. But first the Dem establishment needs to deal with its own corruption problem. It’s not enough to be right about Trump’s corruption—they also have to be credible.

Bar chart from Reuters/Ipsos poll of 1,019 U.S. adults, Sept. 19–21, 2025, asking which party has a better plan on 11 issues. Republicans lead on crime (40% vs. 20%), immigration (40% vs. 22%), foreign conflicts (35% vs. 23%), U.S. economy (34% vs. 24%), gun control (32% vs. 28%), political extremism (30% vs. 26%), and corruption (27% vs. 21%). Democrats lead on respect for democracy (31% vs. 29%), healthcare (34% vs. 25%), women’s rights (38% vs. 25%), and environment (37% vs. 23%). Margin of error ±3 points; third-party and non-responses excluded.

New R package on CRAN! ONNX is a runtime & file format for ML models. 'onnxr' lets you load & run models in about 2 lines of code! E.g. image detection running in ms from a pretrained model. Perfect for embeddings, smaller models, etc. And lots of .onnx available online! corymccartan.com/onnxr/

R code snippet:
library(onnxr)
model = onnx_model(model_path)
res = onnx_run(model, img, simplify = TRUE)Vermeer's "The Milkmaid" with person, bowls, dining table detected by ML model running via onnxr

This take may ruffle some feathers. Not only do political parties expel members all the time, but the ability to do so is absolutely crucial for maintaining political cohesion and protecting against demagogic capture. Like would anyone on this platform object if the Democrats expelled Fetterman?

Post nicht verfügbar.

So I urge you to look at the folks who will inevitable call for us to privatize social security in the wake of this report and check for the overlap with the folks who brought us the surging inequality and a sluggish economic recovery that got us here. 4/4

Another clear centering of Trump’s emotions from a few days ago: the number of headlines and articles that referred to E. Jean Carroll that referred to her as an “enemy” or “adversary” rather than “victim.”

Larry Glickman@larryglickman.bsky.social · 2mo ago

Yet another article centering Trump’s emotions. We learn here that he is ”angry,” “irate” and that he issued a “scathing rebuke” of the decision. www.nytimes.com/2026/05/30/a...

Newly updated simple congressional model (tinyurl.com/cmchousemodel), now with LA map & recalibrated election model fit only to 2026-2024 data! Ds have lost 4 seats on avg to redistricting, but the loss grows to 5.6 seats in a D+6 environment. Ds 98% to win house in a D+6 environment, though

BildBildBild

A small 🧵 about frustrations and opportunities w/ open-source research software for the social sciences: Came across a new paper in a top methods journal, explaining/demo'ing a new R package. Sounded v. cool, so was excited to dig in. What I found was largely a mess...