shira mitchell
@shiraamitchell.bsky.social
survey statistician at blue rose research 🏕
blog post: more on SynthMargins and Bayes-Raking We want to poststratify (MRP) but only have partial information (margins) about the poststratification variables in the population. More ideas from Bob Carpenter (in #mcmc_stan), @yajuansi.bsky.social, and @shirokuriwaki.bsky.social
Survey Statistics: more on SynthMargins and Bayes-Raking statmodeling.stat.columbia.edu/2026/08/25/s...
blog post: Modeling Complex Contingency Tables We want to poststratify (MRP) but only have partial information about the poststratification variables in the population. Max Goplerud, @shirokuriwaki.bsky.social, Jens Wiederspohn, @adamcs.bsky.social, and Philip Greengard have ideas !
Survey Statistics: Modeling Complex Contingency Tables statmodeling.stat.columbia.edu/2026/08/18/s...
blog post: wanting workflow We review a covid survey case study from the new Bayesian Workflow book. We model: - measurement: test sensitivity and specificity - representation: differences between sample and population (MRP) How can workflow help us here ?
Survey Statistics: wanting workflow statmodeling.stat.columbia.edu/2026/08/11/s...
blog post: structured MRP to smooth survey weights Adjusting for lots of variables can lead to very large weights. So @yajuansi.bsky.social, @trangucc.bsky.social, Jonah Sol Gabry, and Andrew Gelman turned to a structured MRP and its equivalent weights.
Survey Statistics: structured MRP to smooth survey weights statmodeling.stat.columbia.edu/2026/08/04/s...
blog post: equivalent models, equivalent weights (locally) with survey-weighting methods, we can ask: under which outcome models do they do well ? with outcome-model methods (like MRP), we can ask: what are the (locally) equivalent weights ?
Survey Statistics: equivalent models, equivalent weights (locally) statmodeling.stat.columbia.edu/2026/07/28/s...
blog post: poststratification without population level information Poststratification uses population data on X to help estimate a population mean E(Y). But sometimes population data on X isn't available: In 2016 Andrew asked pollsters to poststratify on party ID, but how ?
Survey Statistics: poststratification without population level information statmodeling.stat.columbia.edu/2026/07/21/s...
blog post: quantifying uncertainty in ranked choice voting polls RCV uses rankings to get a winner by instant runoff. Polls estimate rank probabilities with uncertainty. Unlike with non-RCV, even in random samples a plurality of uncertainty mass can get an incorrect winner.
Survey Statistics: quantifying uncertainty in ranked choice voting polls statmodeling.stat.columbia.edu/2026/07/14/s...
blog post: toy example for energy balancing weights How do energy balancing weights (used now by the NYT/Siena Poll) handle unsampled population groups ? Let's work thru a toy example and compare to Poststratification, Raking, and MRP.
Survey Statistics: toy example for energy balancing weights statmodeling.stat.columbia.edu/2026/07/07/s...
blog post: Big Changes in the Times/Siena Poll 2 changes to their survey weights: 1. new weighting variable: support score 2. new weighting method: energy balancing
Survey Statistics: Big Changes in the Times/Siena Poll statmodeling.stat.columbia.edu/2026/06/30/s...
blog post: perfect collinearity in the sample but not in the population Two variables are perfectly collinear in your sample, so you drop one. You use your model to predict in the population. What can go wrong ? Let's talk thru a Census Bureau toy example from BDA2.
Survey Statistics: perfect collinearity in the sample but not in the population statmodeling.stat.columbia.edu/2026/06/23/s...
blog post: using MRP in later analyses (pride edition) happy pride ! 🌈 @jeffreylax.bsky.social & Phillips 2009 used MRP to estimate state-level public opinion about policies affecting gays and lesbians. They then use this as a predictor of whether the state adopts the policies.
Survey Statistics: using MRP in later analyses (pride edition) statmodeling.stat.columbia.edu/2026/06/16/s...
blog post: should MRP workflow include LOCO-CV ? Individual-level loss orders models differently than the population-level loss we want judging MRP. To get population-level loss, use out-of-sample classical poststratification to compare with MRP. How to split data ? LOCO = leave one cell out.
Survey Statistics: should MRP workflow include LOCO-CV ? statmodeling.stat.columbia.edu/2026/06/09/s...
blog post: it is (still) the people Survey Statistics blog series' 1st birthday 🥳 Andrew Gelman's 60-ish Birthday 🥳 and NYT weights with synthetic past vote 🗳️
Survey Statistics: it is (still) the people statmodeling.stat.columbia.edu/2026/06/02/s...
blog post: double-plus robustness Meng (2022): GREG is not only “double robust” (consistent if either the outcome model or response model are correct), but “double-plus robust” (consistent if what is left of the outcome model and response model are uncorrelated).
Survey Statistics: double-plus robustness statmodeling.stat.columbia.edu/2026/05/26/s...
blog post: GREG GREG is Generalized REGression estimator. We can think of it either as: 1. Adjusting an estimate based on the model with a Horvitz-Thompson estimate of the error, or 2. On the flip side, adjusting the Horvitz-Thompson estimate with the model.
Survey Statistics: GREG statmodeling.stat.columbia.edu/2026/05/19/s...
blog post: relevant alternatives ? We saw that the multinomial logit model implies independence from irrelevant alternatives (IIA). Let’s expand the model to include choice set C within the logits f(X_ic,C), allowing for non-IIA.
Survey Statistics: relevant alternatives ? statmodeling.stat.columbia.edu/2026/05/12/s...
Survey Statistics: Blue Rose Research is (still) hiring ! statmodeling.stat.columbia.edu/2026/05/05/s...
Survey Statistics: Blue Rose Research is (still) hiring ! | Statistical Modeling, Causal Inference, and Social Science
statmodeling.stat.columbia.edu
blog post: work with us at Blue Rose ! use cutting edge statistics, machine learning, and engineering to study public opinion, forecast elections, and advise Democrats.
Survey Statistics: Blue Rose Research is hiring again ! statmodeling.stat.columbia.edu/2026/03/10/s...
blog post: exploded logit ! a common choice model is multinomial logit. this model implies that rankings follow an exploded logit !
Survey Statistics: exploded logit ! statmodeling.stat.columbia.edu/2026/04/28/s...
blog post: irrelevant alternatives ? a common choice model is multinomial logit. this model implies Independence of Irrelevant Alternatives (IIA), e.g. the ratio of Left-vs-Right preference is the same in round 1 as in the runoff.
Survey Statistics: irrelevant alternatives ? statmodeling.stat.columbia.edu/2026/04/14/s...
blog post: improving with structure We’ve met Mr. P (Multilevel Regression and Poststratification). We’ve met Mrs. P (Multilevel Regression with Synthetic Poststratification). Now let’s meet Ms. P (Multilevel Structured regression with Poststratification).
Survey Statistics: improving with structure statmodeling.stat.columbia.edu/2026/04/07/s...
blog post: design-based cross validation how to split train and test sets to respect survey design ? what lessons carry over to nonprobability samples ?
Survey Statistics: design-based cross validation (dCV) statmodeling.stat.columbia.edu/2026/03/31/s...
blog post: Individualism and the CV Noise Problem Politically meaningful differences among models can be swamped by cross-validation noise.
Survey Statistics: Individualism and the CV Noise Problem statmodeling.stat.columbia.edu/2026/03/24/s...
blog post: individualism doesn't work (even when weighted) individual-level loss (even weighted to the population) orders models differently than the population-level loss of interest to folks using MRP
Survey Statistics: individualism doesn’t work (even when weighted) statmodeling.stat.columbia.edu/2026/03/17/s...
blog post: work with us at Blue Rose ! use cutting edge statistics, machine learning, and engineering to study public opinion, forecast elections, and advise Democrats.
Survey Statistics: Blue Rose Research is hiring again ! statmodeling.stat.columbia.edu/2026/03/10/s...
blog post: sampling-weighted loss we use sampling weights to estimate a population mean E(Y). what about to estimate a conditional mean E(Y|X) ? the best-fit model in the sample may not be the best-fit model in the population.
Survey Statistics: sampling-weighted loss statmodeling.stat.columbia.edu/2026/03/03/s...
blog post: sampling to assess data quality @bhedtgauthier.bsky.social et al. (2012) used sampling to assess and improve data quality in Malawi
Survey Statistics: sampling to assess data quality statmodeling.stat.columbia.edu/2026/02/24/s...
blog post: Gallup's Presidential Approval Ratings Gallup will no longer track presidential approval after 88 years Let's look at their sampling, mode, and weighting (still used for other survey questions)
Survey Statistics: Gallup’s Presidential Approval Ratings statmodeling.stat.columbia.edu/2026/02/17/s...
blog post: more on recalled vote we've talked about measurement error in recalled vote in the US. how does this change in multiparty states ?
Survey Statistics: more on recalled vote statmodeling.stat.columbia.edu/2026/02/10/s...