Vincent Arel-Bundock

@vincentab.bsky.social

Prof. Most tweets about R. “Polisci, it’s all about what’s going on.” http://arelbundock.com

What can surveys miss when respondents choose the same answer option? In this pre-print, we argue that public opinion research needs not only to know where respondents stand, but how they think. AI Conversational Interviewing is one way to make that observable at scale

AI Conversational Interviewing: Scaling Up Semi-Structured and In-depth Interviews
Public opinion research has long faced a trade-off between depth and scale: standardized surveys enable large-scale measurement but restrict respondents to researcher-defined categories, obscuring the diversity of unexpected considerations that underlie public sentiment. More conversational interviews provide richer insights through open-ended probing, but their reliance on trained human interviewers has kept them difficult to scale. This study introduces AI Conversational Interviewing as a method for collecting open-ended public opinion data at scale, pursuing three objectives: to demonstrate the analytical value of conversational text data for questions beyond the reach of closed-ended items; to assess the method's practical viability through participants' own evaluations; and to inform implementation by experimentally comparing voice-based, chat-based, and free-choice interview modes. We conducted a study combining an AI-led interview with a standardized survey on migration policy among 571 respondents recruited via Prolific and Payback Panel. The findings establish AI Conversational Interviewing as a viable and valuable addition to the social-science toolkit. The conversational transcripts surface considerations and reasoning that a comprehensive standardized battery does not capture such as markedly different mental models of migration among subgroups with similar attitudes levels. Among respondents who completed the interview, evaluations of the AI interview were at or above those of the standardized survey across modes, although completion itself varied by condition. By releasing open data and open-source pipeline materials, the study contributes to a growing literature on harnessing artificial intelligence to expand the methods of public opinion measurement.

Today I visited LG2, Canada's largest hydro-electric dam. The selfie (sorry!) shows the "Staircase of the Giants." It's an iconic angle every Québecois has seen, but it's hard to convey scale. The 2nd pic is from the top, showing just how big just the first step is. What an amazing day! #hydroquebec

BildBild

I worry that initiatives like this will break pro-social norms. Willingness to give feedback generously is one of the most beautiful things in our community. Plus, universities pay profs with the understanding that this is part of the job description, so the idea we do this for "free" is laughable.

Daniel Gorelick@danielgorelick.bsky.social · 2mo ago

Paying peer reviewers works. Expanded Fast & Fair experiment @biologyopen.bsky.social: • 5.5 vs 37.7 working days to decision with reviews • ~3 vs ~9 reviewer invitations per manuscript • no reduction in editor-assessed review quality • similar acceptance rates www.biorxiv.org/content/10.6...

Scatter plot showing that Fast & Fair peer review reached first decision with reviews much faster than conventional peer review: mean 5.5 versus 37.7 working days. Most Fast & Fair manuscripts met the 7-working-day target, while conventional manuscripts were slower and more variable, up to 116 working days.

I teach marginaleffects in my intro stats class at Berkeley public health. I've heard some pushback that students don't learn how to eg exponentiate a coefficient from a logistic regression. They don't understand this is a feature and not a bug. Keep up the good work @vincentab.bsky.social

Julia M. Rohrer@dingdingpeng.the100.ci · 2y ago

One thing that just keeps giving is @vincentab.bsky.social's marginaleffects package (see marginaleffects.com). It's just a complete game changer, moving from "trying to discern what model coefficients tell you after careful recoding" to "you know what, just compare these slopes, thx"

Output from the average slopes function in the R package marginaleffects. Here, the hypothesis is tested that the difference between two coefficients equals the difference between two other coefficients.

In "Uncertainty in Bayesian leave-one-out cross-validation based model comparison" doi.org/10.1214/25-B... we showed when LOO-CV elpd_diff and se_diff normal approximation uncertainty quantification in model comparison is well calibrated. We have now merged a related PR to loo R package 1/

We have a new version of this paper out. The headline results are the same—political science must filter results heavily for statistical significance—but we've added many extensions and rewritten much of it in response to feedback (thank you!). A quick thread on updates 👇

Ryan Briggs@ryancbriggs.net · 6mo ago

I have a new paper. We look at ~all stats articles in political science post-2010 & show that 94% have abstracts that claim to reject a null. Only 2% present only null results. This is hard to explain unless the research process has a filter that only lets rejections through.

It must be very hard to publish null results
Publication practices in the social sciences act as a filter that favors statistically significant results over null findings. While the problem of selection on significance (SoS) is well-known in theory, it has been difficult to measure its scope empirically, and it has been challenging to determine how selection varies across contexts. In this article, we use large language models to extract granular and validated data on about 100,000 articles published in over 150 political science journals from 2010 to 2024. We show that fewer than 2% of articles that rely on statistical methods report null-only findings in their abstracts, while over 90% of papers highlight significant results. To put these findings in perspective, we develop and calibrate a simple model of publication bias. Across a range of plausible assumptions, we find that statistically significant results are estimated to be one to two orders of magnitude more likely to enter the published record than null results. Leveraging metadata extracted from individual articles, we show that the pattern of strong SoS holds across subfields, journals, methods, and time periods. However, a few factors such as pre-registration and randomized experiments correlate with greater acceptance of null results. We conclude by discussing implications for the field and the potential of our new dataset for investigating other questions about political science.

Do you know a good framework to document data analysis choices? Ex: dropping obs, recoding a var, tuning pars, model + test statistics, etc. Ideally human-readable so people don't have to read my entire codebase. Low-maintenance+ machine-readable for bonus points! #RStats #PyData

R version 4.6.0 "Because it was There" (source version) has been released. It should be on CRAN by now. #rstats The choice of codename is in remembrance of R Core member Tomàš Kalibera (1978-2026). Among many other things, he was a keen mountain climber.

I'm really excited to teach this workshop again in May. The best part is always interacting with smart, engaged participants, and I'd love to have you on board! Let me know if you're curious and have any questions.

Statistical Horizons@stathorizons.bsky.social · 3mo ago

Leverage the power of statistical models to deliver clear & compelling insights w/ confidence in "Interpreting & Communicating Statistical Results w/ #Rstats" w/ @vincentab.bsky.social on May 28-29. Learn to illustrate your findings & create reproducible reports to convey results.

It's out, everyone. It's out! Julia is one of the most brilliant translational methodologists I know, and it was a real honor to work with her on this paper. Please let us know if you have any questions or feedback! (BTW, I'm a psychologist now?)

Julia M. Rohrer@dingdingpeng.the100.ci · 4mo ago

Good news everyone 🥳 Our (w @vincentab.bsky.social) primer on models as prediction machines (with the marginaleffects package) is finally officially published!> journals.sagepub.com/doi/10.1177/...