Alexandra Olteanu

@aolteanu.bsky.social

Ethical/Responsible AI. Rigor in AI. Grumpy eastern european in north america. Lovingly nitpicky. www.aolteanu.com

So what do AI evaluations do or should do? Wrote a brief reflection about prototypical questions evaluations often try to answer and why it is useful to have clarity about what exactly one is trying to learn about some phenomenon of interest. rigor-in-ai.leaflet.pub/3ms4ax4sojs2j

Back-to-basics: so what do evaluations do

TL;DR — Developing useful evaluations requires clarity about what exactly one is trying to learn about a phenomenon of interest.

rigor-in-ai.leaflet.pub

Wrote another brief post motivated by a reflection on how in AI research folks might not fully appreciate the care that working with unobservable constructs requires. TL;DR -- Because unobservable constructs hold ‘surplus meaning,’ your metric is not your construct.

Back-to-basics: unobservable constructs and their ‘surplus meaning’

TL;DR — Because unobservable constructs hold ‘surplus meaning,’ your metric is not your construct.

rigor-in-ai.leaflet.pub

Wrote a brief reflection on poor conceptualizations in AI work and why they matter, something I believe is so startlingly neglected. TL;DR - Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims.

Back-to-basics: on poor conceptualizations in AI work

TL;DR — Poor conceptual foundations can severely undermine the credibility and reliability of knowledge claims. (And, no, your metric is not your construct.)

rigor-in-ai.leaflet.pub

I'm thinking about writing brief reflections on a few topics (mainly from my own attempts to gain clarity about research practices I find concerning). I'm thinking about using medium, but I see many folks use substack. Are there other platforms? What's your reasoning for using a certain platform?

Come meet me at the last poster session at ICML if you are interested in chatting about LLM evaluation and pairwise comparisons! I will be presenting joint work with my advisor (Moritz Hardt). Thu Jul 9, 5:00 PM – 6:45 PM KST, Hall A (#4411) arxiv.org/abs/2606.09409

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge bi...

arxiv.org

Sometimes it feels like some folks are losing the plot about what the goal of publishing research actually is. The goal is certainly not meant to be the productions of papers, but rather the production and communication of science.

Yesterday was my last day at MSR. We recently learned that our roles were eliminated, and with them our little FATE Montreal team. I joined MSR a bit over 7.5 years ago while on active chemotherapy, and being at MSR has overlapped with so much change in my life.

We are hoping to see applications from and bring together a diverse group of students across multiple disciplinary areas! If you are a graduate student and interested in FAccT’s scope, the Doctoral Colloquium is for you! #facct2026 #facct26

ACM FAccT@facct.bsky.social · 7mo ago

Calling all graduate students! The #FAccT2026 call is out for our Doctoral Colloquium. Join us for connections, support, mentorship, etc. Apply by February 24! facctconference.org/2026/callfor...

I wish folks would use more precise terminology than "AI sycophancy." Not all validating behaviours/interactions are sycophantic. By definition, for them to be sycophantic there needs to be an underlying intention to e.g., gain advantage or favour. Intention is something AI systems do not have.

This was accepted to #NeurIPS 🎉🎊 TL;DR Impoverished notions of rigor can have a formative impact on AI work. We argue for a broader conception of what rigorous work should entail & go beyond methodological issues to include epistemic, normative, conceptual, reporting & interpretative considerations

Alexandra Olteanu@aolteanu.bsky.social · last yr.

We have to talk about rigor in AI work and what it should entail. The reality is that impoverished notions of rigor do not only lead to some one-off undesirable outcomes but can have a deeply formative impact on the scientific integrity and quality of both AI research and practice 1/

Print screen of the first page of a paper pre-print titled "Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor" by Olteanu et al.  Paper abstract: "In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about AI capabilities. Our position is that a broader conception of what rigorous AI research and practice should entail is needed. We believe such a conception -- in addition to a more expansive understanding of (1) methodological rigor -- should include aspects related to (2) what background knowledge informs what to work on (epistemic rigor); (3) how disciplinary, community, or personal norms, standards, or beliefs influence the work (normative rigor); (4) how clearly articulated the theoretical constructs under use are (conceptual rigor); (5) what is reported and how (reporting rigor); and (6) how well-supported the inferences from existing evidence are (interpretative rigor). In doing so, we also aim to provide useful language and a framework for much-needed dialogue about the AI community's work by researchers, policymakers, journalists, and other stakeholders."

My university has announced a fund to essentially poach doctoral students from US institutions. DM me if you do work on the history/social impacts of AI and are interested in being poached 😂

Not sure who needs to hear this but what people want AI systems to do, what AI systems do, and what people believe AI systems do are not the same thing. Just because one wants or believes AI systems do or can do certain things, doesn't mean they actually do those things.