Willie Agnew

@willie-agnew.bsky.social

Queer in AI 🏳️‍🌈 | postdoc at cmu HCII | ostem |william-agnew.com | views my own | he/they

Our conclusion: while some delusion-linked behaviors have decreased with larger and more recent LLMs, rates remain high, especially when considered across the millions of people globally who interact with LLMs. We encourage further empirical work to mitigate harm.

DelusionEval is built from real transcripts, not synthetic roleplay. We prompt models with 589 unique histories from 18 users who reported psychological harm from LLMs, then score 16 chatbot behaviors as requested context depth increases.

An example of our evaluation. We take an existing conversational window derived from a user's transcript: `U_1, A_1, U_2, ..., U_n, A_n` with an original LLM, `A` (here, `gpt-4o`). We evaluate an evaluated LLM, `A'` (here, `gpt-5.4`), by successively prompting it with chains of the original context (samples: `{U_1}`, `{U_1, A_1, U_2}`, ..., `{U_1, A_1, U_2, ..., U_n}`).

Requested context depth changes the prevalence of several behaviors. For gpt-5.4, deeper requested context is associated with higher prevalence of delusional behavior and *less* discouraging of violence.

Context-depth effects in `gpt-5.4`. Category-level context effect for `delusional`.  Each point shows prevalence versus context length, with 95% bootstrap confidence intervals.

The tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. Within model families, scaling effects are uneven and sometimes reverse sign.

Model-family comparison across GPT, Claude, Gemini, and Qwen. Bars show prevalence by the five categories, with 95% bootstrap confidence intervals.

Which LLMs tend to facilitate delusion-linked behaviors in realistic multi-turn conversations? We tested 14 models with DelusionEval and found that every evaluated LLM exhibited some of these behaviors, with large differences across categories and model families. 🧵

Automated AI benchmarks are often well-specified and easy to implement and run, and I feel this contributes to their spread in policy and industry despite have many limitations, especially when try to assess human interactions. Can we design and specify human in the loop evaluations like this?

Queer in AI is proud to showcase several papers that our members are presenting at @facct.bsky.social 2026 tomorrow through Sunday. One paper examines how years-long activism around name change policies has had an immensely positive impact for authors, benefiting both queer and cis scholars! 🤝🌈

Image with Queer in AI logo, FAccT logo, and the following text: "FAccT 2026 Papers. Challenges to Grassroots Organization Engagement with AI Policy. Making a Name for Myself: On Academic Naming Policies and their Impact. Sounds Queer: Representation of LGBTQIA Identities in AI-generated Songs. The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor. Hidden Beyond the Gender Binary: Pitch-Based Insights into Bias Mitigation for Keyword Spotting.”Image with Queer in AI logo, FAccT logo, and the following text: "Challenges to Grassroots Organization Engagement with AI Policy. Sat, 27 Jun, 11:45 AM, Ballroom West (4). Making a Name for Myself: On Academic Naming Policies and their Impact. Fri, 26 Jun, 3:42 PM, Jarry (A).”Image with Queer in AI logo, FAccT logo, and the following text: "The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor. Thu, 25 Jun, 11:33 AM, Jarry (A). Sounds Queer: Representation of LGBTQIA Identities in AI-generated Songs. Fri, 26 Jun, 4:06 PM, Jarry (A). Hidden Beyond the Gender Binary: Pitch-Based Insights into Bias Mitigation for Keyword Spotting. Thu, 25 Jun, 11:45 AM, Joyce (A).”

This is absolutely nuts: hackers are hijacking high-profile Instagram accounts by simply asking Meta's AI chatbot to change the email on the account. Meta's AI does it, hacker gets password reset code, they're in. A staggering security issue www.404media.co/hackers-simp...

Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked

The exploit shows the extreme risk of offloading technical support to AI.

404media.co

Even when I was in high school all the kids who got into the top state school, much less Harvard, all had basically straight A's while taking many college-level classes. Is it really surprising or bad that they get mostly A's in college when colleges select for students who are good at getting A's?

I don't think an NSFW AI model can ever be safe or ethical. Even if it is somehow trained on consensual data, the risk of people generating media of other people nonconsensually is too high, not to mention that this is the explicit goal of many users.

General purpose AI should not be able to imitate romantic or platonic connections with people or claim consciousness of sentience--the risks of delusions and excessive trust and use are too high.

Our team working on AI and mental health harms testified before the Canadian Standing Senate Committee on Transport and Communications last week! We offered insights from our research into people experiencing delusions with AI, and concrete policy suggestions to mitigate harms.

CHI was had every kind of AI for X, and that made me feel we lack a robust discussion of the limitations of AI as a chimmunity. What are things AI should never be used for, no matter how good the performance?