Jana Jung

@janajung.bsky.social

PhD student in Social Data Science @ University of Mannheim | jasoju.github.io

Are you using survey-style questionnaires designed for humans to measure characteristics of LLMs? In our #EACL2026 paper, we evaluate both the reliability and validity of such tests and found that their scores do not reflect real-world model behavior. In fact, they can be deceptive! 🧵1/3

Bild

🚨New paper alert🚨 🤔 Ever wondered how the way you write a persona prompt affects how well an LLM simulates people? In our #EMNLP2025 paper, we find that using interview-style persona prompts makes LLM social simulations less biased and more aligned with human opinions. 🧵1/7

Bild

Thrilled to talk about how seemingly small decisions in silicon sampling can have a large impact on simulated survey responses 👀 Join us on Oct 29th! 👈

Joshua Claassen@jclaass.bsky.social · 10mo ago

🚨 Upcoming #CS3Meeting 🚨 @wanlo.bsky.social talks about analytic flexibility in silicon samples on October 29, 3:15 to 4:00 PM CET). Great opportunity to gain novel insights into how survey responses can be generated with #LLMs. Sign up now: ww3.unipark.de/uc/cs3_meeti...

LLMs can generate synthetic survey responses, e.g. for imputation, but how reliable are they? 📋 At #IC2S2, I'll be sharing our research on the robustness of AI-generated responses to perturbations and if they mirror human survey biases. 🤖 Come by my poster on Tuesday between 1:30 and 3:30 p.m.

Bild

Really excited to also present this work at #IC2S2 next week in Norrköping! 🎉 I'd love to discuss how to produce LLM survey responses at my poster on Wed at 13:30 (Poster Session 2, Poster ID 68) 📊

Georg Ahnert @IC2S2@ahnert.eurosky.social · last yr.

LLMs are trained to produce open-ended responses 📝, but most survey items require closed-ended responses instead 📊 This Wed 11:00–12:30 at #ESRA25, I'll discuss the large impact that Answer Production Methods have on prediction results + share recommendations for methods and parameters. 👈

A research setup for the evaluation of Answer Production Methods for closed-ended survey responses from LLMs. An LLM is prompted with a survey and an optional instruction, before a Answer Production Method is applied. These methods range from token-probabilities to open-ended text generation + classification. I then evaluated them against human survey answers and calculate individual-level accuracy as well as distribution alignment for sub-populations.

Very excited to head to #IC2S2 next week! 🎉 In our project, we tested whether a psychological assessment can measure sexism in LLMs, and found that applying such tools to LLMs is not as straightforward as it seems. Find me and my poster at Poster Session 1 (Tue 12:30-14:30) — hope to see you there