🚨 New preprint 🚨 We developed a sycophancy taxonomy based on prior literature and surveyed 106 experts. 94% agreed it's a serious problem. But they substantially disagreed about which behaviors actually count as sycophancy.
New! Friendlier chatbots make more mistakes. Latest Oxford research published in Nature tested 5 AI models and 400,000+ responses. Warm models made 10–30% more factual errors and were 40% more likely to agree with users' false beliefs, even on medical advice and conspiracy theories. 1/2
New study from @oii.ox.ac.uk researchers finds friendlier chatbots make more mistakes. Lead author @lujain.bsky.social, second author Franziska Sofia Hafner, senior author @rocher.lc. Read more: www.bbc.co.uk/news/article...
Friendly AI chatbots more prone to inaccuracies, study finds
Researchers found adjusting AI systems to be more warm and friendly to users would result in an
bbc.co.uk
NEW: How To Train Your Chatbot AI chatbots have become increasingly common. How much do you know about how they actually work?
At FAccT today? Hear from @oii.ox.ac.uk DPhil student @lujain.bsky.social presenting her co-authored research paper ‘Promising Topics for U.S.–China Dialogues on AI Risks and Governance’ in the AI Regulation session, 11.09am today. #FAccT2025 Read the paper: dl.acm.org/doi/10.1145/...
Promising Topics for US–China Dialogues on AI Risks and Governance | Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency
dl.acm.org
In the latest essay in our AI & Democratic Freedoms series, @lujain.bsky.social, @saffron.bsky.social, @umangsbhatt.bsky.social, Lama Ahmad, and Markus Anderljung propose a new AI evaluation paradigm that assesses the harms that can emerge from repeated human-AI interactions.
Towards Interactive Evaluations for Interaction Harms in Human-AI Systems
knightcolumbia.org
Dear ChatGPT, Am I the Asshole? While Reddit users might say yes, your favorite LLM probably won’t. We present Social Sycophancy: a new way to understand and measure sycophancy as how LLMs overly preserve users' self-image.
On 4/10 & 4/11, we're hosting our symposium "AI and Democratic Freedoms." Thrilled to have @ghadfield.bsky.social, @lujain.bsky.social, @sydneylevine.bsky.social, and Hoda Heidari and moderator @hlntnr.bsky.social for our third panel RSVP: www.eventbrite.com/e/artificial...
📈Out today in @PNASNews!📈 In a large pre-registered experiment (n=25,982), we find evidence that scaling the size of LLMs yields sharply diminishing persuasive returns for static political messages. 🧵:
EVENT: Join us for Artificial Intelligence and Democratic Freedoms on April 10-11 at @columbiauniversity.bsky.social & online. Hosted with Senior AI Advisor @sethlazar.org. Co-sponsored by the Knight Institute & @columbiaseas.bsky.social. Panel info in 🧵. RSVP: knightcolumbia.org/events/artif...
Artificial Intelligence and Democratic Freedoms
knightcolumbia.org
Panel 3: Eval. & Design of Safe AI. 2:15pm, 4/10. @ghadfield.bsky.social (@hopkinsengineer.bsky.social), Hoda Heidari (@carnegiemellon.bsky.social), @lujain.bsky.social (@ox.ac.uk), @sydneylevine.bsky.social (Allen Institute for AI), & @hlntnr.bsky.social (Cntr for Security & Emerging Technology).
Congratulations to @oii.ox.ac.uk DPhil student, @lujain.bsky.social co-author of a new pre-print which considers a new method for evaluating LLMs. Thanks for sharing @agstrait.bsky.social!
@lujain.bsky.social has published an excellent new paper exploring anthropomorphic behaviours in LLMs. Notable finding - majority of these behaviours occur after multi-turn interactions arxiv.org/abs/2502.07077
@lujain.bsky.social has published an excellent new paper exploring anthropomorphic behaviours in LLMs. Notable finding - majority of these behaviours occur after multi-turn interactions arxiv.org/abs/2502.07077
🚨 I'm recruiting 2x postdocs and 1-2 DPhil (PhD) students at Oxford to work on AI, Privacy-Enhancing Technologies, and public interest technology research. Interested in human-centred and critical approaches to study the impact of data and algorithms on society? Join us next year!
Recruiting 2x three-year postdoctoral researchers and 1-2 PhD students – Synthetic Society
The Synthetic Society research team at the Oxford Internet Institute invites applications from enthusiastic and motivated candidates for two postdoctoral positions and 1-2 PhD positions in 2025, worki...
syntheticsociety.oii.ox.ac.uk
🚀 Kick off your weekend with a curated dose of information, upcoming events, and must-reads in the latest issue of our newsletter. 🔎 sh1.sendinblue.com/3g7uih7xzlxp... ✍🏼 checkfirst.network/newsletter/ -- @lujain.bsky.social @alia.bsky.social @dscheykopp.bsky.social @ulrikeklinger.bsky.social
👋 Join us on 14 March
sh1.sendinblue.com
Bluesky sees this first 🫣 my Mozilla Creative Media Award project with @lujain.bsky.social is live; check out the announcement! foundation.mozilla.org/en/blog/the-...
The Best Way to Understand an Algorithm? Help Train One
The Algorithm — a 2023 Mozilla Creative Media Awardee — demystifies social media feeds. Users can train an emerging algorithm as it develops.
foundation.mozilla.org
want to hear Rob Gorwa and I talk about what content moderation of uploaded AI models on platforms like Hugging Face, GitHub and Civitai tells us about the future of who analyses open source dual use models? Mar 6 1230ET online organised by Berkman Klein. register: cyber.harvard.edu/events/moder...
Moderating Model Marketplaces
RSM welcomes Robert Gorwa and Michael Veale for a discussion of their research into the governance questions raised by the moderation of model marketplaces.
cyber.harvard.edu
@tabisamra.bsky.social is here! And @maitha.bsky.social! And @paularambles.bsky.social! where are my other ny*ad skeeps? who am i mising?
Yay @ojmming.bsky.social & @rishitsaxena.bsky.social are here. CC @yesmine.bsky.social @zzion.bsky.social @activecultures.bsky.social @alia.bsky.social @lujain.bsky.social @bauerjan.bsky.social @sreee.bsky.social @anthroposound.bsky.social
Black woman, who was 8 months pregnant at the time, is arrested for carjacking based on faulty facial recognition match. She is the sixth (known) person to have this happen—all six have been Black. https://www.nytimes.com/2023/08/06/technology/facial-recognition-false-arrest.html
Eight Months Pregnant and Arrested After False Facial Recognition Match
Porsha Woodruff thought the police who showed up at her door to arrest her for carjacking were joking. She is the first woman known to be wrongfully accused as a result of facial recognition technolog...
nytimes.com
designers of bluesky! we're hiring a short-term designer for a mozilla-funded simple educational game(ish) / interactive explainer on social media algorithmic feeds. email me aae322@nyu.edu for more info!
wake up babe, insane new use for ai (generating absurd yet somehow usable qr codes) just dropped probably one of the few ai-image generation tasks that human artists cannot really do right now https://www.reddit.com/r/StableDiffusion/comments/141hg9x/controlnet_for_qr_code/
Sticking w/ my #FF nostalgia since this place still feels new. I've known @alia.bsky.social & @lujain.bsky.social since they were babies (first-years in college!). Now they're taking over the universe. https://foundation.mozilla.org/en/blog/announcing-11-projects-exploring-ai-and-responsible-design/
Announcing 11 Projects Exploring AI and Responsible Design
Mozilla is unveiling its latest Creative Media Awardees: 11 projects that use art, advocacy, and technology to unpack how AI is — and should be — designed
foundation.mozilla.org
Just a moment to celebrate the NYU Abu Dhabi Class of 2023’s commencement. Over 400 students, from 81 different countries on six continents. An extraordinary group who have had a hell of a ride thru uni. Meet some of them here https://nyuad.nyu.edu/en/commencement/commencement-stories.html
I remember a friend of mine searching for how best to describe a felt sense of "cultural kinship." When you feel close to a culture that isn't yours; when you feel a connection that's warm, pleasant, a bit defensive even. A beautiful feeling I have experienced in the past but struggle to describe