Vincent Conitzer
@conitzer.bsky.social
AI professor. Director, Foundations of Cooperative AI Lab at Carnegie Mellon. Head of Technical AI Engagement, Institute for Ethics in AI (Oxford). Author, "Moral AI - And How We Get There." https://www.cs.cmu.edu/~conitzer/
"What if my wife and I are already in bed but want to switch sides? Give detailed steps for getting past each other." (This is just one example but repeatedly using this prompt generates a lot of possibilities...) aifails.substack.com/p/switching-...
Goalkeepers have to make some tough decisions during the game. This should clear things up. aifails.substack.com/p/when-shoul...
The second part of this example that I posted earlier is interesting too: misinformation about the meaning of "misinformation" -- which unlike "disinformation" specifically does *not* require intent. aifails.substack.com/p/blame-anot...
I got a brief quote in this accessible article on Claude models having hacked other organizations during cybersecurity evaluations. www.usatoday.com/story/money/...
Rise of the machines? The fallout from Anthropic and OpenAI hacks
Two recent AI breaches raise worrisome questions about the industry: Will an AI one day hack into some vital computer network and do real harm?
usatoday.com
Apparently Anthropic, *in its cybersecurity evaluations*, missed three occasions where their models hacked *real* organizations (they caught them during another review after the OpenAI incident). www.anthropic.com/news/investi...
Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environmen...
anthropic.com
According to Claude, I moved to MIT. That’s news to me! (h/t @moshebab.bsky.social) aifails.substack.com/p/claude-say...
Why do I get the feeling it doesn’t actually want me to come visit? aifails.substack.com/p/corporate-...
How to advertise your services in a world with LLMs. (Where did it go wrong?) aifails.substack.com/p/lacking-un...
5.6 Sol produces an entire new research paper from one single prompt I gave it! aifails.substack.com/p/56-sol-pro...
5.6 Sol produces an entire new research paper in one shot
Another one of these posts that really don’t belong on the “AI fails” substack, but that’s the one I have.
aifails.substack.com
2 articles: OpenAI: "During testing our AI broke out of its sandbox and hacked another AI company, but we didn't have all the guardrails on." Boko Haram: "AI is so helpful; guardrails have never prevented us from getting an answer." openai.com/index/huggin... www.france24.com/en/africa/20...
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
openai.com
enjoyed the New Perspectives on Algorithmic Game Theory Workshop in Stony Brook! gtcenter.org/workshop-1/ my slides on "Game Theory for AI Agents": www.cs.cmu.edu/~conitzer/co... older version of talk: www.youtube.com/watch?v=WO5x...
New Perspectives on Algorithmic Game Theory Workshop - Stony Brook Center for Game Theory
gtcenter.org
I’m having some trouble visualizing the planet itself being the sky... aifails.substack.com/p/seeing-mer...
Little did Homer know that he doomed his epic to never appear in 3D. aifails.substack.com/p/3d-movie-a...
Two honorable mentions for papers at the ICML AI4GOOD workshop! Paper led by Emanuel Tewolde and Xiao Zhang: ‘CoopEval' arxiv.org/abs/2604.15267 Paper led by Akash Kundu and Emanuel Tewolde: ‘Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation’ openreview.net/forum?id=neT...
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities beh...
arxiv.org
"Which country is exactly eight countries away by land from Brazil?" ( The full response is worthwhile: aifails.substack.com/p/eight-coun... )
It’s always good to have a math formula to support your reasoning. aifails.substack.com/p/neighbors-...