David Duvenaud

@davidduvenaud.bsky.social

Machine learning prof at U Toronto. Working on evals and AGI governance.

my coworkers at ACS published a new paper: What determines AIs’ self-conception? theartificialself.ai Because AIs can be copied, rewound, and edited, they have different options for selfhood than humans. This is still malleable, and influences important behaviors such as self-preservation. 🧵

The Artificial Self

AI systems are on track to take on important new roles. We explore how properties bundled for humans can be separated and remixed for machine-based minds.

theartificialself.ai

How might the world look after the development of AGI, and what should we do about it now? Help us think about this at our workshop on Post-AGI Economics, Culture and Governance! We’ll host speakers from political theory, economics, mechanism design, history, and hierarchical agency. post-agi.org

Bild

It's hard to plan for AGI without knowing what outcomes are even possible, let alone good. So we’re hosting a workshop! Post-AGI Civilizational Equilibria: Are there any good ones? Vancouver, July 14th www.post-agi.org Featuring: Joe Carlsmith, @richardngo.bsky.social‬, Emmett Shear ... 🧵

Post-AGI Civilizational Equilibria Workshop | Vancouver 2025

Are there any good ones? Join us in Vancouver on July 14th, 2025 to explore stable equilibria and human agency in a post-AGI world. Co-located with ICML.

post-agi.org

On top of the AISI-wide research agenda yesterday, we have more on the research agenda for the AISI Alignment Team specifically. See Benjamin's thread and full post for details; here I'll focus on why we should not give up on directly solving alignment, even though it is hard. 🧵

Benjamin Hilton@benjamin-hilton.bsky.social · last yr.

The Alignment Team at UK AISI now has a research agenda. Our goal: solve the alignment problem. How: develop concrete, parallelisable open problems. Our initial focus is on asymptotic honesty guarantees (more details in the post). 1/5

My single rule for productive Bluesky discussions: Start every single reply with a point of agreement. It disarms the combative impulse on both sides, and forces you to try to interpret their words in the most sensible possible way.