Anna Tsvetkov

@annatsv.bsky.social

Postdoc @ Princeton AI Natural and Artificial Minds Prev: Philosophy PhD @ Brown, MIT FutureTech Website: https://annatsv.github.io/

This is a beautiful paper! The first third helpfully labels a stream of recent work in philosophy of AI as "propositional interpretability". The idea is to use propositional attitudes like belief, desire, and intention, to help explain AI in a way that we can understand. 1/n

DDavid Chalmers@davidchalmers.bsky.social · 2y ago

a draft paper (for an invited talk at AAAI next month) with a philosophical analysis of work on mechanistic interpretability, with special attention to methods for propositional interpretability. arxiv.org/abs/2501.15740

"The AI risk repository, which includes over 700 AI risks grouped by causal factors (e.g. intentionality), and domains (e.g. discrimination), was born out of a desire to understand the overlaps and disconnects in AI safety research" #AIEthics techcrunch.com/2024/08/14/m...

MIT researchers release a repository of AI risks | TechCrunch

A group of researchers at MIT and elsewhere have compiled what they claim is the most thorough databases of possible risks around AI use.

techcrunch.com