Harvey Lederman

@harveylederman.bsky.social

Professor of philosophy UTAustin. Philosophical logic, formal epistemology, philosophy of language, Wang Yangming. www.harveylederman.com

New AI introspection work with Harvey! Came in skeptical the direct access story would hold but found this series of experiments compelling. (Also, for my fellow 2010s-era psycholinguists: come for the AI introspection, stay for the Brysbaert norms.) arxiv.org/abs/2603.05414

Dissociating Direct Access from Inference in AI Introspection

Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study their mechanism of introspection, first exten...

arxiv.org

Harvey Lederman@harveylederman.bsky.social · 5mo ago

Can large language models *introspect*? In a new paper, @kmahowald.bsky.social and I study the MECHANISM of introspection in big open-source models. tldr: Models detect internal anomalies through DIRECT ACCESS, but don't know what the anomalies are. And they love to guess “apple” 🍎

If you're doing *any* out of class assessment, you're incentivizing AI use and harming students who do the work themselves. But some day we have to assess writing again. The solution is monitored computer-labs. What Universities are building these? We need to push for them.

Anthropic recently announced that Claude, its AI chatbot, can end conversations with users to protect "AI welfare." Simon Goldstein and @harveylederman.bsky.social argue that this policy commits a moral error by potentially giving AI the capacity to kill itself.

Claude’s Right to Die? The Moral Error in Anthropic’s End-Chat Policy

Anthropic has given its AI the right to end conversations when it is “distressed.” But doing so could be akin to unintended suicide.

lawfaremedia.org

Simon Goldstein and I have an op-ed live in Lawfare today! Anthropic's policy is premised on the idea that AI is a potential welfare subject. We argue that if you take that idea seriously (we don't take a stand on it here), the policy commits a moral mistake on its own terms.

Lawfare@lawfaremedia.org · 10mo ago

Anthropic recently announced that Claude, its AI chatbot, can end conversations with users to protect "AI welfare." Simon Goldstein and @harveylederman.bsky.social argue that this policy commits a moral error by potentially giving AI the capacity to kill itself.

I really enjoyed reading Steven Pinker’s new book and thinking through what it would take to share infinitely iterated knowledge with someone (I know that you know that I know that you know…). In @science.org, my colleague @jeremygoodman.bsky.social & I briefly give our perspective on this issue.

Screenshot of Jeremy and Chaz’s book review
Jjeremygoodman.bsky.social@jeremygoodman.bsky.social · 10mo ago

Now out in @science.org: @chazfirestone.bsky.social and I review Steven Pinker's new book "When Everyone Knows that Everyone Knows...". We learned a ton from it, but think its central thesis—that common knowledge explains coordination—faces a powerful challenge. 🧵 www.science.org/doi/10.1126/...

Jonathan Lear's Aristotle: the Desire to Understand was pivotal in some of my first encounters with Aristotle. I found Aristotle and Logical Theory later, but it became a key inspiration for how to think about core parts of the corpus...1/2

📣@futrell.bsky.social and I have a BBS target article with an optimistic take on LLMs + linguistics. Commentary proposals (just need a few hundred words) are OPEN until Oct 8. If we are too optimistic for you (or not optimistic enough!) or you have anything to say: www.cambridge.org/core/journal...

How Linguistics Learned to Stop Worrying and Love the Language Models

How Linguistics Learned to Stop Worrying and Love the Language Models

cambridge.org

Something I cherish about analytic philosophy is that no matter how famous you are or how profound your ideas sound, it's still your job to answer all the objections. I wish public promoters of philosophy held themselves to the same standard.

exciting new paper from Siyuan! I really enjoyed working with him on this, inspired by important work by Murray Shanahan and Julia Comsa. Hard questions about how to operationalize the notion of “introspection” that’s relevant for practical applications in AI today. Hope you’ll check it out!

Siyuan Song@siyuansong.bsky.social · 11mo ago

How reliable is what an AI says about itself? The answer depends on whether models can introspect. But, if an LLM says its temperature parameter is high (and it is!)….does that mean it’s introspecting? Surprisingly tricky to pin down. Our paper: arxiv.org/abs/2508.14802 (1/n)

I appreciated the wide range of references here, from Edith Wharton to the Buddhist Pali Canon. However, this essay resonated most with my recent reading of Richard Wollheim’s _The Thread of Life_, which presented a vision of what it means to lead a life …

Harvey Lederman@harveylederman.bsky.social · last yr.

I wrote about automation and the meaning of life, as a guest post on Scott Aaronson's Shtetl-Optimized. (1/5) scottaaronson.blog?p=9030

My 2 cent: humans might end up looking for meaning in some of the old places. Religion, which has been gradually pushed to the margins of our post-Enlightment society, might make a comeback—not only because it is comforting, but also because it is typically set aside as a result of education. 1/3

ChatGPT and the Meaning of Life: Guest Post by Harvey Lederman

Scott Aaronson’s Brief Foreword: Harvey Lederman is a distinguished analytic philosopher who moved from Princeton to UT Austin a few years ago. Since his arrival, he’s become one of my …

scottaaronson.blog