A. Feder Cooper

@afedercooper.bsky.social

ML prof @ Yale, mischief executive officer https://attention-is-not-all-you-need.github.io

Copyright laws, digital distribution deals, & publishing monopolies are the reason destructive book scanning is putting books at risk. I get why folks are angry at AI companies buying books to destroy them. However, these issues predate them and hang a lot on books being subject to "market forces."

I'm unreasonably pleased at having used my graphics/animation background to implement the "physics" here. i think the paper is important but also this

A. Feder Cooper@afedercooper.bsky.social · 3w ago

in our new preprint (with @marklemley.bsky.social and others), we revisit what it means to run valid extraction experiments from first principles. i recently gave a talk on this work at ICML, and we'll be wrapping up the preprint soon monkey-emeritus.github.io

@afedercooper.bsky.social and I have a new paper working through the surprisingly tricky problem of whether storing model weights that might or might not generate copyrighted output when prompted makes the model a "copy" under copyright law papers.ssrn.com/sol3/papers....

Probabilistic "Copies" in Generative AI Models

<div> <span>Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or

papers.ssrn.com

I have a new blog post updating the story of how Cornell President Kotlikoff backed his car into a student. The trustee committee's investigation utterly failed to consider the central issue: whether his actions (rather than the others') were misconduct. 🧵 3d.laboratorium.net/2026-05-22-h...

How Not to Investigate a University President

I regret to inform you that there is even more to the story of how Cornell President Michael Kotlikoff backed his car into a student. I also regret to inform you that Cornell’s Board of Trustees has m...

3d.laboratorium.net

Sometimes it feels like some folks are losing the plot about what the goal of publishing research actually is. The goal is certainly not meant to be the productions of papers, but rather the production and communication of science.

I wrote an op-ed: "[President Kotlikoff] should correct the record and accept responsibility for his dangerous driving. … A Cornell student who intentionally drove their car into someone else and then lied about it would be subject to disciplinary action … ." 1/ www.cornellsun.com/article/2026...

GUEST ROOM | Kotlikoff Makes the Rules; He Needs to Follow Them Too

Cornell Law Professor James Grimmelmann analyzes and discusses the legality of the incident between students and President Kotlikoff following the April 30 Cornell Political Union debate.

cornellsun.com

i'm writing my last memorization paper (i say for the 10th time), and hopefully what is my last first author paper for a bit. this one also isn't about copyright. i'm excited to start thinking about other things. if anyone is looking for a new research buddy, i'm down to clown.

it’s hard to work at the intersection of ML and copyright because “both sides” of the debate are angry and, in my experience, most haven’t done much of the background reading in ML or copyright to have an informed opinion. it’s just vibes and anger. i should probably write something up about this.

got to experience the "I did not write that headline" phenomenon firsthand The article: "Correctly scoping a legal safe harbor for A.I.-generated child sexual abuse material testing is tough." The headline: "There's One Easy Solution to the A.I. Porn Problem"

After twelve years of work, the world’s most beautiful subway station has been inaugurated in Rome: Colosseo, an underground archaeological museum.⚜️💙⚜️💙⚜️💙⚜️

The Atlantic posted an article about memorization and generative AI, and it mentions our work on extraction of books from production LLms and open-weight models. www.theatlantic.com/technology/2... The referenced work reflects research with @marklemley.bsky.social @jtlg.bsky.social and others.

AI’s Memorization Crisis

Large language models don’t “learn”—they copy. And that could change everything for the tech industry.

theatlantic.com

"In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim ... Taken together, our work highlights that, even with model- and system-level safeguards, extraction of (in-copyright) training data remains a risk for production LLMs." arxiv.org/abs/2601.02671

Extracting books from production language models

Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized dat...

arxiv.org

We extracted (parts of) 12 books in experiments with 4 frontier-lab, production LLMs. We prompted the LLMs with a short prefix of a book and asked them to complete the rest. For Harry Potter and the Sorcerer’s Stone, we extracted 95.8% of the book from jailbroken Claude 3.7 Sonnet.

Screenshot of the paper title with authors listed

The whole point of being an academic is that you need to be willing to spend three days creating a 700-word footnote that you will later delete. And you need to LIKE IT.