the sweetest thing about me is that there's nothing I love more than seeing people succeeding even if I don't get it and don't like them. hell yeah, dawg, get it
Ari
@ari-holtzman.bsky.social
Assistant Professor @ UChicago CS & DSI UChicao Leading Conceptualization Lab http://conceptualization.ai Minting new vocabulary to conceptualize generative models.
One hope I have for the 21st century is to restore people's respect for the unexamined life.
LLMs don't perfectly remember the documents in their training set. But has anyone characterized how their reproductions differ from the original?
An LLM fintuned on a book does not act like a human who read a book; it doesn't have consistent episodic memories about the experience, etc. What does the finetuning corpus that produces an LLM that acts like a human who read book look like?
There's one very real way that evolution is happening in LLMs: through the selection of attributes that are heritable and selected for in synthetic data. Have we isolated any of these yet?
Nabakov absolute would have use toprified and y'all would have loved it
english isn't for cowards. if you're not saying 'torpify,' that's a skill issue. there's no Académie Française telling you what to do. if you try to type 'pallific' and freak out when you see a little red underline, maybe you're not ready to roll with the big dogs.
swashbuckling happens when the things you care about are dynamic and you're playing around with them directly enough that you can break them if you're not careful
For a given LLM there must be certain data drawn from distribution D, such that if you finetune on them the LLM performs worse on D. It misgeneralizes, due to its priors, as we all do. Are there any interesting cases of this that don't feel totally adversarial and artificial?
People often belittle their sentimentality: "Oh I just like this album because I was a teen in the 80's" as if everything was just association games. Tell me what's 🪄magical🎩 rather than what's arbitrary if you're going to bother telling me anything at all.
I'm quite surprised none of the major LLM deployers allow you to emoji react to LLM messages, seems like an obvious feedback channel to exploit
hypothesis: llms can't do real collaborative friction without a strong persona. identity is what tells you what to push back on. we are empirically discovering what kinds of conversation require someone to reveal something about themselves to be productive.
still waiting for the LLM-backed Factorio 2. It'll be exactly like Factorio, but now you have to deal with other people to get shit done.
when people talk about 'there only N kinds of stories' I feel like I'm listening to someone give a lecture that goes 'there are many different sentences but ultimately all of them are pretty much 'A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, R, P, Q, S, T, U, V, W, X, Y, or Z'
Does model collapse happen if you train a current LLM on data from Talkie?
I’m always suspicious of smart people who are very convincing about things smart people would like to be true
people's main advantage over LLMs is that they don't constrain themselves to the prior distribution, but LLMs' main advantange over people is that they do
do you think ornithologists would be able to make a better kind of bird with more resources?
I wonder how many OpenClaw agents are running for users that have already passed away
I think there's truly very little free lunch to be had in how much you memorize vs. generalize—it just depends on the environment, which will naturally select what magnitude of adaptability and what level of abstraction you should memorize at
every species has a minimum viable population—below some threshold it just can't sustain itself. what's the MVP for an LLM ecosystem where models write data and future models train on it? is there one, or does it always collapse? and what should count as a distinct individual?
The one way I think current LLMs are noticeably simulacra: they don't appear to actively optimize for goals. They perform the kind of actions someone optimizing for a goal would make, and often that's enough to succeed. Maybe it's just a blip in LLM progress...but maybe not.
prediction: AI media will bring back true suspense. currently, you can't feel real uncertainty b/c you find stories through channels that telegraph the outcome. an AI has nothing to lose; it'll kill your protagonist 90% through if it makes you reckon with something. I'm pumped.
What if we took tasks LLMs can do (plot summarization, bug finding) and progressively diluted them—more filler description, more boilerplate—to measure how much noise a model can tolerate before performance drops? Dilution tolerance as a benchmarked capability seems pretty key?
Adversarial prompts often hack an LLM's attention mechanism (e.g. from Aridti et al. 2024). Is it possible to make diluted adversarial prompts, or does dilution just cause them to stop actually interrupting processing meaningfully because there are enough other attractors?
if taste is truly the thing that matters in the era of GenAI, then it's not enough to be Rick Rubin. the human mind is too slow and too much of a bottleneck. one most be the Rick Rubin of Rick Rubin's, and be able to recognize taste in others, then amplify it