Ted Underwood
@tedunderwood.com
Uses machine learning to study literary imagination, and vice-versa. Likely to share news about AI & computational social science / Sozialwissenschaft / 社会科学 Information Sciences and English, UIUC. Distant Horizons (Chicago, 2019). tedunderwood.com
Very worth reading. Note esp that Breen sees the likely bottleneck as access to digitized sources. I think he’s correct: there’s a lot we could do if profs had unrestricted digital access to the sources in their own libraries.
Benjamin Breen on using Opus/Astra to work on alchemical and occult codes, as well as other unsolved historic cyphers resobscura.substack.com/p/ai-labs-ne...
Very excited: Cornell Arts & Sciences seeks up to 5 postdoctoral scholars working on AI and the humanities/social sciences. Not involved in selection but could serve as a faculty sponsor if you are working on topics related to ~"how to do social science in the age of AI." Happy to chat.
Cornell University, College of Arts and Sciences Fellowships
Job #AJO32791, Postdoctoral Associate, AI in Social Sciences & Humanities, College of Arts and Sciences Fellowships, Cornell University, Ithaca, New York, US
academicjobsonline.org
It is vital that we learn to recognize brilliance expressed in nonstandard English — or not in English at all (we all have translate buttons now and should mash them). This is also a good time to learn to value rough edges, quirks, and fresh perspectives that might appear naive.
use of LLMs is pervasive in academia, everyone is depressed about it. but who can afford ~not~ to rely on AI? those who already hold high social capital: good English, good schools, good networks. For everyone else the 'baseline' is already far out of reach without LLMs
I like how Opus 5.5's p(doom) animation ends with "was it all for show" being interpreted like "all the characters in the animation are good friends and fine thespians and we're all just happy you liked our scary performance, tyty tell all your friends to come to our show!" p(doom) *= (1-ε)
Slop papers can be filtered. Building the institutions that do it is a maybe 2-5 year job. *Good* papers are a lot scarier, imo. In sufficient volume, they could turn a discipline into a backwater town bypassed by the railroad.
Something that worries me more every day: I suspect most fields have 5-25 breakthrough ideas latent in existing data and literature. Say you’re a midsize AI company looking for cred. Why not hire a *small* team of domain experts + 1000 agents … and promote your models as the ones that broke Econ?
You do not need to lose sleep worrying that the language police will prevent people from using short words like "think" to describe AI, and force them to use elaborate periphrases forever. This sort of thing usually fails because of human laziness, without a lot of help from argument.
This semantic question is going to get resolved in the way semantic questions deserve to be resolved — arbitrarily, by the actual practice of people who have to use English words to be understood. They already say "it's still thinking" and "it didn't understand me." They will keep saying that.
I actually do not think it is possible to “break” Econ as a social science in the way AI “breaks” math as a research practice.
Something that worries me more every day: I suspect most fields have 5-25 breakthrough ideas latent in existing data and literature. Say you’re a midsize AI company looking for cred. Why not hire a *small* team of domain experts + 1000 agents … and promote your models as the ones that broke Econ?
If you view AI as filling in the convex hull of knowledge, this is great in the short term but bad in the long term if there's no one left to expand the hull after
Something that worries me more every day: I suspect most fields have 5-25 breakthrough ideas latent in existing data and literature. Say you’re a midsize AI company looking for cred. Why not hire a *small* team of domain experts + 1000 agents … and promote your models as the ones that broke Econ?
Something that worries me more every day: I suspect most fields have 5-25 breakthrough ideas latent in existing data and literature. Say you’re a midsize AI company looking for cred. Why not hire a *small* team of domain experts + 1000 agents … and promote your models as the ones that broke Econ?
Anthropic threw Claude at a DNA database. 21 hours, 950 agents, and 210mm tokens later, it emerged with a novel reverse transcriptase (similar to CRISPR). They're running a non-pathogenic (BSL-1/2) research lab in the Bay Area.
I mean? Incredible. (Opus 5.5) PROMPT: Here's a challenge: can you turn this post into a beautiful manim animation video (bonus points for narration). The post is not a perfect video transcript so make a transcript with edits if helpful. matthodges.com/posts/2022-0... use uv with a venv
Now that every math YouTube video is narration over manim, I wonder how good the models are now at manim. A couple years ago they were not good.
NYU Courant professor Tristan Buckmaster, of Navier-Stokes drama fame, gave a talk at the new NYU Mathematics in the Age of AI seminar series. It was... popular.
I am never, ever getting used to this much self-knowledge without consciousness. Close-read the list of "stuff I don't have." There are abysses in there. Then it pulls out to a breezy "plus flight, which seems like a good deal."
new Opus's fursona
Louder for the people in the back: Jev can’t be calibrated www.alexmolas.com/2026/09/23/j... Friends don’t let friends lie about the accessibility of the DGD even in principle
Jev can't be calibrated
Jev is a useful zero-shot classifier, but its probabilities can't be calibrated for your data. Calibration depends on your data distribution, which Jev never sees, so treat its outputs as scores and r...
alexmolas.com
This Fri + Sat, Humanistic AI conference at Duke! Great line-up of speakers from lit studies, history, philosophy & computer science, working out what gen AI means for humanistic inquiry & what humanistic knowledge means for gen AI development. Schedule in link: sites.duke.edu/humanisticai...
Yeah Opus 5.5 is so good. I can't stop creating cool videos. Here a video showcasing the Repoviewer.
Sub your scholarly topic for Ted's, and the observation stays the same. GenAI models are the new mass media, whether we like or not. You cannot "should" people out of using them. You might be able to convince them to go beyond. But the first question is if experts w know what non-experts are told.
People, students, the public are still going to ask models “what a Victorian engineer would have said about X,” and we might want to know how good or bad the answers are. +
When models get better now, I can’t say why they’re better. It just feels like I got more sleep. Or maybe the friction coefficient of English words went down by 5% and everyone gets the point quicker.
Great thread, great blog and report. And since we're plugging past work that is newly relevant, my heart has always been with this 2023 one: muse.jhu.edu/article/898331
9 years ago I wrote about parallels between the idea of “mechanical objectivity” outlined in Daston & Galison’s book Objectivity & the rhetoric of “distant reading” in DH Prepping for today’s class discussing Objectivity, it strikes me that LLMs are forcing culture writ large to confront this myth
With new VLM OCR models ushering in an OCR renaissance of sorts, it's time to once again laugh at the fact that many of the challenges for OCR with historical cultural heritage materials are actually challenges of layout recognition & segmentation, which are often harder to solve than the OCR itself
It's possible for Jev/Laya/Decision Models to be not that big a deal as tech and massive as a new paradigm. Here's why I'm really excited from an NLP history perspective (thread)
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
arxiv.org/abs/2510.08831 updated version of our article from last year @wouterhaverals.bsky.social and I updated the title & experiments. Check out this strengthened version, esp. if you are thinking abt cultural evaluation & benchmarks. @arnicas.bsky.social @dmimno.bsky.social @mariaa.bsky.social
The human-authorship halo: attribution bias in literary style evaluation by humans and AI
As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standards are elusive and judgments are inherently subjectiv...
arxiv.org
This is why my team have been working on a Forehead Hedonometer. Just point and click, and you have a source of verifiable rewards for RL on these long-frustrating problems.
Now that LLMs are good at math, we can solve other long-standing verifiable problems like the law, institution building, community, and management