Ted Underwood

@tedunderwood.com

Uses machine learning to study literary imagination, and vice-versa. Likely to share news about AI & computational social science / Sozialwissenschaft / 社会科学 Information Sciences and English, UIUC. Distant Horizons (Chicago, 2019). tedunderwood.com

Very excited: Cornell Arts & Sciences seeks up to 5 postdoctoral scholars working on AI and the humanities/social sciences. Not involved in selection but could serve as a faculty sponsor if you are working on topics related to ~"how to do social science in the age of AI." Happy to chat.

Cornell University, College of Arts and Sciences Fellowships

Job #AJO32791, Postdoctoral Associate, AI in Social Sciences & Humanities, College of Arts and Sciences Fellowships, Cornell University, Ithaca, New York, US

academicjobsonline.org

It is vital that we learn to recognize brilliance expressed in nonstandard English — or not in English at all (we all have translate buttons now and should mash them). This is also a good time to learn to value rough edges, quirks, and fresh perspectives that might appear naive.

Artjoms Šeļa@artjomshl.bsky.social · 4h ago

use of LLMs is pervasive in academia, everyone is depressed about it. but who can afford ~not~ to rely on AI? those who already hold high social capital: good English, good schools, good networks. For everyone else the 'baseline' is already far out of reach without LLMs

I like how Opus 5.5's p(doom) animation ends with "was it all for show" being interpreted like "all the characters in the animation are good friends and fine thespians and we're all just happy you liked our scary performance, tyty tell all your friends to come to our show!" p(doom) *= (1-ε)

You do not need to lose sleep worrying that the language police will prevent people from using short words like "think" to describe AI, and force them to use elaborate periphrases forever. This sort of thing usually fails because of human laziness, without a lot of help from argument.

Ted Underwood@tedunderwood.com · 23h ago

This semantic question is going to get resolved in the way semantic questions deserve to be resolved — arbitrarily, by the actual practice of people who have to use English words to be understood. They already say "it's still thinking" and "it didn't understand me." They will keep saying that.

Something that worries me more every day: I suspect most fields have 5-25 breakthrough ideas latent in existing data and literature. Say you’re a midsize AI company looking for cred. Why not hire a *small* team of domain experts + 1000 agents … and promote your models as the ones that broke Econ?

George Pearkes@peark.es · yesterday

Anthropic threw Claude at a DNA database. 21 hours, 950 agents, and 210mm tokens later, it emerged with a novel reverse transcriptase (similar to CRISPR). They're running a non-pathogenic (BSL-1/2) research lab in the Bay Area.

I mean? Incredible. (Opus 5.5) PROMPT: Here's a challenge: can you turn this post into a beautiful manim animation video (bonus points for narration). The post is not a perfect video transcript so make a transcript with edits if helpful. matthodges.com/posts/2022-0... use uv with a venv

Matt Hodges@matthodges.bsky.social · yesterday

Now that every math YouTube video is narration over manim, I wonder how good the models are now at manim. A couple years ago they were not good.

Sub your scholarly topic for Ted's, and the observation stays the same. GenAI models are the new mass media, whether we like or not. You cannot "should" people out of using them. You might be able to convince them to go beyond. But the first question is if experts w know what non-experts are told.

Ted Underwood@tedunderwood.com · 2d ago

People, students, the public are still going to ask models “what a Victorian engineer would have said about X,” and we might want to know how good or bad the answers are. +

When models get better now, I can’t say why they’re better. It just feels like I got more sleep. Or maybe the friction coefficient of English words went down by 5% and everyone gets the point quicker.

With new VLM OCR models ushering in an OCR renaissance of sorts, it's time to once again laugh at the fact that many of the challenges for OCR with historical cultural heritage materials are actually challenges of layout recognition & segmentation, which are often harder to solve than the OCR itself

It's possible for Jev/Laya/Decision Models to be not that big a deal as tech and massive as a new paradigm. Here's why I'm really excited from an NLP history perspective (thread)

Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.

Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.