David Bamman

@dbamman.bsky.social

Associate Professor, School of Information, UC Berkeley. NLP, computational social science, digital humanities.

Thrilled to share that I am joining UC Berkeley as an Assistant Professor in the School of Information! I start in Fall 2027, and I am recruiting PhD students this cycle. List me in your application if you want to work with me! More on what I'm looking for (and a form to indicate interest) 🧵

Image of the UC Berkeley campus

What’s in, or what isn’t in, a close-up? Kinolab is working with computer scientist @dbamman.bsky.social (UC Berkeley, School of Information) and other collaborators this year to optimize analytical AI for film and media research, bridging the gap between distant and close viewing practices. 📉🎥👀

This has been a desideratum since the day we first published this model -- I'm so happy Patrick figured this out!

Patrick J. Burns@diyclassics.bsky.social · 5mo ago

So many of you have asked—happy to announce... Latin BERT v1 now available on HuggingFace 📦 Model: huggingface.co/latincy/lati... 📝 Preprint: arxiv.org/abs/2009.10053 Original weights, experimental repackaging—leave issues/etc. in the HF discussions #nlproc #digiclass cc: @dbamman.bsky.social

Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵

Title, author list, and two figures from the paper. 
Title: The Aftermath of DrawEduMath: Vision Language Models
Underperform with Struggling Students and Misdiagnose Errors
Authors: Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo
Figure 1: On the left is a math problem, where students are asked to draw x < 5/2 on a number line. The right side shows two example student responses that differ in correctness. DrawEduMath pairs each math problem with one student response, and prompts VLMs to answer questions about the student response.
Figure 2: VLMs consistently perform worse on answering DrawEduMath benchmark questions pertaining to erroneous student responses. Performance on non-erroneous student responses is labeled with specific VLMs’ names; that same model’s performance on erroneous student responses is directly below.

This program brought together such a wonderful group of ~30 PhD students for two weeks at Berkeley last year with backgrounds in sociology, information, law, computer science and more — same idea but in Oxford next summer; apply by Feb 2 if you’re looking to make connections across fields

UC Berkeley School of Information@berkeleyischool.bsky.social · 8mo ago

📢 Ph.D. STUDENTS: The app is now open to join us for the Oxford-Berkeley Summer Doctoral Program, an opportunity for students to learn from & engage with leading academics in the field of information and internet studies! #AcademicSky Deadline ⏰: 2/2 #OIISDP #OIIBerkeleySDPhttps://bit.ly/4p4CK2n

Berkeley wrote up our work on measuring the stories in contemporary songs and had me reflect on the origin of that work in Bruce Springsteen’s “Thunder Road”

UC Berkeley School of Information@berkeleyischool.bsky.social · 8mo ago

A new study by @dbamman.bsky.social created a machine learning algorithm that can identify narrative storytelling elements in song lyrics. 🎵 📖 "There’s been less work on measuring narrativity or even operationalizing it within songs," said Bamman.

why intern at Ai2? 🐟interns own major parts of our model development, sometimes even leading whole projects 🐡we're committed to open science & actively help our interns publish their work reach out if u wanna build open language models together 🤝 links 👇

Bild

Unfortunately, I expect that the fact they won a pretty resounding fair use victory is going to be lost in much of the coverage. They lost on downloading & storing the LibGen dataset (straight infringement!), but the act of training on copyrighted material (+ making their own ebooks) was a win.

Will Oremus@willoremus.com · 11mo ago

Breaking: In landmark agreement, AI firm Anthropic will pay $1.5 billion to settle a copyright lawsuit brought by book authors and publishers. Story to follow.

One of the fun parts of policy work is that you have a handful of extremely important, esoteric, but load-bearing legal concepts (like fair use) that only ever make it into public consciousness when they're being targeted for destruction, and public opinion about them is entirely outcome-dependent.

Creative people and people of good will who are allies and advocates for creativity need to come to grips with the fact that AI training is generally fair use for the same deep reasons that critique, parody, teaching, scholarship, and other daily creative practices are fair use.

Go to WI and work with Lucy! And stay for the farmer's market, memorial union, orpheum, willy st coop, chocolate shoppe, american players theater, and the million other things that make Madison an amazing place to live. PhD app deadline in December.

Lucy Li@lucy3.bsky.social · last yr.

I'm sadly not at #IC2S2 😭, but I will be at #ACL2025 in Vienna ☕️ next week!! Please spread the word that I'm recruiting prospective PhD students: lucy3.notion.site/for-prospect...

I’m a week late with my #thanksbrett but just want to add my appreciation too — in thinking of all the careers that ODH has given space to and supported (including my own), words can’t suffice. @brettbobley.bsky.social maybe at least it’s now permitted to let us buy you a beer.

Jennifer Serventi@jenserventi.bsky.social · last yr.

*sigh* It's too soon, but on the day of his retirement from public service, I think that it's important to recognize the contributions of Brett Bobley to the work of the National Endowment for the Humanities and the field. #ThanksBrett for your leadership, your good cheer & your home-baked goods.