📢 New blog post describing the COLM 2026 PCs’ analyses of AI use in submitted papers. gregdurrett.github.io/colm2026-blo... (temporary home, new COLM website coming soon!)
Greg Durrett
@gregdnlp.bsky.social
CS professor at NYU. Large language models and NLP. he/him
Research highlight! CosmicAI Researchers Wenxuan Ding (NYU), @gregdnlp.bsky.social (NYU) as well as external collaborator Nicholas Tomlin (NYU, TTIC) investigated whether LLM agents like Claude Code and OpenAI Codex can navigate cost-benefit tradeoffs in their actions. youtube.com/shorts/GMZ5z...
Can LLM agents like Claude Code & OpenAI Codex can navigate cost-benefit tradeoffs in their actions?
YouTube video by NSF-Simons AI Institute for Cosmic Origins
youtube.com
Check out Ramya's work on analyzing why and how LLM-generated stories feel homogeneous: the setting you prompt with might be novel but the plot unfolds in a very conventional way. Thread for how we quantified this & compare to existing metrics: 👇
Are LLM-generated stories novel? They can have unique characters and cliché plots, or the other way around. A holistic score doesn’t help distinguish the two 😔. Meet GENIE 🧞 – a fine-grained novelty metric that tells you where and why a response is original!
❗The full paper submission deadline for COLM is ~14 hours from now (11:59pm AOE)! Please submit your final PDFs on the same page where you uploaded your abstracts. And please use the provided LaTeX templates; do not handwrite your manuscript like this llama is! Good luck!
ICML reviews have you considering this? Please look at our final submission instructions for COLM below!
~45 hours until the abstract deadline! Submit abstracts on OpenReview by 3/26 11:59pm AOE, full papers 3/31. Final reminders & submission instructions for COLM are below. Note that as of the March 31 deadline, papers must not be under review for ICML or committed to ACL colmweb.org/submission-i...
Check out Manya's benchmark for LLM creativity! Inspired by work on creativity in graphs (@adtraghunathan.bsky.social 's "roll the dice" paper), CREATE isolates testing of creative insights for discovery. Future: understanding how LLMs derive insights & how they can be better creative partners!
⚛️ Introducing CREATE, a benchmark for creative associative reasoning in LLMs. Making novel, meaningful connections is key for scientific & creative works. We objectively measure how well LLMs can do this. 🧵👇
Check out Wenxuan's work on cost-uncertainty tradeoffs in agents! Providing different info to the model (here, estimates of its uncertainty) triggers a really different type of reasoning. Agent behavior under a harness reflects what info is included, not just the instructions.
Agents interact with environments to get information. But exploration (tools, retrieval, user interaction) is costly. Calibrate-Then-Act allows LLM agents to balance exploration and cost: 📐 Estimate uncertainty about the environment 💭 Reason about cost-uncertainty tradeoffs ⚙️ Act accordingly
Still accepting applications for this postdoc position in my lab at NYU! Applications due Feb 1. Please see the posting below for more information and apply on Interfolio: cims.nyu.edu/taur/postdoc... apply.interfolio.com/178940
cims.nyu.edu
📢 Postdoc position 📢 I’m recruiting a postdoc for my lab at NYU! Topics include LM reasoning, creativity, limitations of scaling, AI for science, & more! Apply by Feb 1. (Different from NYU Faculty Fellows, which are also great but less connected to my lab.) Link in 🧵
Submit to COLM! Deadline of March 31. This llama gets to enjoy his holidays and isn't stressed out just yet...
COLM 2026 is just around the corner! Mark your calendars for: 💡 Abstract deadline: Thursday, March 26, 2026 📄 Full paper submission deadline: Tuesday, March 31, 2026 Call for papers (website coming soon): docs.google.com/document/d/1...
Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop
📢 Postdoc position 📢 I’m recruiting a postdoc for my lab at NYU! Topics include LM reasoning, creativity, limitations of scaling, AI for science, & more! Apply by Feb 1. (Different from NYU Faculty Fellows, which are also great but less connected to my lab.) Link in 🧵
Two brief advertisements! TTIC is recruiting both tenure-track and research assistant professors: ttic.edu/faculty-hiri... NYU is recruiting faculty fellows: apply.interfolio.com/174686 Happy to chat with anyone considering either of these options
TTIC Faculty Opportunities at TTIC
ttic.edu
Unfortunately I won't be at #COLM2025 this week, but please check out our work being presented by my collaborators/advisors! If you are interested in evals of open-ended tasks/creativity please reach out and we can schedule a chat! :)
Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster
Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster
Excited to present this at #COLM2025 tomorrow! (Tuesday, 11:00 AM poster session)
One of the ways that LLMs can be inconsistent is the "generator-validator gap," where LLMs deem their own answers incorrect. 🎯 We demonstrate that ranking-based discriminator training can significantly reduce this gap, and improvements on one task often generalize to others! 🧵👇
Check out this feature about AstroVisBench, our upcoming NeurIPS D&B paper about code workflows and visualization in the astronomy domain! Great testbed for the interaction of code + VLM reasoning models.
Exciting news! Introducing AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy! A new benchmark developed by researchers at the NSF-Simons AI Institute for Cosmic Origins is testing how well LLMs implement scientific workflows in astronomy and visualize results.
News🗞️ I will return to UT Austin as an Assistant Professor of Linguistics this fall, and join its vibrant community of Computational Linguists, NLPers, and Cognitive Scientists!🤘 Excited to develop ideas about linguistic and conceptual generalization (recruitment details soon!)
Great to work on this benchmark with astronomers in our NSF-Simons CosmicAI institute! What I like about it: (1) focus on data processing & visualization, a "bite-sized" AI4Sci task (not automating all of research) (2) eval with VLM-as-a-judge (possible with strong, modern VLMs)
How good are LLMs at 🔭 scientific computing and visualization 🔭? AstroVisBench tests how well LLMs implement scientific workflows in astronomy and visualize results. SOTA models like Gemini 2.5 Pro & Claude 4 Opus only match ground truth scientific utility 16% of the time. 🧵
The end of US leadership in science, technology, and innovation. All in one little table. A tremendous gift to China, courtesy of the GOP. nsf-gov-resources.nsf.gov/files/00-NSF...
Super excited Marin is finally out! Come see what we've been building! Code/platform for training fully reproducible models end-to-end, from data to evals. Plus a new high quality 8B base model. Percy did a good job explaining it on the other place. marin.community x.com/percyliang/s...
Percy Liang on X: "What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3" / X
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3
x.com
Check out Anirudh's work on a new benchmark for C-to-Rust transpilation! 100 realistic-scale C projects, plus target Rust interfaces + Rust tests that let us validate the transpiled code beyond what prior benchmarks allow.
🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]
🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]
Check out Manya's work on evaluation for open-ended tasks! The criteria from EvalAgent can be plugged into LLM-as-a-judge or used for refinement. Great tool with a ton of potential, and there's LOTS to do here for making LLMs better at writing!
Evaluating language model responses on open-ended tasks is hard! 🤔 We introduce EvalAgent, a framework that identifies nuanced and diverse criteria 📋✍️. EvalAgent identifies 👩🏫🎓 expert advice on the web that implicitly address the user’s prompt 🧵👇
Check out Ramya et al.'s work on understanding discourse similarities in LLM-generated text! We see this as an important step in quantifying the "sameyness" of LLM text, which we think will be a step towards fixing it!
Have that eerie feeling of déjà vu when reading model-generated text 👀, but can’t pinpoint the specific words or phrases 👀? ✨We introduce QUDsim, to quantify discourse similarities beyond lexical, syntactic, and content overlap.
Our final South by Semantics lecture at UT Austin is happening on Wednesday April 23!
Check out @juand-r.bsky.social and @wenxuand.bsky.social 's work on improving generator-validator gaps in LLMs! I really like the formulation of the G-V gap we present, and I was pleasantly surprised by how well the ranking-based training closed the gap. Looking forward to following up in this area!
One of the ways that LLMs can be inconsistent is the "generator-validator gap," where LLMs deem their own answers incorrect. 🎯 We demonstrate that ranking-based discriminator training can significantly reduce this gap, and improvements on one task often generalize to others! 🧵👇
If you're scooping up students off the street for writing op-eds, you're secret police, and should be treated accordingly.
I'm excited to announce two papers of ours which will be presented this summer at @naaclmeeting.bsky.social eting.bsky.social and @iclr-conf.bsky.social ! 🧵
Excited about Proofwala, @amitayush.bsky.social's new framework for ML-aided theorem-proving. * Paper: arxiv.org/abs/2502.04671 * Code: github.com/trishullab/p... Proofwala allows the collection of proof-step data from multiple proof assistants (Coq and Lean) and multilingual training. (1/3)
Popular or not Dems cannot bend on the need for trans people to be treated with basic humanity and respect. If we give up that because the right made trans people unpopular, we give up everything. They’ll dice us group by group like a salami. We die on this hill or we die alone in a ditch