Greg Durrett

@gregdnlp.bsky.social

CS professor at NYU. Large language models and NLP. he/him

Research highlight! CosmicAI Researchers Wenxuan Ding (NYU), @gregdnlp.bsky.social (NYU) as well as external collaborator Nicholas Tomlin (NYU, TTIC) investigated whether LLM agents like Claude Code and OpenAI Codex can navigate cost-benefit tradeoffs in their actions. youtube.com/shorts/GMZ5z...

Can LLM agents like Claude Code & OpenAI Codex can navigate cost-benefit tradeoffs in their actions?

YouTube video by NSF-Simons AI Institute for Cosmic Origins

youtube.com

Check out Ramya's work on analyzing why and how LLM-generated stories feel homogeneous: the setting you prompt with might be novel but the plot unfolds in a very conventional way. Thread for how we quantified this & compare to existing metrics: 👇

Ramya Namuduri@ramyanamuduri.bsky.social · 2mo ago

Are LLM-generated stories novel? They can have unique characters and cliché plots, or the other way around. A holistic score doesn’t help distinguish the two 😔. Meet GENIE 🧞 – a fine-grained novelty metric that tells you where and why a response is original!

❗The full paper submission deadline for COLM is ~14 hours from now (11:59pm AOE)! Please submit your final PDFs on the same page where you uploaded your abstracts. And please use the provided LaTeX templates; do not handwrite your manuscript like this llama is! Good luck!

A llama sweating while writing a paper at a desk. A sign says "Deadline! March 31 11:59pm AOE"

Check out Manya's benchmark for LLM creativity! Inspired by work on creativity in graphs (@adtraghunathan.bsky.social 's "roll the dice" paper), CREATE isolates testing of creative insights for discovery. Future: understanding how LLMs derive insights & how they can be better creative partners!

Manya Wadhwa@manyawadhwa.bsky.social · 5mo ago

⚛️ Introducing CREATE, a benchmark for creative associative reasoning in LLMs. Making novel, meaningful connections is key for scientific & creative works. We objectively measure how well LLMs can do this. 🧵👇

Check out Wenxuan's work on cost-uncertainty tradeoffs in agents! Providing different info to the model (here, estimates of its uncertainty) triggers a really different type of reasoning. Agent behavior under a harness reflects what info is included, not just the instructions.

Wenxuan Ding@wenxuand.bsky.social · 5mo ago

Agents interact with environments to get information. But exploration (tools, retrieval, user interaction) is costly. Calibrate-Then-Act allows LLM agents to balance exploration and cost: 📐 Estimate uncertainty about the environment 💭 Reason about cost-uncertainty tradeoffs ⚙️ Act accordingly

Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop

📢 Postdoc position 📢 I’m recruiting a postdoc for my lab at NYU! Topics include LM reasoning, creativity, limitations of scaling, AI for science, & more! Apply by Feb 1. (Different from NYU Faculty Fellows, which are also great but less connected to my lab.) Link in 🧵

Bild

Unfortunately I won't be at #COLM2025 this week, but please check out our work being presented by my collaborators/advisors! If you are interested in evals of open-ended tasks/creativity please reach out and we can schedule a chat! :)

Greg Durrett@gregdnlp.bsky.social · 10mo ago

Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster

News🗞️ I will return to UT Austin as an Assistant Professor of Linguistics this fall, and join its vibrant community of Computational Linguists, NLPers, and Cognitive Scientists!🤘 Excited to develop ideas about linguistic and conceptual generalization (recruitment details soon!)

Picture of the UT Tower taken by me on my first day at UT as a postdoc in 2023!

Great to work on this benchmark with astronomers in our NSF-Simons CosmicAI institute! What I like about it: (1) focus on data processing & visualization, a "bite-sized" AI4Sci task (not automating all of research) (2) eval with VLM-as-a-judge (possible with strong, modern VLMs)

Sebastian Joseph@sebajoe.bsky.social · last yr.

How good are LLMs at 🔭 scientific computing and visualization 🔭? AstroVisBench tests how well LLMs implement scientific workflows in astronomy and visualize results. SOTA models like Gemini 2.5 Pro & Claude 4 Opus only match ground truth scientific utility 16% of the time. 🧵

🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]

BildBild

Check out Manya's work on evaluation for open-ended tasks! The criteria from EvalAgent can be plugged into LLM-as-a-judge or used for refinement. Great tool with a ton of potential, and there's LOTS to do here for making LLMs better at writing!

Manya Wadhwa@manyawadhwa.bsky.social · last yr.

Evaluating language model responses on open-ended tasks is hard! 🤔 We introduce EvalAgent, a framework that identifies nuanced and diverse criteria 📋✍️. EvalAgent identifies 👩‍🏫🎓 expert advice on the web that implicitly address the user’s prompt 🧵👇

Popular or not Dems cannot bend on the need for trans people to be treated with basic humanity and respect. If we give up that because the right made trans people unpopular, we give up everything. They’ll dice us group by group like a salami. We die on this hill or we die alone in a ditch

Post nicht verfügbar.