Anirudh Khatry

@anirudhkhatry.bsky.social

CS PhD @utaustin.bsky.social

How good are LLMs at 🔭 scientific computing and visualization 🔭? AstroVisBench tests how well LLMs implement scientific workflows in astronomy and visualize results. SOTA models like Gemini 2.5 Pro & Claude 4 Opus only match ground truth scientific utility 16% of the time. 🧵

BildBild

News🗞️ I will return to UT Austin as an Assistant Professor of Linguistics this fall, and join its vibrant community of Computational Linguists, NLPers, and Cognitive Scientists!🤘 Excited to develop ideas about linguistic and conceptual generalization (recruitment details soon!)

Picture of the UT Tower taken by me on my first day at UT as a postdoc in 2023!

Evaluating language model responses on open-ended tasks is hard! 🤔 We introduce EvalAgent, a framework that identifies nuanced and diverse criteria 📋✍️. EvalAgent identifies 👩‍🏫🎓 expert advice on the web that implicitly address the user’s prompt 🧵👇

BildBild

🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]

BildBild

A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏

Bild

encourage postdocs to apply 👇 @soldaini.net, myself and others from @ai2.bsky.social have been helping in project & also learning a ton---continued pretraining, creating domain-specific training data & evals---to build foundation models that scientists can use. promising area for open source LMs!

Jessy Li@jessyjli.bsky.social · last yr.

🌟Job ad🌟 We (@gregdnlp.bsky.social, @mattlease.bsky.social and I) are hiring a postdoc fellow within the CosmicAI Institute, to do galactic work with LLMs and generative AI! If you would like to push the frontiers of foundation models to help solve myths of the universe, please apply!

@ayushkhaitan.bluesky.social, Amitayush Thakur, and I are organizing an #AI4Math panel at the Joint Mathematics Meeting this month. Please spread the word among your math friends! We will post a summary of the discussion after the event.

Ayush Khaitan@ayushkhaitan.bsky.social · 2y ago

Looking forward to the #jmm2025 panel on the "Use of AI tools for Mathematics research" that we are co-organizing with @swarat.bsky.social and Amitayush Thakur. The panelists are Alex Kontorovich, Rishi Mehta, Emily Wenger and Kaiyu Yang. See you there!

We got an 🥂 Outstanding Paper Award!! Cannot be more grateful 🥹 This is super validating for our long pursuit of computational work on QUD. Congrats to the amazing @yatingwu.bsky.social, Ritika Mangla, Alex Dimakis, @gregdnlp.bsky.social

Jessy Li@jessyjli.bsky.social · 2y ago

Wednesday at #EMNLP: @yatingwu.bsky.social will present our work connecting curiosity and questions in discourse. We built strong models to predict salience, outperforming large LLMs. 👉[Oral] Discourse+Phonology+Syntax2 10:30-12:00 @ Flagler also w/ Ritika Mangla @gregdnlp.bsky.social Alex Dimakis

I won't be at EMNLP, but come and see: 🔍 Detecting factual errors from LLMs (Liyan Tang) 🛠️ Detect, critique, & refine pipeline (Manya Wadhwa and Lucy Zhao) 🏭 Synthetic data generation (Abhishek Divekar) 📄 Fact-checking (Aniruddh Sriram) at FEVER t.co/fQbl0G7m23 (1st real post in the bluer skies!)

Image of the linked website listing EMNLP paper titles, authors, and locations