Swarat Chaudhuri

@swarat.bsky.social

Professor of Computer Science at UT Austin and Visiting Researcher at Google Deepmind, London. Automated Reasoning + Machine Learning + Formal Methods. https://www.cs.utexas.edu/~swarat

Harvard has set an example for other higher-ed institutions - rejecting an unlawful and ham-handed attempt to stifle academic freedom, while taking steps to make sure students can benefit from an environment of intellectual inquiry, rigorous debate and mutual respect. Let’s hope others follow suit.

Upon learning that yesterday would be my last day as a program officer at the National Science Foundation, I shared this parting message with my colleagues. The next few months will be frenetic and stressful for them. Here are some things that you can do to help them with the mission ahead. (1)

Bild

DARPA released a Request for Information (RFI) that seeks community feedback on the draft DARPA Guide to Formal Methods to Deliver Resilient Systems for Proposals (“the FMDRS Guide”). You can find the RFI here on Sam.gov. Details in the image...

Bild

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation “We introduce the Formally Verified Automated Programming Progress Standards, or FVAPPS, a benchmark of 4715 samples […] including 1083 curated and quality controlled samples” arxiv.org/abs/2502.05714

Proving the Coding Interview: A Benchmark for Formally Verified Code Generation

We introduce the Formally Verified Automated Programming Progress Standards, or FVAPPS, a benchmark of 4715 samples for writing programs and proving their correctness, the largest formal verification ...

arxiv.org

Can LLMs be used to discover interpretable models of human and animal behavior?🤔 Turns out: yes! Thrilled to share our latest preprint where we used FunSearch to automatically discover symbolic cognitive models of behavior. 1/12

Bild

This is the most relevant article to NIH and research cuts I’ve seen. Imagine if this was today , how many people would be saying “Why are we studying Gila Monsters and their impact on diabetes ? That’s wasted money !” globalnews.ca/news/9793403...

How a Canadian scientist and a venomous lizard helped pave the way for Ozempic - National | Globalnews.ca

In 1984, Dr. Daniel Drucker, an endocrinologist from the University of Toronto, discovered a hormone that helped pave the way for popular diabetes drugs such as Ozempic.

globalnews.ca

@ayushkhaitan.bluesky.social, Amitayush Thakur, and I are organizing an #AI4Math panel at the Joint Mathematics Meeting this month. Please spread the word among your math friends! We will post a summary of the discussion after the event.

Ayush Khaitan@ayushkhaitan.bsky.social · 2y ago

Looking forward to the #jmm2025 panel on the "Use of AI tools for Mathematics research" that we are co-organizing with @swarat.bsky.social and Amitayush Thakur. The panelists are Alex Kontorovich, Rishi Mehta, Emily Wenger and Kaiyu Yang. See you there!

An excellent post by Kevin Buzzard on informal reasoning methods like o3. The key point, one I wholeheartedly agree with, is that informal methods continue to struggle with proof even when they give the correct answers, and this is a critical liability. xenaproject.wordpress.com/2024/12/22/c...

Can AI do maths yet? Thoughts from a mathematician.

So the big news this week is that o3, OpenAI’s new language model, got 25% on FrontierMath. Let’s start by explaining what this means.

xenaproject.wordpress.com

Delighted to share our new position paper: arxiv.org/abs/2412.16075! The o1/o3 path to math reasoning is based on LLMs and large-scale test-time search. We argue for a different path that uses formal proof assistants for ✅ creating high-quality synthetic data ✅ rigorous test-time feedback. (1/2)

"Formal mathematical reasoning: A new frontier in AI"Block diagram of a neural theorem prover

Good 🦋 samaritans: any of you have a friend at X who can get my X account back? If so, DM/email me! The thread below tells what happened. I couldn't ever reach a human at X; the bots kept saying X couldn't help. Every new account I've created since has been auto-suspended. bsky.app/profile/swar...

Swarat Chaudhuri@swarat.bsky.social · 2y ago

I was thinking about what to write in my first real 🦋 post. Didn't expect it to be so easy. My X account has been hacked. The hacker changed the account email, and X won't return access to me because I don't know what it was changed to. 🤡 Oh well, good riddance and sorry I didn't quit sooner.

Samuele created a fantastic benchmark for studying Neurosymbolic Reasoning Shortcuts! Reasoning shortcuts are a common failure mode of Neurosymbolic models. 🚀 Good diagnostic tools will really help the field forward! Glad to have contributed to this.

Samuele Bortolotti@samubortolotti.bsky.social · 2y ago

📣 Does your model learn high-quality #concepts, or does it learn a #shortcut? Test it with our #NeurIPS2024 dataset & benchmark track paper! rsbench: A Neuro-Symbolic Benchmark Suite for Concept Quality and Reasoning Shortcuts What's the deal with rsbench? 🧵