Sebastian Schuster

@sebschu.bsky.social

Computational semantics and pragmatics, interpretability and occasionally some psycholinguistics. he/him. 🦝 https://sebschu.com

🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10)

Bild

Can coding agents autonomously implement AI research extensions? We introduce RExBench, a benchmark that tests if a coding agent can implement a novel experiment based on existing research and code. Finding: Most agents we tested had a low success rate, but there is promise!

Screenshot of the RExBench preprint title page.