Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff? We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
Thrilled to be sharing my latest work at #CogSci2026! Many of the spaces we move through, from kitchens to airports, were designed with specific uses in mind. How do people create such environments, and how do users figure out what they were designed for? 📃 osf.io/preprints/ps...
Excited to share our new work at #CogSci2025! We explore how people plan deceptive actions, and how detectives try to see through the ruse and infer what really happened based on the traces left behind. 🕵️♀️ Paper: osf.io/preprints/osf/vqgz5_v1 Code: github.com/cicl-stanford/recursive_deception 1/
Excited to share our paper “Evaluating LLM Agent Collusion in Double Auctions”! We put LLMs in a simulated market and find that collusion increases when they are able to communicate via natural language, differs across models, and is influenced by urgency and oversight. 1/