Verona Teo

@veronateo.bsky.social

Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff? We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.

Excited to share our paper “Evaluating LLM Agent Collusion in Double Auctions”! We put LLMs in a simulated market and find that collusion increases when they are able to communicate via natural language, differs across models, and is influenced by urgency and oversight. 1/