This looks awesome! We need benchmarks that consider coordination between agents. Also super cool to see cross comparison between LLM/ VLM and MARL agents! Congrats to @kale-ab.bsky.social & co. ๐
Can LLM agents coordinate in long-horizon, open-ended worlds? We test 13 LLMs in Alem, a new benchmark. Most struggle, averaging ~6% normalised return. Yet on Hard, zero-shot Gemini 3.1 Pro matches the best MARL agent after 1B training steps. Our ablations show communication matters most. ๐งต