Presenting PhantomWiki with @albertgong.bsky.social and Johann at @icmlconf.bsky.social on Tuesday 11am + an oral talk at Long Context Workshop on Saturday! Come say hi/chat about LLM reasoning and retrieval evaluation!
Anmol Kabra
@anmolkabra.com
anmolkabra.com ML PhD at @cornellbowers.bsky.social: LLM reasoning, agents, and AI for Science. Can cycle, run, juggle. Currently trying combinations.
🚨 Our paper PhantomWiki is accepted to ICML 2025 @icmlconf.bsky.social OG: bsky.app/profile/anmo... 🧑💻We designed it as a future-proof LLM reasoning benchmark. And it shows: new Qwen3-32B model with auto-thinking-mode struggles with higher difficulty questions, like DeepSeek-R1 from Jan
🚀 📢 Releasing PhantomWiki, a reasoning + retrieval benchmark for LLM agents! If I asked you "Who is the friend of father of mother of Tom?", you'd simply look up Tom -> mother -> father -> friend and answer. 🤯 SOTA LLMs, even DeepSeek-R1, struggle with such simple reasoning!
🎉 PhantomWiki is accepted to the @iclr-conf.bsky.social DATA-FM workshop! Come chat with us in Singapore 🦁 🧠 The reasoning + retrieval benchmark comes right on the heels of new @realaaai.bsky.social presidential report: AI Reasoning and Agents research front and center!
🚀 📢 Releasing PhantomWiki, a reasoning + retrieval benchmark for LLM agents! If I asked you "Who is the friend of father of mother of Tom?", you'd simply look up Tom -> mother -> father -> friend and answer. 🤯 SOTA LLMs, even DeepSeek-R1, struggle with such simple reasoning!
🚀 📢 Releasing PhantomWiki, a reasoning + retrieval benchmark for LLM agents! If I asked you "Who is the friend of father of mother of Tom?", you'd simply look up Tom -> mother -> father -> friend and answer. 🤯 SOTA LLMs, even DeepSeek-R1, struggle with such simple reasoning!