🤨"But why linguistics" is the most common question when talking about linguistic reasoning benchmarks. Last year we organized a shared task at WMT...and no one participated 🤣 🤯Let me change your mind why this is one of the most challenging, focused and best reasoning benchmarks right now.
For the launch of the 1-month open challenge that precedes the in-person Olympiad, @juliakreutzer.bsky.social is hosting @danmirea.bsky.social & Eduardo Sanchez to discuss the appeal of these problems, solution challenges, and potential insights from AI-human evaluation. Join us: luma.com/jk8uv7zs