Great news! Our workshop what-llms-can-not-do.org is going to COLING 2027 in Macau! Hope to see you there!
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
what-llms-can-not-do.org
Lukas Edman
@lukasnlp.bsky.social
Post-Doc at TU Munich, doing cool work and that's all you need to know
Great news! Our workshop what-llms-can-not-do.org is going to COLING 2027 in Macau! Hope to see you there!
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
what-llms-can-not-do.org
Ever feel like it's too hard to keep track of what LLMs cannot do as well as humans? We're making your life easier over at: what-llms-can-not-do.github.io We're compiling a list of papers testing the abilities of LLMs against humans. Check it out! And you can help contribute too!
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
what-llms-can-not-do.github.io
Happy to announce two papers at #ACL2025! 1. We made an extension to our CUTE benchmark, EXECUTE! We find that LLMs actually *do* have the ability to do character-level manipulation, but language bias, and tokenization too, get in the way. openreview.net/forum?id=m95... 1/2
EXECUTE: A Multilingual Benchmark for LLM Token Understanding
The CUTE benchmark showed that LLMs struggle with character understanding in English. We extend it to more languages with diverse scripts and writing systems, introducing EXECUTE. Our simplified...
openreview.net