LLMs can retrieve knowledge — but can they connect it in *creative* ways to solve problems? Introducing CresOWLve 🦉, a new benchmark that evaluates creative problem-solving over real-world knowledge, using puzzles that require multiple creative thinking strategies.👇
💡Can we optimize LLMs to be more creative? Introducing Creative Preference Optimization (CrPO) and MuCE (Multi-task Creativity Evaluation Dataset). Result: More novel, diverse, surprising text—without losing quality! 📝 Appearing at #EMNLP2025
I am attending @naaclmeeting.bsky.social this week to present our paper. Come check out our poster at 14:00, Apr 30 in Hall 3 . @defnecirci.bsky.social and Hale Sirin will also be there to answer your questions!
Are LLMs linguistically productive and systematic in morphologically-rich languages as good as humans? No 🤨 Our new NAACL 2025 paper (arxiv.org/abs/2410.12656) reveals a significant performance gap between LLMs and humans in linguistic creativity and morphological generalization.
Lots of great news out of the EPFL NLP lab these last few weeks. We'll be at @iclr-conf.bsky.social and @naaclmeeting.bsky.social in April / May to present some of our work in training dynamics, model representations, reasoning, and AI democratization. Come chat with us during the conference!
Are LLMs linguistically productive and systematic in morphologically-rich languages as good as humans? No 🤨 Our new NAACL 2025 paper (arxiv.org/abs/2410.12656) reveals a significant performance gap between LLMs and humans in linguistic creativity and morphological generalization.
Evaluating Morphological Compositional Generalization in Large Language Models
Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks. However, their linguistic generalization capabilities remain questio...
arxiv.org