Is the needle-in-a-haystack test still meaningful given the giant green heatmaps in modern LLM papers? We create ONERULER 💍, a multilingual long-context benchmark that allows for nonexistent needles. Turns out NIAH isn't so easy after all! Our analysis across 26 languages 🧵👇
✨I am on the faculty job market in the 2024-2025 cycle!✨ My research centers on advancing Responsible AI, specifically enhancing factuality, robustness, and transparency in AI systems. If you have relevant positions, let me know! lasharavichander.github.io Please share/RT!
Abhilasha Ravichander - Home
lasharavichander.github.io
Long-form text generation with multiple stylistic and semantic constraints remains largely unexplored. We present Suri 🦙: a dataset of 20K long-form texts & LLM-generated, backtranslated instructions with complex constraints. 📎 arxiv.org/abs/2406.19371
🌊Heading to #EMNLP2024 tmr, presenting PostMark on Tue. morning! 🔗 arxiv.org/abs/2406.14517 Aside from this, I'd love to chat about: • long-context training • realistic & hard eval • synthetic data • tbh any cool projects people are working on Also, I'm on the lookout for a summer 2025 internship!