Jekaterina Novikova

@j-novikova-nlp.bsky.social

Principal AI research scientist @Vanguard_Group | Founder&Host @WiAIR_podcast | Evaluation and reliability of foundation models | own opinions only 🇨🇦🇪🇺🏳️‍🌈

SUPER thrilled that our #NAACL2025 paper got the runnerup BEST paper award 😍😍🏆🏆🏆🚀🚀 We show that people rely 30% more on LLMs when they use emphatic expressions (eg "Sure, happy to help") even though the answer is wrong and 10% more when the task involves math questions 😵 📜 arxiv.org/pdf/2407.07950

arxiv.org

Kaitlyn Zhou@kaitlynzhou.bsky.social · last yr.

Thrilled that our paper won 🏆 Best Paper Runner-Up 🏆 at #NAACL25!! Our work (REL-A.I.) introduces an evaluation framework that measures human reliance on LLMs and reveals how contextual features like anthropomorphism, subject, and user history can significantly influence user reliance behaviors.

🚀 Our new episode is LIVE! 🎙️ In Episode 3, we talk with @aparnabee.bsky.social about: 🏥⚠️ Unique challenges of applying AI in medical contexts 📊🧑🏽‍🤝‍🧑🏻 Data quality and bias 👩‍⚕️🩺 Importance of collaboration with clinicians Watch and subscribe! youtu.be/DEdJltlFg4I #MLforHealth #WiAIR #WomenInAI

Responsible AI for Health, with Aparna Balagopalan

YouTube video by Women in AI Research WiAIR

youtu.be

The latest happenings in open models - Eagerly awaiting Qwen 3 - Llama 4 uptake is slow - Reasoning models seem to be saturating - Multimodal models are being slept on - China is still dominating - Oh yeah, and a reminder that my RLHF book online version0 is done! Artifacts Log #9. buff.ly/F6lapGF

The latest open artifacts (#9): RLHF book draft, where the open reasoning race is going, and unsung heroes of open LM work

Artifacts Log 9.

interconnects.ai