Excited to be presenting our new paper at #CogSci2026 this coming week, which asks: Does Contextual Informativeness Predict Preschoolers’ Word Learning from Stories? I will be in poster session 3 on Friday evening. Full paper will be published in proceedings at the conclusion of the conference! :)
Maria Valentini
@mvalentini.bsky.social
computer science/cognitive science PhD student @ CU Boulder • computational psycholinguistics, NLP for education, AI ethics
Congratulations to @mginn.bsky.social, @covetedfish.bsky.social, Ali Marashian, @mvalentini.bsky.social, @alexispalmer.bsky.social & friends for receiving an Outstanding Paper award at #ACL2026 for the paper "Massively Multilingual Joint Segmentation and Glossing". aclanthology.org/2026.acl-lon...
Massively Multilingual Joint Segmentation and Glossing
Michael Ginn, Lindia Tjuatja, Enora Rice, Ali Marashian, Maria Valentini, Jasmine Xu, Graham Neubig, Alexis Palmer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguist...
aclanthology.org
Excited to announce that the PolyGloss paper has been accepted to @aclmeeting.bsky.social! Previously, we trained models to help in endangered language documentation workflows by automatically predicting interlinear glosses. But real-world user studies revealed crucial issues...
This is the third story I've read in a month about how AI chatbots are leading people into psychological crises. Gift link
They Asked an A.I. Chatbot Questions. The Answers Sent Them Spiraling.
nytimes.com
I don’t really have the energy for politics right now. So I will observe without comment: Executive Order 14110 was revoked (Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence)
Excited to be presenting my work with @teaywright.bsky.social at #COLING2025 next week in Abu Dhabi! Find us in poster session 6/E on Jan 22nd (11 AM in the atrium). Paper: arxiv.org/abs/2412.17427
Measuring Contextual Informativeness in Child-Directed Text
To address an important gap in creating children's stories for vocabulary enrichment, we investigate the automatic evaluation of how well stories convey the semantics of target vocabulary words, a tas...
arxiv.org
1. Can you stop companies from training generative AI using your data? No, not currently. 2. Is this dataset meant for training generative AI? 🤷♀️ but more likely for research and statistical analysis. 3. Is it ok to duplicate and distribute people’s data without agency to opt out? I’d argue no.
First dataset for the new @huggingface.bsky.social @bsky.app community organisation: one-million-bluesky-posts 🦋 📊 1M public posts from Bluesky's firehose API 🔍 Includes text, metadata, and language predictions 🔬 Perfect to experiment with using ML for Bluesky 🤗 huggingface.co/datasets/blu...
So many people, CS researchers included, think that you can explore how an LLM works by simply asking it to tell you what it is doing or "thinking". Here @jennhu.bsky.social provides an excellent illustration of how that approach fails even at the most basic level.
To researchers doing LLM evaluation: prompting is *not a substitute* for direct probability measurements. Check out the camera-ready version of our work, to appear at EMNLP 2023! (w/ @rplevy.bsky.social) Paper: arxiv.org/abs/2305.13264 Original thread: twitter.com/_jennhu/stat...