Excited to announce that the PolyGloss paper has been accepted to @aclmeeting.bsky.social! Previously, we trained models to help in endangered language documentation workflows by automatically predicting interlinear glosses. But real-world user studies revealed crucial issues...
Enora
@covetedfish.bsky.social
NLP/Computational Linguistics PhD @lecslab.bsky.social and @bouldernlp.bsky.social typologically robust multilingual NLP technology for language documentation https://covetedfish.github.io/
Really excited about this work w/ my long-time collaborators at Boulder! We address limitations in existing morphosyntactic annotation systems for digitally under-resourced languages and show how *jointly* predicting morphological segmentation helps with glossing performance
Excited to announce that the PolyGloss paper has been accepted to @aclmeeting.bsky.social! Previously, we trained models to help in endangered language documentation workflows by automatically predicting interlinear glosses. But real-world user studies revealed crucial issues...
nintendo is going to drone strike the data center responsible for this
“I asked ChatGPT-“ okay well I asked the big oak tree at the center of the woods and she said you’re a lazy dork
I’ll be presenting this work at #NAACL at the LM4UC workshop on Sunday. Come say hi! bsky.app/profile/arxi...
Enora Rice, Ali Marashian, Hannah Haynie, Katharina von der Wense, Alexis Palmer Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging https://arxiv.org/abs/2503.19979
Shout out CU boulder for pulling up the little native plant garden outside our apartment for no reason
I don't want to be taken seriously. I want to be taken sillily.
Every time I see this crane I wonder if someone will take it as a sign to breakup with their girlfriend
I'm looking for work on how NLP tools can be misaligned with the real and present needs of documentary linguists. Specifically, cases where the tools *seem* useful but aren't practical or miss the mark in some other way. Any ideas?
Enora Rice, Ali Marashian, Hannah Haynie, Katharina von der Wense, Alexis Palmer Untangling the Influence of Typology, Data and Model Architecture on Ranking Transfer Languages for Cross-Lingual POS Tagging https://arxiv.org/abs/2503.19979
I'm giving a lecture on language model debiasing to my undergrad NLP course on Friday but I'm not super up to date on the research. Does anyone have any suggestions for papers/topics to cover?
The AmericasNLP 2025 Shared Task on Machine Translation Metrics for Indigenous languages is live! Standard metrics like BLEU and ChrF are not ideal for evaluating Indigenous American languages. How can we do better? Due March 5th, 2025. turing.iimas.unam.mx/americasnlp/...
First Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP)
The goal of the workshop is to encourage and increase the visibility of work on indigenous languages of the Americas. It aims to encourage research on NLP, computational linguistics, corpus linguistic...
turing.iimas.unam.mx
it's weird how "performative BS" never includes "posts excoriating other people's performative BS"
I am so excited to share with you all the 2025 edition of our #AmericasNLP workshop! Do not hesitate in submitting your amazing research paper on indigenous and low resource languages. Submission deadline: March 7, 2025 turing.iimas.unam.mx/americasnlp/...
Our paper was accepted to #COLING! If you work on low-resource MT and have ever found yourself limited to bible data, you might find this interesting.
people will take any excuse to use the word nascent (I'm people)
I feel like reviewers often expect short papers to be long papers condensed into 4 pages. They should really be a venue to showcase focused and incremental work.
Why do people still use iso 639-1 codes instead of iso 639-3? Honest question