Andrea Piergentili

@apierg.bsky.social

NLP researcher PhD at the University of Trento and @fbk-mt.bsky.social, working on gender-inclusive machine translation Applied Scientist Intern at Amazon (he/him) apierg.github.io #NLP #NLProc #MT

Our #PickOfTheWeek by @bsavoldi.bsky.social: "Attention to Non-Adopters" by @kaitlynzhou.bsky.social, @gligoric.bsky.social, @myra.bsky.social, @mlam.bsky.social, @vyoma-raman.bsky.social, Boluwatife Aminu, Caeley Woo, Michael Brockman, @hannah-cha.bsky.social, @jurafsky.bsky.social (2025).

Beatrice Savoldi@bsavoldi.bsky.social · 6mo ago

🗞️ Pick of the week @fbk-mt.bsky.social: Most LLM dev centers current adopters, so what are we missing? Worth a read! 👇https://arxiv.org/pdf/2510.15951 #NLP #LLMs #HCI

Impressive work by the Glitter team: a new human-made benchmark for German gender-inclusive MT with long passages and multiple inclusive approaches + experiments showing that MT systems and LLMs still fall short in generating inclusive outputs. aclanthology.org/2025.finding...

Glitter: A Multi-Sentence, Multi-Reference Benchmark for Gender-Fair German Machine Translation

A Pranav, Janiça Hackenbuchner, Giuseppe Attanasio, Manuel Lardelli, Anne Lauscher. Findings of the Association for Computational Linguistics: EMNLP 2025. 2025.

aclanthology.org

Super interesting paper by Subramonian et al: "Agree to Disagree? A Meta-Evaluation of LLM Misgendering" arxiv.org/abs/2504.17075 Turns out, misgendering is messier than just pronouns. I'd love to see this analysis extended to grammatical gender languages! #LLM #AI #ethics @fbk-mt.bsky.social

Agree to Disagree? A Meta-Evaluation of LLM Misgendering

Numerous methods have been proposed to measure LLM misgendering, including probability-based evaluations (e.g., automatically with templatic sentences) and generation-based evaluations (e.g., with automatic heuristics or human validation). However, it has gone unexamined whether these evaluation methods have convergent validity, that is, whether their results align. Therefore, we conduct a systematic meta-evaluation of these methods across three existing datasets for LLM misgendering. We propose a method to transform each dataset to enable parallel probability- and generation-based evaluation. Then, by automatically evaluating a suite of 6 models from 3 families, we find that these methods can disagree with each other at the instance, dataset, and model levels, conflicting on 20.2% of evaluation instances. Finally, with a human evaluation of 2400 LLM generations, we show that misgendering behaviour is complex and goes far beyond pronouns, which automatic evaluations are not currently designed to capture, suggesting essential disagreement with human evaluations. Based on our findings, we provide recommendations for future evaluations of LLM misgendering. Our results are also more widely relevant, as they call into question broader methodological conventions in LLM evaluation, which often assume that different evaluation methods agree.

arxiv.org

🔍 Stiamo studiando come l'AI viene usata in Italia e per farlo abbiamo costruito un sondaggio! 👉 bit.ly/sondaggio_ai... (è anonimo, richiede ~10 minuti, e se partecipi o lo fai girare ci aiuti un sacco🙏) Ci interessa anche raggiungere persone che non si occupano e non sono esperte di AI!

Qualtrics Survey | Qualtrics Experience Management

The most powerful, simple and trusted way to gather experience data. Start your journey to experience management and try a free account today.

bit.ly

👀 Wanted: #Italian or #Dutch native speakers to take a survey on audiovisual translation for a master thesis student: watch a short video, answer some questions, help academic research 😎 ⏩ Sharing = nice! ❤️ NL link: ugent.qualtrics.com/jfe/form/SV_... IT link: ugent.qualtrics.com/jfe/form/SV_...

a woman is standing in front of a bookshelf in a bookstore and talking about research .

ALT: a woman is standing in front of a bookshelf in a bookstore and talking about research .

media.tenor.com

💭Dreaming of attending #GITT2025 but need a little extra 💸 boost? 📣 Bursary applications to support participation are now open at tinyurl.com/gitt25 📆 Deadline May 9th 🙏Thanks to our incredible sponsors DCA at Tilburg University tinyurl.com/tudca25 and FLW at Ghent University www.ugent.be/lw/en

a man in a suit is making a funny face with the words dreams are expensive behind him

ALT: a man in a suit is making a funny face with the words dreams are expensive behind him

media.tenor.com

BREAKING NEWS: CDC orders mass retraction and revision of submitted research across all science and medicine journals. Banned terms must be scrubbed. Goes beyond MMWR +other CDC pubs. Applies to research already submitted to top medical journals. Take a look. open.substack.com/pub/insideme...

BREAKING NEWS: CDC orders mass retraction and revision of submitted research across all science and medicine journals. Banned terms must be scrubbed.

Any unpublished manuscript mentioning certain topics, including gender and "LGBT," must be pulled or revised.

open.substack.com