🥳 Variation Matters is at #AMTA2026! Translating a term two different ways: an error? Humans do it on purpose. Consistency metrics say yes 😬 Our answer: variation-aware evaluation. 🗣️ Oral talk in Québec 🇨🇦, Mon 31 Sept · w/ @ziqianpeng.bsky.social, @rachelbawden.bsky.social & @yvofr.bsky.social ✨
Today, we release a corpus of 300k parallel abstracts derived from the [French national thesis archive](these.fr) and 450k parallel keywords, covering all science domains. Data & Code: github.com/ANR-MaTOS/Re.... @ziqianpeng.bsky.social @lichaozhu.bsky.social Maxime Bouthors @yvofr.bsky.social
GitHub - ANR-MaTOS/Resources
Contribute to ANR-MaTOS/Resources development by creating an account on GitHub.
github.com
🔹 Main Conference The MaTOS Pipeline for the Translation of Scientific Abstracts on the HAL Platform Panagiotis Tsolakis, Ziqian Peng, Laurent Romary, François Yvon and Rachel Bawden 📅 17th June | 15:30-15:55 | oral
🆕 Une terminologie bilingue du traitement automatique des langues (TAL) est désormais disponible sur Loterre, dans le cadre du projet ANR MaTOS ! 📖 À lire sur ISTEX 👉 www.istex.fr/une-nouvelle... #ISTEX #TAL #ScienceOuverte #MaTOS #Loterre
Very happy to be in Geneva for the #MTSummit2025. Tomorrow we present our recent progresses on the translation of scholarly documents. Check our resources and publications on the project website (anr-matos.github.io).
Merci à l'équipe organisatrice des journées #istex pour l'invitation ! Très heureux d'avoir pu présenter les avancées de MaTOS à Nancy devant un public de connaisseurs. @cnrs-inist.bsky.social
Really appreciate the feedback on this paper! It was mainly inspired by Valentin Hofmann et al. DagoBERT/"Superbizarre" papers
2: Unlike “Likely”, “Unlike” is Unlikely: always with @yvofr.bsky.social Because BPE makes a difference between tokens at the beginning and the end of words, LLMs are unable to generate prefixations
Hope you enjoyed the presentation!
1: Towards the Machine Translation of Scientific Neologisms with @yvofr.bsky.social : ever struggled to translate a new term such as pretraining or Reinforcement Learning from Human Feedback? We aim to leverage the definitions of terms to translate them more accurately