Voice "cloning" is style transfer. Across three widely used systems — ElevenLabs V3, Coqui-XTTS, Chatterbox — clones don't just copy speakers, they reshape them to be warmer, more authoritative, more native English-like, and even more “humanlike”.
Martijn Bartelds
@mbartelds.bsky.social
Researcher at Together AI | Formerly at Stanford NLP, University of Groningen, TU Delft and UPenn
🌟 Next chapter! 🌟 After an incredible time at @stanfordnlp.bsky.social (huge thanks to my advisor @jurafsky.bsky.social), I've joined Together AI! Working with @jameszou.bsky.social and a world-class team to solve the hardest problems in speech and multimodal AI. Let's build! 🦾🎉
Accepted at #ICLR2026! 🎉🇧🇷 Deep learning models often fail on specific subgroups. Group DRO was designed to help, but fails when group losses aren't comparable. This is common in speech. We introduce CTC-DRO: up to 47.1% lower worst-language errors in multilingual ASR 👇
🎙️ Speech recognition is great - if you speak the right language. Our new @stanfordnlp.bsky.social paper introduces CTC-DRO, a training method that reduces worst-language errors by up to 47.1%. Work w/ Ananjan, Moussa, @jurafsky.bsky.social, Tatsu Hashimoto and Karen Livescu. Here’s how it works 🧵
✨Meet OLMoASR✨ By pairing our curated 1M-hour dataset with a powerful architecture, we've built open ASR models that achieve competitive performance with models like Whisper. We're open-sourcing data, code and models to help the community build more robust and transparent ASR.
🎙️ Say hello to OLMoASR—our fully open, from-scratch speech-to-text (STT) model. Trained on a curated audio-text set, it boosts zero-shot ASR and now powers STT in the Ai2 Playground. 👇
Now that school is starting for lots of folks, it's time for a new release of Speech and Language Processing! Jim and I added all sorts of material for the August 2025 release! With slides to match! Check it out here: web.stanford.edu/~jurafsky/sl...
Speech and Language Processing
Speech and Language Processing
web.stanford.edu
Big THANK YOU to the amazing #Interspeech2025 Organizing Committee! 💙 🎤 Odette Scharenborg, Catharine Oertel, Khiet Truong 💰 Martijn Bartelds 🌐 Dragoș Bălan 🗂️ Saskia Peters 🤝 Ginny Ruiter, Marie Louise Verhagen, Natascha Voskuijl
🎙️ Speech recognition is great - if you speak the right language. Our new @stanfordnlp.bsky.social paper introduces CTC-DRO, a training method that reduces worst-language errors by up to 47.1%. Work w/ Ananjan, Moussa, @jurafsky.bsky.social, Tatsu Hashimoto and Karen Livescu. Here’s how it works 🧵
I am excited to announce that I will join the University of Zurich as an assistant professor in August this year! I am looking for PhD students and postdocs starting from the fall. My research interests include optimization, federated learning, machine learning, privacy, and unlearning.
📢 Join us for the Conversational AI Reading Group meeting on Thursday, January 16th, 11 AM-12 PM EST. Martijn Bartelds will present "Improving Universal Access to Modern Speech Technology". Details here: poonehmousavi.github.io/rg
Happy New Year everyone! Jim and I just put up our January 2025 release of Speech and Language Processing! Check it out here: web.stanford.edu/~jurafsky/sl...
Speech and Language Processing
Speech and Language Processing
web.stanford.edu
Natural Language Processing—artificial intelligence that uses human language—has been on a roll lately. You’ve probably noticed! So the Stanford NLP Group has been growing, and diversifying into lots of new topics, including agents, language model programs, and socially aware #NLP. nlp.stanford.edu
Excited to announce the launch of our ML-SUPERB 2.0 challenge @interspeech.bsky.social 2025! Join us in pushing the boundaries of multilingual ASR and LID! 🚀 💻 multilingual.superbbenchmark.org
SUPERB: Speech processing Universal PERformance Benchmark
A comprehensive and reproducible benchmark for Self-supervised Speech Representation Learning
multilingual.superbbenchmark.org
We are excited to announce the launch of ML SUPERB 2.0 (multilingual.superbbenchmark.org) as part of the Interspeech 2024 official challenge! We hope this upgraded version of ML SUPERB advances universal access to speech processing worldwide. Please join it! #Interspeech2025
Hi speech people, super exciting news here! We are running another "Multimodal information based speech (MISP)" Challenge at @interspeech.bsky.social Participate! Spread the word! More info 👇 mispchallenge.github.io/mispchalleng...
Multimodal Information Based Speech Processing (MISP) 2025 Challenge
mispchallenge.github.io
I've started putting together a starter pack with people working on Speech Technology and Speech Science: go.bsky.app/BQ7mbkA (Self-)nominations welcome!
I wanted to contribute to "Starter Pack Season" with one for Stanford NLP+HCI: go.bsky.app/VZBhuJ5 Here are some other great starter packs: - CSS: go.bsky.app/GoEyD7d + go.bsky.app/CYmRvcK - NLP: go.bsky.app/SngwGeS + go.bsky.app/JgneRQk - HCI: go.bsky.app/p3TLwt - Women in AI: go.bsky.app/LaGDpqg