Alexandre Défossez

@honualx.bsky.social

Chief Exploration Officer @kyutai-labs.bsky.social in Paris.

We just released unmute.sh 🔇🔊 It is a text LLM wrapper, based on in-house streaming ASR, TTS, semantic VAD to reduce latency. ⏱️ Unlike Moshi 🟢, Unmute 🔊 is turn base, but allows customization in two clicks🖱️: voice and prompt! Paper and open source coming soon.

Just back from holidays, so a bit late, to announce MoshiVis, extending Moshi's multimodal capabilities to take in images 📷. Only 200M weights were added to plug a ViT through cross attention with gating 🖼️🔀🎤 Training relies on a mix of text only and text+audio synthetic data (~20k hours) 💽

Kyutai@kyutai-labs.bsky.social · 2y ago

Meet MoshiVis🎙️🖼️, the first open-source real-time speech model that can talk about images! It sees, understands, and talks about images — naturally, and out loud. This opens up new applications, from audio description for the visual impaired to visual access to information.

I'll start my presentation in 10 minutes, you can join in Zoom: concordia-ca.zoom.us/j/81541793947 See you there!

Alexandre Défossez@honualx.bsky.social · 2y ago

I'll present a dive into Moshi 🟢 and our translation model Hibiki 🇫🇷♻️🇬🇧 as part of the next @convai-rg.bsky.social reading group 👨‍🏫📗. 📅 13th of March 🕰️ 11am ET, 4pm in Paris. I'll discuss Mimi 🗜️ and multi-stream audio modeling 🔊. Join on Zoom, replay on YT. ⬛ ⬛ 🟧 🟧 🟨 🟨 🟩 🟩 🟩 ⬛ ⬛ 🟧 🟧 🟨 🟨 🟩 🟩 🟩 ⬛ ⬛

I'll present a dive into Moshi 🟢 and our translation model Hibiki 🇫🇷♻️🇬🇧 as part of the next @convai-rg.bsky.social reading group 👨‍🏫📗. 📅 13th of March 🕰️ 11am ET, 4pm in Paris. I'll discuss Mimi 🗜️ and multi-stream audio modeling 🔊. Join on Zoom, replay on YT. ⬛ ⬛ 🟧 🟧 🟨 🟨 🟩 🟩 🟩 ⬛ ⬛ 🟧 🟧 🟨 🟨 🟩 🟩 🟩 ⬛ ⬛

@convai-rg.bsky.social · 2y ago

📢 Join our Conversational AI Reading Group! 📅 Thursday, March 13 | 11 AM - 12 PM EST 🎙Speaker: Alexandre Defossez 📖 Topic: "Moshi: a speech-text foundation model for real-time dialogue" 🔗 Details: (poonehmousavi.github.io/rg) ▶️ Missed a session? Watch on YouTube: (www.youtube.com/@CONVAI_RG) 🚀

We just released the Helium-1 model , a 2B multi-lingual LLM which @exgrv.bsky.social and @lmazare.bsky.social have been crafting for us! Best model so far under 2.17B params on multi-lingual benchmarks 🇬🇧🇮🇹🇪🇸🇵🇹🇫🇷🇩🇪 On HF, under CC-BY licence: huggingface.co/kyutai/heliu...

Bild
Kyutai@kyutai-labs.bsky.social · 2y ago

Meet Helium-1 preview, our 2B multi-lingual LLM, targeting edge and mobile devices, released under a CC-BY license. Start building with it today! huggingface.co/kyutai/heliu...