Karsten Roth

@confusezius.bsky.social

Large Models, Multimodality, Continual Learning | ELLIS ML PhD with Oriol Vinyals & Zeynep Akata | Previously Google DeepMind, Meta AI, AWS, Vector, MILA 🔗 karroth.com

Can we enhance the performance of T2I models without any fine-tuning? We show that with our ReNO, Reward-based Noise Optimization, one-step models consistently surpass the performance of all current open-source Text-to-Image models within the computational budget of 20-50 sec! #NeurIPS2024

Bild

How far can you push model merging over time, as more experts and options to model-merge arise? We comprehensively and systematically investigate this in our new work, check it out!

Sebastian Dziadzio@dziadzio.bsky.social · 2y ago

📄 New Paper: "How to Merge Your Multimodal Models Over Time?" arxiv.org/abs/2412.06712 Model merging assumes all finetuned models are available at once. But what if they need to be created over time? We study Temporal Model Merging through the TIME framework to find out! 🧵

😵‍💫 Continually pretraining large multimodal models to keep them up-to-date all-the-time is tough, covering everything from adapters, merging, meta-scheduling to data design and more! So I'm really happy to present our large-scale study at #NeurIPS2024! Come drop by to talk about all that and more!

Bild

Read our paper: Context-Aware Multimodal Pretraining Now on ArXiv Can you turn vision-language models into strong any-shot models? Go beyond zero-shot performance in SigLixP (x for context) Read @confusezius.bsky.social thread below… And follow Karsten … a rising star!

Karsten Roth@confusezius.bsky.social · 2y ago

🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist? Turns out you can, and here is how: arxiv.org/abs/2411.15099 Really excited to this work on multimodal pretraining for my first bluesky entry! 🧵 A short and hopefully informative thread:

More than zero-shot generalization, few-shot *adaptation* is critical for many applications. We find simple changes to multimodal pretraining are sufficient to yield outsized gains on a wide range of few-shot tasks. Congratulations @confusezius.bsky.social on a very successful internship!

Karsten Roth@confusezius.bsky.social · 2y ago

🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist? Turns out you can, and here is how: arxiv.org/abs/2411.15099 Really excited to this work on multimodal pretraining for my first bluesky entry! 🧵 A short and hopefully informative thread:

We maintain strong zero-shot transfer of CLIP / SigLIP across model size and data scale, while achieving up to 4x few-shot sample efficiency and up to +16% performance gains! Fun project with @confusezius.bsky.social, @zeynepakata.bsky.social, @dimadamen.bsky.social and @olivierhenaff.bsky.social.

Karsten Roth@confusezius.bsky.social · 2y ago

🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist? Turns out you can, and here is how: arxiv.org/abs/2411.15099 Really excited to this work on multimodal pretraining for my first bluesky entry! 🧵 A short and hopefully informative thread:

🤔 Can you turn your vision-language model from a great zero-shot model into a great-at-any-shot generalist? Turns out you can, and here is how: arxiv.org/abs/2411.15099 Really excited to this work on multimodal pretraining for my first bluesky entry! 🧵 A short and hopefully informative thread:

BildBildBild