Pretrained ViTs usually come in rigid sizes (S, B, L, H). But your hardware constraints don't We built a way to make DINO or CLIP fully elastic in <5 mins without any retraining ⚡️ Get the exact model size you need, not just what was released Find Walter at #NeurIPS Poster 4709 | Thu 4:30-7:30 PM
Yuki Asano
@yukimasano.bsky.social
Professor at University of Technology Nuremberg Head of Fundamental AI Lab
On the occasion of the 1000th citation of our Sinkhorn-Knopp self-supervised representation learning paper, I've written a whole post about the history and the key bits of this method that powers the state-of-the-art SSL vision models. Read it here :): docs.google.com/document/d/1...
Today, we release Franca, a new vision Foundation Model that matches and often outperforms DINOv2. The data, the training code and the model weights are open-source. This is the result of a close and fun collaboration @valeoai.bsky.social (in France) and @funailab.bsky.social (in Franconia)🚀
1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.
🚀🚀PaliGemma 2 is our updated and improved PaliGemma release using the Gemma 2 models and providing new pre-trained checkpoints for the full cross product of {224px,448px,896px} resolutions and {3B,10B,28B} model sizes. 1/7
Pls RT Permanent Assistant Professor (Lecturer) position in Computer Vision @bristoluni.bsky.social [DL 6 Jan 2025] This is a research+teaching permanent post within MaVi group uob-mavi.github.io in Computer Science. Suitable for strong postdocs or exceptional PhD graduates. t.co/k7sRRyfx9o 1/2
https://tinyurl.com/BristolCVLectureship
t.co
Today we had a joint workshop between our FunAI Lab, UTN and AIST Japan. 13 talks, 1 cake and lots of Bavarian food really get research discussions going! Towards more collaborations in AI between 🇩🇪 & 🇯🇵. @hirokatukataoka.bsky.social
Also @phdcomics.bsky.social is on 🦋 👏. slowly nesting here.
Marriage vs PhD
Nice 👏! We love small (M)LLMs :) will training code also be released?
Releasing SmolVLM, a small 2 billion parameters Vision+Language Model (VLM) built for on-device/in-browser inference with images/videos. Outperforms all models at similar GPU RAM usage and tokens throughputs Blog post: huggingface.co/blog/smolvlm
LoRA et al. enable personalised model generation and serving, which is crucial as finetuned models still outperform general ones in many tasks. However, serving a base model with many LoRAs is very inefficient! Now, there's a better way: enter Prompt Generation Networks, presented today #BMVC
Hello world! Is there any tool to sync twitter and bluesky posting?
My growing list of #computervision researchers on Bsky. Missed you? Let me know. go.bsky.app/M7HGC3Y
The thingie that brings over your twitter followers worked jolly well for me. Very cool! I am following another 500 people now thanks to that… chromewebstore.google.com/detail/sky-f...
Sky Follower Bridge - Chrome Web Store
Instantly find and follow the same users from your Twitter follows on Bluesky.
chromewebstore.google.com