TimDarcet

@timdarcet.bsky.social

PhD student, SSL for vision @ MetaAI & INRIA tim.darcet.fr

(3/3) LUDVIG uses a graph diffusion mechanism to refine 3D features, such as coarse segmentation masks, by leveraging 3D scene geometry and pairwise similarities induced by DINOv2.

Brilliant talk by Ilya, but he's wrong on one point. We are NOT running out of data. We are running out of human-written text. We have more videos than we know what to do with. We just haven't solved pre-training in vision. Just go out and sense the world. Data is easy.

Bild

Along with INQUIRE, we introduce iNat24, a new dataset of 5 million research-grade images from @inaturalist with 10,000 species labels. This is one of the largest publicly available natural world image repositories!

Bild

A fun thesis experiment: ResNet, DETR, and CLIP tackle Saint-Bernards. 🐶 ResNet focused on **fur** patterns, DETR too but also use **paws** (possibly because it helps define bounding boxes), and CLIP **head** concept oddly included human heads — language shaping learned concepts?

An image showing how three model top concept look like to classify st bernard, resnet use head and fur, while detr also use paws (maybe it help him delimitate the boundary). Clip use the head of the st bernard, but oodly the head seems to also react to human head...

These opportunities are mostly reserved for the rest of the world. We need similar Industry-Academia PhD programs in the US too! We need an american version of the CIFRE.

Jakob Foerster@jfoerst.bsky.social · 2y ago

Hello BlueSky! Joao Henriques (joao.science) and I are hiring a fully funded PhD student (UK/international) for the FAIR-Oxford program. The student will spend 50% of their time @UniofOxford and 50% @MetaAI (FAIR) in London, while completing a DPhil (Oxford PhD). Deadline: 2nd of Dec AOE!!

Sidenote: TMLR is such a pleasant journal. It's fast and reviews are (mostly) insightful, detailed and helpful. Kind of how conference reviews were before the big rush, for the youngsters who thought It's always been that way.

DinoV2 is without a doubt one of the most important Self Supervised Learning (SSL) methods right now. But training it takes 32 80Gb GPUs which is not easy to come by for small labs. What if we could train a comparable high-res model on 24Gb of VRAM? That's what I hope to show you here soon!🤞🧵 #mlsky

Vision transformers need registers! Or at least, it seems they 𝘸𝘢𝘯𝘵 some… ViTs have artifacts in attention maps. It’s due to the model using these patches as “registers”. Just add new tokens (“[reg]”): - no artifacts - interpretable attention maps 🦖 - improved performances! arxiv.org/abs/2309.16588

Bild