How do we teach vision models to infer visual concepts from just a handful of example images? Check out our new ECCV paper, where we explore learning directly from image sets 👇
Most representation learning methods produce generic embeddings that try to preserve everything in an image. We introduce VICIS: a way to use example sets to define a tailored embedding space for what matters. This is useful whenever the desired visual signal is easier to show than to describe. 🧵👇