Spyros Gidaris

@spyrosgidaris.bsky.social

Senior Research Scientist at Valeo.ai (@valeoai.bsky.social) https://gidariss.github.io/

Discovered that our RangeViT paper keeps being cited in what might be LLM-generated papers. Number of citations increased rapidly in the last weeks. Too good to be true. Papers popped up on different platforms, but mainly on ResearchGate with ~80 papers in just 3 weeks. [1/]

1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.

Bild

I am at #CVPR2025 this week in Nashville! Presenting "Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers" on multi-modal semantic future prediction. Come discuss! Fri 13 Jun 10:30-12:30, poster #345 bsky.app/profile/sta8...

@sta8is.bsky.social · 2y ago

🧵 Excited to share our latest work: FUTURIST - A unified transformer architecture for multimodal semantic future prediction, is accepted to #CVPR2025! Here's how it works (1/n) 👇 Links to the arxiv and github below

1/n Introducing ReDi (Representation Diffusion): a new generative approach that leverages a diffusion model to jointly capture – Low-level image details (via VAE latents) – High-level semantic features (via DINOv2)🧵

Bild

The @valeoai.bsky.social team is presenting a few exciting works @iclr-conf.bsky.social this year on masked generative transformers, adaptation of VLMs, self-supervised representation learning, neural solvers. #iclr2025 Check them out 👇

valeo.ai@valeoai.bsky.social · last yr.

Our recent research will be presented at #ICLR2025 @iclr_conf: VLMs, LLMs, diffusion models, self-supervised learning, physics-informed learning… Find out more below 🧵 valeoai.github.io/posts/2025-0...

Still mesmerized by this work and its results: a mid-to-end driving agent trained with self-play on just 8 maps on 1.6B km of driving (9500 years of subjective driving experience) smashes in off-the-shelf manner all existing benchmarks (nuPlan, CARLA, Waymax) 😮

Andrei Bursuc@abursuc.bsky.social · 2y ago

Crazily amazing work by @eugenevinitsky.bsky.social @senerozan.bsky.social & team, setting the bar so high for anyone working in autonomous driving these days. Check it out arxiv.org/abs/2502.03349

EQ-VAE: Such a simple & cool trick to regularize multiple kinds of autoencoders: align reconstruction of transformed latents w/ the corresponding transformed inputs. 🚀REPA: 4x training speedup 🚀MaskGIT: 2x training speedup 🚀DiT-XL/2: 7x faster convergence Kudos @nicolabourbaki.bsky.social et al.

Thodoris Kouzelis@nicolabourbaki.bsky.social · 2y ago

1/n🚀If you’re working on generative image modeling, check out our latest work! We introduce EQ-VAE, a simple yet powerful regularization approach that makes latent representations equivariant to spatial transformations, leading to smoother latents and better generative models.👇

The things I've found hardest about research have all been non-technical: maintaining confidence and self-esteem, not abandoning the work when it's too hard or stressful, finding time to learn new things. In comparison, the technical parts are much easier

1/n🚀If you’re working on generative image modeling, check out our latest work! We introduce EQ-VAE, a simple yet powerful regularization approach that makes latent representations equivariant to spatial transformations, leading to smoother latents and better generative models.👇

Bild

1/n 🚀 Excited to share our latest work: DINO-Foresight, a new framework for predicting the future states of scenes using Vision Foundation Model features! Links to the arXiv and Github 👇

Bild