We needed region annotations. We looked at what already existedโฆ and we decided to build our own dataset. Meet STRAP ๐ ๐ค huggingface.co/datasets/vrg-prague/STRAP ๐ klarajanouskova.github.io/STRAP Regionโtext annotations for 2M web images, from a single pass of a frozen open MLLM. ๐งต
Vladan Stojniฤ
@stojnicv.xyz
Ph.D. student at Visual Recognition Group, Czech Technical University in Prague ๐ https://stojnicv.xyz
Our #ECCV2026 Whareformer paper tackles long-term object tracking in Egocentric videos ๐https://arxiv.org/abs/2607.08537 see Dima's thread below ๐ฝ joint work w Jacob Chalk, @saptarshisinha.bsky.social @dimadamen.bsky.social from Bristol & @skamalas.bsky.social from @naverlabseurope.bsky.social
We just released our #ECCV2026 paper on Model Merging for Computer Vision ๐ arxiv.org/abs/2604.12935 Joint work w @pdejorge.bsky.social Cesar De Souza @bjoernmichele.bsky.social @mbsariyildiz.bsky.social @weinzaepfelp.bsky.social Florent Perronnin & @skamalas.bsky.social See Pau's thread below ๐ฝ
๐ข Deadline Extended! We have extended the Student Support Grant application deadline for the ILR+G Workshop @eccv.bsky.social by 2 days. If you haven't applied yet, there's still time! ๐๏ธ New deadline: August 2, 2026 ๐ Apply now: forms.gle/qvqn2LyWXCvN... We look forward to seeing you in Malmรถ!
forms.gle
๐ Student Support Grants are now available! Are you a student working on instance-level recognition and generation? We offer Student Support Grants for participants attending the ILR+G Workshop @eccv.bsky.social ๐ Application deadline: July 31, 2026 ๐ Apply here: forms.gle/qvqn2LyWXCvN...
Check out Pau's thread on our new #ECCV2026 paper! ๐งต We propose a proxy for making model merging scalable for real-world computer vision tasks (merging encoders across diverse domains like 2D, 3D, and human understanding)
1/6 Excited to share that our paper on model merging was accepted at ECCV 2026! ๐ We introduce an efficient, decoder-free proxy that makes model selection faster, simpler and practical across vision tasks. ๐ arxiv.org/abs/2604.12935 ๐ europe.naverlabs.com/task-alignment ๐งต๐
1/6 Excited to share that our paper on model merging was accepted at ECCV 2026! ๐ We introduce an efficient, decoder-free proxy that makes model selection faster, simpler and practical across vision tasks. ๐ arxiv.org/abs/2604.12935 ๐ europe.naverlabs.com/task-alignment ๐งต๐
๐ Student Support Grants are now available! Are you a student working on instance-level recognition and generation? We offer Student Support Grants for participants attending the ILR+G Workshop @eccv.bsky.social ๐ Application deadline: July 31, 2026 ๐ Apply here: forms.gle/qvqn2LyWXCvN...
forms.gle
๐จ Presenting today at #CVPR2026! Our highlight โญ paper Retrieve and Segment (RNS) asks a simple question: Can a few examples bridge the supervision gap in open-vocabulary segmentation? โ Turns out they can. ๐ Poster #578 โฐ 16:45โ18:45 ๐ arxiv.org/abs/2602.23339
๐น๐๐๐๐๐๐๐ ๐๐๐ ๐บ๐๐๐๐๐๐ (๐น๐ต๐บ) at CVPR 2026 โ selected as a ๐ ๐ฏ๐๐๐๐๐๐๐๐ ๐ (top ~4% of submissions)! ๐ช๐ต๐ฒ๐ฟ๐ฒ ๐๐ผ ๐ณ๐ถ๐ป๐ฑ ๐๐ ๐ฎ๐ ๐๐ฉ๐ฃ๐ฅ ๐ฎ๐ฌ๐ฎ๐ฒ (๐๐ฒ๐ป๐๐ฒ๐ฟ): ๐น ๐ด๐๐๐ ๐ช๐๐๐๐๐๐๐๐๐ Poster Session 4 (#578) โ June 6, 16:45โ18:45 ๐น ๐พ๐๐๐ ๐๐ ๐ต๐๐๐ ๐๐ ๐ด๐๐๐๐๐๐๐ ๐๐ ๐ญ๐๐๐๐ ๐๐๐๐๐ ๐ด๐๐ ๐๐๐? Workshop โ June 3, 14:30โ16:00
Vision Transformers are brittle to variable input resolutions. In dense prediction tasks, the standard fix is sliding-window inference with heavy overlap, which is effective but painfully slow. SPAR (Single-Pass Any-Resolution ViT), takes a different approach.
๐จ Call for Papers ๐จ 8th Instance-Level Recognition and Generation (ILR+G) Workshop at @eccv.bsky.social ๐ Malmรถ, Sweden ๐ September 8โ9, 2026 ๐ ilr-workshop.github.io/ECCVW2026/ Submission deadline: June 26, 2026 Notification of acceptance: July 24, 2026 #ECCV2026 #ComputerVision
Do vision and vision-language foundation models suffer from shortcut learning? Is this only a curse or also a blessing? I'm excited to discuss our recent work at the AMD AI Research Club session with Pier Luigi Dovesi. ๐ May 28, 2026 โ 2PM Prague time Register here ๐ amd.zoom.us/webinar/regi...
Welcome! You are invited to join a webinar: Invisible Shortcuts: Why Vision Encoders Know Your Camera | AMD AI Research Club. After registering, you will receive a confirmation email about joining the...
Join us for the third session of the AMD AI Research Club, a peer-to-peer paper spotlight series connecting AI researchers, engineers, and practitioners shaping the future of AI. In this session, Gio...
amd.zoom.us
๐ง๐ท Presenting our ICLR 2026 paper โEfficient Probingโ (EP) today! โWhat if linear probing is asking the wrong question? ๐ฅณ EP is a lightweight attention probing method that better evaluates local, patch-level representations from models like MIM. ๐Friday 24 April, P4-#3713, 15:15โ17:45
๐จ Efficient Local Visual Similarity (ELViS) @ #ICLR 2026 ๐ง๐ท ELViS is a fast, lightweight, and interpretable module for estimating image-to-image similarity that generalizes well to many image domains. Paper: arxiv.org/abs/2603.28603 Code: github.com/pavelsuma/ELViS Come see poster today @ P4-#3715
The Visual Recognition Group at CTU in Prague organizes the 51st Pattern Recognition and Computer Vision Colloquium with Roman Bednarik, Istvรกn Sรกrรกndi, Andreas Geiger, Jan ล kvrna, Yuki Asano and Dรกniel Barรกth. cmp.felk.cvut.cz/colloquium/#... Happening today, Thursday Apr 23, 11:00-17:00.
Got a few labeled images lying around? You can use them to drastically improve your open-vocabulary segmentation! Check out RnS, which boosts OVS baselines by up to 34%. ๐๐
1/n #CVPR2026 Accepted Paper๐ ๐จ๐๐ ๐ ๐ญ๐๐ ๐ฌ๐๐๐๐๐๐๐ ๐ฌ๐๐๐๐๐ ๐๐ ๐ฉ๐๐๐ ๐๐ ๐๐๐ ๐บ๐๐๐๐๐๐๐๐๐๐ ๐ฎ๐๐ ๐๐ ๐ถ๐๐๐-๐ฝ๐๐๐๐๐๐๐๐๐ ๐บ๐๐๐๐๐๐๐๐๐๐๐? ๐น๐๐๐๐๐๐๐ ๐๐๐ ๐บ๐๐๐๐๐๐ (๐น๐ต๐บ) answers this question. Paper/code at the end๐๐ผ
Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation? Tilemachos Aravanis @stojnicv.xyz @billpsomas.bsky.social Nikos Komodakis @gtolias.bsky.social tl;dr: almost yes if use 1-3 images, no if more(fig 6) arxiv.org/abs/2602.23339 #CVPR2026
Excited to share that our paper "Global-Aware Edge Prioritization for Pose Graph Initialization" has been accepted to CVPR 2026! #CVPR2026 See you soon in Denver!๐ฅณ๐ฅณ Code is coming soon๐ง โHow would you do an accurate and efficient pose graph initialization in a global manner? arxiv.org/abs/2602.21963
Global-Aware Edge Prioritization for Pose Graph Initialization
The pose graph is a core component of Structure-from-Motion (SfM), where images act as nodes and edges encode relative poses. Since geometric verification is expensive, SfM pipelines restrict the pose...
arxiv.org
Global-Aware Edge Prioritization for Pose Graph Initialization @weitong8591.bsky.social, @gtolias.bsky.social, Jiri Matas, @danielbarath.bsky.social tl;dr: rank pose graph edges->global consistency->improve SfM arxiv.org/abs/2602.21963
1/n Attention, Please! ๐ Our work โRevisiting Attentive Probing Through the Lens of Efficiencyโ has been accepted at #ICLR2026. We introduce Efficient Probing (EP) โ a lightweight, multi-query attentive probing method for frozen encoders. Paper + code at the end ๐
What if position encodings were designed for vision from scratch? We introduce PaPEโParabolic Position Encoding. Outperforms RoPE on 7/8 datasets and extrapolates to higher resolutions without fine-tuning or position interpolation. Paper, code, and website in thread ๐งต
I have an opening for a two years post-doc position on instance-level (personalized) visual generation. Eligibility: (i) <=7 years from Ph.D. (ii) studies or 1 year outside of Czechia (ii) >=3 journal with IF or CORE A*/A conference papers. Deadline: 15 Feb. Details: www.euraxess.cz/jobs/399390
Postdoctoral research position in Instance-level visual generation
Czech Technical University in Prague (CTU) offers a fellowship program, the CTU Global Postdoc Fellowship. This new and attractive two-year fellowship-program offers excellent researchers who have rec...
euraxess.cz
1/n REGLUE Your Latents! ๐ We introduce REGLUE: a unified framework that entangles VAE latents โ Global โ Local semantics for faster, higher-fidelity image generation. Links (paper + code) at the end๐
Announcing the first AI for Peace Workshop @ ICLR 2026. The workshop aims to provide a forum for examining the relationship between AI research and its military, surveillance, and conflict-related applications. aiforpeaceworkshop.github.io
Home
Building Bridges, Not Weapons: AI for Peaceful Progress
aiforpeaceworkshop.github.io
Are you at BMVC tomorrow and interested in VLMs? Come see our poster "Image Recognition with Vision and Language Embeddings of VLMs". ๐๏ธ๐ TLDR: We benchmark VLMs on language- and vision-based classification, and propose a simple, training-free vision-language fusion. Link: arxiv.org/pdf/2509.09311
We have PhD opportunity (start date Sep 2026) at the University of Edinburgh at the intersection of biodiversity mapping and zoonotic disease prediction. It is part of the UKRI AI Centre for Doctoral Training in Biomedical Innovation based in the School of Informatics: ai4bi-cdt.ed.ac.uk
23rd of October. @iccv.bsky.social We present our work on โLarge-scale Pretraining for Grounded Video Caption Generationโ with Cordelia Schmid and @josef-sivic.bsky.social at Exhibit Hall I #434 on the morning session. Weโll have a search demo for our dataset as well! See you there!! ๐
Are you at @iccv.bsky.social #ICCV2025? Come by our poster. ย ๐ October 22, 2025, 14:30 โ 16:30 HST ย ๐ Location: Exhibit Hall I, Poster #207
Have you ever asked yourself how much your favorite vision model knows about image capture parameters (e.g., the amount of JPEG compression, the camera model, etc.)? Furthermore, could these parameters influence its semantic recognition abilities?
๐บ Just 4 days to go! Join us in Honolulu for the Instance-Level Recognition and Generation Workshop at #ICCV2025 ๐ ๐๏ธ Oct 19, 8:30amโ12:30pm ๐ Room 306 A Weโll have amazing keynotes, plus oral and poster sessions featuring accepted and invited papers. Donโt miss it! ilr-workshop.github.io/ICCVW2025/
The Visual Recognition Group at CTU in Prague organizes the 50th Pattern Recognition and Computer Vision Colloquium with Torsten Sattler, Paul-Edouard Sarlin, Vicky Kalogeiton, Spyros Gidaris, Anna Kukleva, and Lukas Neumann. On Thursday Oct 9, 11:00-17:00. cmp.felk.cvut.cz/colloquium/
Super happy that QuARI: Query Adaptive Retrieval Improvement was accepted at #NeurIPS2025. You can significantly boost retrieval performance for very hard retrieval tasks by learning query-specific transformations of your encoders. w/ @jacobsn.bsky.social @pless.bsky.social arxiv.org/pdf/2505.21647