Klara Janouskova
@klara-cz.bsky.social
Computer vision PhD student from Prague. https://klarajanouskova.github.io/
Attending #ECCV2026? Want to know why your favorite vision models learn to encode camera acquisition and processing traces? Come meet our team! 👇 @eccv.bsky.social
One more thing: we are not treating this as finished. Fine-grained classes are where we would most welcome another pair of eyes, especially from people who know their birds, fungi or dog breeds better than we do 🐕🍄🐦https://vrg.fel.cvut.cz/reimagenet/?browse=1
ReImageNet
vrg.fel.cvut.cz
🐕 🍄 🎻 🚜 ReImageNet is out. 🚜 🎻🍄� We reannotated the whole ImageNet-1k val set from scratch: 50,000 images, multilabel, 96,051 bounding boxes, new definitions for all 1,000 classes, five semantic attributes. Browse every annotation: vrg.fel.cvut.cz/reimagenet Paper: arxiv.org/abs/2608.13783
I would also recommend our "Doomed to Re-Annotate, Forever: The ImageNet Story" paper to anybody who plans to create a new classification/detection computer vision dataset. We have thought we have it all figured out, "now it will be easy" many times, we were always wrong. arxiv.org/abs/2608.13783
arxiv.org
🐕 🍄 🎻 🚜 ReImageNet is out. 🚜 🎻🍄� We reannotated the whole ImageNet-1k val set from scratch: 50,000 images, multilabel, 96,051 bounding boxes, new definitions for all 1,000 classes, five semantic attributes. Browse every annotation: vrg.fel.cvut.cz/reimagenet Paper: arxiv.org/abs/2608.13783
🐕 🍄 🎻 🚜 ReImageNet is out. 🚜 🎻🍄� We reannotated the whole ImageNet-1k val set from scratch: 50,000 images, multilabel, 96,051 bounding boxes, new definitions for all 1,000 classes, five semantic attributes. Browse every annotation: vrg.fel.cvut.cz/reimagenet Paper: arxiv.org/abs/2608.13783
Your camera leaves fingerprints on every photo. Vision encoders learned to read them. Read about the bad and the good sides of it. "Invisible Shortcuts: Why Vision Encoders Know Your Camera" has been accepted at ECCV 2026. Paper: arxiv.org/pdf/2608.05424 @eccv.bsky.social #ECCV2026
🎉 Our #ICLR26 paper on efficient probing (EP) now has a home that outlives the paper. ⏱️ Probing Frozen Encoders — a standing ImageNet-1k benchmark for k-NN, linear (LP) and efficient probing (EP). Built to stay. Issues and PRs welcome 👇 github.com/billpsomas/e...
GitHub - billpsomas/efficient-probing: [ICLR 2026] - Official implementation of "Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency"
[ICLR 2026] - Official implementation of "Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency" - billpsomas/efficient-probing
github.com
Yaaay, the biggest "side project" of my phd is out! More posts coming soon :) In the meanwhile, you can have a look at the result via our preview vrg.fel.cvut.cz/reimagenet - it also serves as a tool to report issues after login.
ReImageNet
vrg.fel.cvut.cz
Doomed to Re-Annotate, Forever: The ImageNet Story Illia Volkov, Nikita Kisel, Tetiana Mishkina, @klara-cz.bsky.social Jiri Matas tl;dr: ImageNet val annotation with all bboxes, careful study of concepts, attributes (is drawing of a cat, a cat?) Year of efforts. 1/ arxiv.org/abs/2608.13783
Doomed to Re-Annotate, Forever: The ImageNet Story Illia Volkov, Nikita Kisel, Tetiana Mishkina, @klara-cz.bsky.social Jiri Matas tl;dr: ImageNet val annotation with all bboxes, careful study of concepts, attributes (is drawing of a cat, a cat?) Year of efforts. 1/ arxiv.org/abs/2608.13783
Something really cool is going on in Prague! 🤩 Also, I have heard they are hiring 👀 You would have a hard time looking for such an exceptional team and a CEO who cares about his team this much. And, of course, Prague is beautiful. Free canistherapy included! x.com/CsabaSzepesv...
We needed region annotations. We looked at what already existed… and we decided to build our own dataset. Meet STRAP 🎉 🤗 huggingface.co/datasets/vrg-prague/STRAP 🔗 klarajanouskova.github.io/STRAP Region–text annotations for 2M web images, from a single pass of a frozen open MLLM. 🧵
Trying to gather some strength for the supplementary deadline.
Fable 1 (chat) getting annoyed by Fable 2's (code) work, cute. I guess this is what happens when you try to have agents implement sth to "quickly" check if it works without properly understanding it :) Don't ask how much usage I have left now 🫠
EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers is now valued at more than $500 million.
The DeepMind trio who built a poker AI, are now making money for quant hedge funds | TechCrunch
EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers is now valued at more than $500 million.
techcrunch.com
And a few pics from walks with doggo 🐾
VRG researches successfully released into the wild. Reintroduction to natural habitat (the lab) scheduled for tomorrow. The summer retreat was great, thanks @gtolias.bsky.social and everybody who helped make it happen! 😊
VRG researches successfully released into the wild. Reintroduction to natural habitat (the lab) scheduled for tomorrow. The summer retreat was great, thanks @gtolias.bsky.social and everybody who helped make it happen! 😊
When mentoring young people (high-school) who don't know much about how to code and maybe computers overall, I am struggling with not knowing what is it that is important for them to learn these days and what knowledge is going to be completely irrelevant in a few years (months??).
I can't believe that, now, the standard interface to a computer is natural language. The disconnect with 9yo me discovering Commodore basic is unfathomably high. It's certainly a blessing but I wonder if I would have been interested had it been not as mysterious as it was at the time.
📍 @c1rcuslegend.bsky.social is presenting our work at #CVPR2026 Findings, find the poster with the most 🐈🐾 🗓 Friday, 07:00–08:30 📌 ExHall A; Poster #166 Stop by if you're curious whether MLLMs make good classifiers or to discuss our "Doomed to Reannotate" ImageNet project (preprint coming soon) 📝
Let me introduce our new paper: Multimodal Large Language Models as Image Classifiers ❓ Multimodal LLMs are increasingly used for visual tasks, but evaluating their image classification ability has produced conflicting conclusions. Link: arxiv.org/html/2603.06...
🚀 Reminder: Our #CVPR2026 highlight ⭐ paper "Retrieve and Segment" will be presented today at the What is Next in Multimodal Foundation Models? Workshop. 🕝 Today, 14:30–16:00 📄 Paper: arxiv.org/abs/2602.23339 💻 Code: github.com/TilemahosAra... Looking forward to the discussions and feedback!
I guess my reviews have still turned out alright 😊 #CVPR2026 @cvprconference.bsky.social
The curse of doing research as an undergrad: publish one paper on topic X, then spend your entire PhD reviewing papers on X. There’s a reason I changed topics. 🫠 At least one paper in my batch actually looks very interesting though :)
NeurIPS paper bidding actually makes me look forward to reviewing :) Interestingly, this time my top suggestions are more about topics I have just started working on but have not published anything yet. I guess it signals there may be too many people working on the same thing 🫣
🐾 He has been the goodest boy for a year already! Time for his first 🐑
Two #CVPR2026 competitions are live: S23DR 2026 and BuildingWorld 2026! Task: reconstruct house roof wireframes from point clouds and segmentations: Total prize fund: $22k Deadline: end of May 2026 1) huggingface.co/spaces/usm3d... 2) huggingface.co/spaces/Build... @cvprconference.bsky.social
Multimodal Large Language Models as Image Classifiers Nikita Kisel, Illia Volkov @klara-cz.bsky.social Jiri Matas tl;dr: if you evaluate good (chatGPT) model on a dirty (ImageNet) test set, it is bad. Yes, ImageNet test is bad nowadays. +insights from labeling. arxiv.org/abs/2603.065...
To study this, we introduce ReGT, a new multilabel reannotation of 625 ImageNet classes that corrects many of these issues. When evaluated on the cleaned labels, multimodal LLMs improve by up to +10.8% accuracy, substantially narrowing the gap with supervised vision models. 📈
Let me introduce our new paper: Multimodal Large Language Models as Image Classifiers ❓ Multimodal LLMs are increasingly used for visual tasks, but evaluating their image classification ability has produced conflicting conclusions. Link: arxiv.org/html/2603.06...