Klara Janouskova

@klara-cz.bsky.social

Computer vision PhD student from Prague. https://klarajanouskova.github.io/

Yaaay, the biggest "side project" of my phd is out! More posts coming soon :) In the meanwhile, you can have a look at the result via our preview vrg.fel.cvut.cz/reimagenet - it also serves as a tool to report issues after login.

ReImageNet

vrg.fel.cvut.cz

Dmytro Mishkin@ducha-aiki.bsky.social · 3w ago

Doomed to Re-Annotate, Forever: The ImageNet Story Illia Volkov, Nikita Kisel, Tetiana Mishkina, @klara-cz.bsky.social Jiri Matas tl;dr: ImageNet val annotation with all bboxes, careful study of concepts, attributes (is drawing of a cat, a cat?) Year of efforts. 1/ arxiv.org/abs/2608.13783

Fable 1 (chat) getting annoyed by Fable 2's (code) work, cute. I guess this is what happens when you try to have agents implement sth to "quickly" check if it works without properly understanding it :) Don't ask how much usage I have left now 🫠

Bild

When mentoring young people (high-school) who don't know much about how to code and maybe computers overall, I am struggling with not knowing what is it that is important for them to learn these days and what knowledge is going to be completely irrelevant in a few years (months??).

David Picard@davidpicard.eurosky.social · 3mo ago

I can't believe that, now, the standard interface to a computer is natural language. The disconnect with 9yo me discovering Commodore basic is unfathomably high. It's certainly a blessing but I wonder if I would have been interested had it been not as mysterious as it was at the time.

📍 @c1rcuslegend.bsky.social is presenting our work at #CVPR2026 Findings, find the poster with the most 🐈🐾 🗓 Friday, 07:00–08:30 📌 ExHall A; Poster #166 Stop by if you're curious whether MLLMs make good classifiers or to discuss our "Doomed to Reannotate" ImageNet project (preprint coming soon) 📝

Klara Janouskova@klara-cz.bsky.social · 6mo ago

Let me introduce our new paper: Multimodal Large Language Models as Image Classifiers ❓ Multimodal LLMs are increasingly used for visual tasks, but evaluating their image classification ability has produced conflicting conclusions. Link: arxiv.org/html/2603.06...

NeurIPS paper bidding actually makes me look forward to reviewing :) Interestingly, this time my top suggestions are more about topics I have just started working on but have not published anything yet. I guess it signals there may be too many people working on the same thing 🫣

To study this, we introduce ReGT, a new multilabel reannotation of 625 ImageNet classes that corrects many of these issues. When evaluated on the cleaned labels, multimodal LLMs improve by up to +10.8% accuracy, substantially narrowing the gap with supervised vision models. 📈

Bild