Stella Frank

@scfrank.bsky.social

Thinking about multimodal representations | Postdoc at UCPH/Pioneer Centre for AI (DK).

1/8 🧵 GPT-5's storytelling problems reveal a deeper AI safety issue. I've been testing its creative writing capabilities, and the results are concerning - not just for literature, but for AI development more broadly. 🚨

My Lab at the University of Edinburgh🇬🇧 has funded PhD positions for this cycle! We study the computational principles of how people learn, reason, and communicate. It's a new lab, and you will be playing a big role in shaping its culture and foundations. Spread the words!

Bild

🚀 DinoV3 just became the new go-to backbone for geoloc! It outperforms CLIP-like models (SigLip2, finetuned StreetCLIP)… and that’s shocking 🤯 Why? CLIP models have an innate advantage — they literally learn place names + images. DinoV3 doesn’t.

Bild

New paper hot off the press www.nature.com/articles/s41... We analysed over 40,000 computer vision papers from CVPR (the longest standing CV conf) & associated patents tracing pathways from research to application. We found that 90% of papers & 86% of downstream patents power surveillance 1/

Computer-vision research powers surveillance technology - Nature

An analysis of research papers and citing patents indicates the extensive ties between computer-vision research and surveillance.

nature.com

I am excited to announce our latest work 🎉 "Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory". We review recent works on culture in VLMs and argue for deeper grounding in cultural theory to enable more inclusive evaluations. Paper 🔗: arxiv.org/pdf/2505.22793

Paper title "Cultural Evaluations of Vision-Language Models
Have a Lot to Learn from Cultural Theory"

as an extra take-away, this implies that our eval tends to be overly precision focused. we should really think of what we lose in terms of recalls, as this directly relates to what we miss out for whom when we build these large-scale, general-purpose models. (4/4)

Bild

🚀 We are excited to introduce Kaleidoscope, the largest culturally-authentic exam benchmark. 📌 Most VLM benchmarks are English-centric or rely on translations—missing linguistic & cultural nuance. Kaleidoscope expands in-language multilingual 🌎 & multimodal 👀 VLMs evaluation

Bild

Today we are releasing Kaleidoscope 🎉 A comprehensive multimodal & multilingual benchmark for VLMs! It contains real questions from exams in different languages. 🌍 20,911 questions and 18 languages 📚 14 subjects (STEM → Humanities) 📸 55% multimodal questions

Bild

Thanks to these insects, we can now study environmental microplastics retrospectively. 🔍 Even before Duprat began his now famous experiments with caddisfly larvae, insects in the wild were already experimenting with plastic... 🐛 14/x

Above: Casing of Ironoquia dubia (RMNH.INS.1544419) collected on May 18th 1971 in Loenen, The Netherlands. b) The label of the specimen.  Depicted on the right: detail of the artificial items.  Photographs: overview: Auke-Florian Hiemstra, details: Pasquale Ciliberti. Below: Caddisfly larvae in the studio of Hubert Duprat, carrying cases made from mostly gold. © Hubert Duprat, adagp, 2024, Courtesy the Artist and Art : Concept, Paris, Photo F. Delpech.

Does anyone have a nice guide to digital privacy aimed at non-citizen residents traveling & returning to the US? I've been trying to check in with scientists I know with upcoming travel and am realizing many folks could use a 101 on removing biometric unlocking, doing a fresh iOS install, etc.