(1/n) Thrilled to share my first paper at Meta FAIR! "EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data" ๐ถ Human infants learn language from sparse, noisy multimodal input. Today's VLMs can't. We built a benchmark + challenge to close that gap. ๐งต
Jiaang Li
@jiaangli.bsky.social
PhD student at University of Copenhagen @belongielab.org | #nlp #computervision | ELLIS student @ellis.eu ๐ https://jiaangli.github.io/
๐ Excited to share our work "๐ฅ๐๐ฉ๐๐ก๐๐: ๐ ๐๐ฒ๐ป๐ฐ๐ต๐บ๐ฎ๐ฟ๐ธ ๐ณ๐ผ๐ฟ ๐ ๐๐น๐๐ถ๐บ๐ผ๐ฑ๐ฎ๐น ๐ฅ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น-๐๐๐ด๐บ๐ฒ๐ป๐๐ฒ๐ฑ ๐ฉ๐ถ๐๐๐ฎ๐น ๐๐๐น๐๐๐ฟ๐ฒ ๐จ๐ป๐ฑ๐ฒ๐ฟ๐๐๐ฎ๐ป๐ฑ๐ถ๐ป๐ด", accepted at #ICLR2026! ๐ง๐ท I'll be attending ICLR in person โ would love to connect and chat there! ๐ค ๐๏ธ Sat, Apr 25, 2026, 10:30 AM โ 1:00 PM GMT-03 ๐ Pavilion 4 P4-# 3618
Feeling overwhelmed by all the recent developments in video understanding? What used to require dozens of modular computational workflows involving SLAM, feature tracking, optical flow, camera calibration, multiview geometric constraints, and resnet backbones is now... (1/3)
Feel free to reach out and chat with Xinyi on July 18th in Vancouver at the #ICML
Excited to present at the #ICML2025 World Models Workshop! ๐ July 18, 15:45โ17:00 ๐ง What if Othello-Playing Language Models Could See? We show that visual grounding improves prediction & internal structure.โ๏ธ
Would you present your next NeurIPS paper in Europe instead of traveling to San Diego (US) if this was an option? Sรธren Hauberg (DTU) and I would love to hear the answer through this poll: (1/6)
NeurIPS participation in Europe
We seek to understand if there is interest in being able to attend NeurIPS in Europe, i.e. without travelling to San Diego, US. In the following, assume that it is possible to present accepted papers ...
docs.google.com
Check out our new preprint ๐๐๐ง๐ฌ๐จ๐ซ๐๐๐๐. We use a robust decomposition of the gradient tensors into low-rank + sparse parts to reduce optimizer memory for Neural Operators by up to ๐๐%, while matching the performance of Adam, even on turbulent NavierโStokes (Re 10e5).
PhD student, Jiaang Li and his collaborators, with insights into cultural understanding of vision-language models ๐
๐New Preprint๐ Can Multimodal Retrieval Enhance Cultural Awareness in Vision-Language Models? Excited to introduce RAVENEA, a new benchmark aimed at evaluating cultural understanding in VLMs through RAG. arxiv.org/abs/2505.14462 More details:๐
I am excited to announce our latest work ๐ "Cultural Evaluations of Vision-Language Models Have a Lot to Learn from Cultural Theory". We review recent works on culture in VLMs and argue for deeper grounding in cultural theory to enable more inclusive evaluations. Paper ๐: arxiv.org/pdf/2505.22793
๐New Preprint๐ Can Multimodal Retrieval Enhance Cultural Awareness in Vision-Language Models? Excited to introduce RAVENEA, a new benchmark aimed at evaluating cultural understanding in VLMs through RAG. arxiv.org/abs/2505.14462 More details:๐
I wonโt be attending #ICLR in person this year๐ข. But feel free to check our paper โRevisiting the Othello World Model Hypothesisโ with Anders Sรธgaard, accepted at ICLR world models workshop! Paper link arxiv.org/abs/2503.04421
Revisiting the Othello World Model Hypothesis
Li et al. (2023) used the Othello board game as a test case for the ability of GPT-2 to induce world models, and were followed up by Nanda et al. (2023b). We briefly discuss the original experiments, ...
arxiv.org
Thrilled to announce "Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation" is accepted as a Spotlight (5%) at #ICLR2025! Our model MM-FSS leverages 3D, 2D, & text modalities for robust few-shot 3D segmentationโall without extra labeling cost. ๐คฉ arxiv.org/pdf/2410.22489 More details๐
Forget just thinking in words. ๐Our New Preprint: ๐ New Era of Multimodal Reasoning๐จ ๐ Imagine While Reasoning in Space with MVoT Multimodal Visualization-of-Thought (MVoT) revolutionizes reasoning by generating visual "thoughts" that transform how AI thinks, reasons, and explains itself.
FGVC12 Workshop is coming to #CVPR 2025 in Nashville! Are you working on fine-grained visual problems? This year we have two peer-reviewed paper tracks: i) 8-page CVPR Workshop proceedings ii) 4-page non-archival extended abstracts CALL FOR PAPERS: sites.google.com/view/fgvc12/...
FGVC12 Workshop - Submission
Call for Papers Workshop - Date TBC (either June 11th or 12th 2025) FGVC12 will have two paper tracks and a nectar track: Proceedings track: 8-page papers that will appear in the official CVPR worksho...
sites.google.com
FGVC12 Workshop accepted to CVPR 2025, Nashville! CALL FOR PAPERS: sites.google.com/view/fgvc12/... We discuss domains where expert knowledge is typically required and investigate artificial systems that can efficiently distinguish a large number of very similar visual concepts. #CVPR #CVPR2025 #AI
Hereโs a short film produced by the Danish Royal Academy of Sciences, showcasing the WineSensed ๐ท project of รรณranna Bender et al. thoranna.github.io/learning_to_...
VidenSkaber | Min AI forstรฅr mig ikke - professor Serge Belongie
YouTube video by Videnskabernes Selskab
youtu.be
From San Diego to New York to Copenhagen, wishing you Happy Holidays!๐
With @neuripsconf.bsky.social right around the corner, weโre excited to be presenting our work soon! Hereโs an overview (1/5)
Hereโs a starter pack with members of our lab that have joined Bluesky
Belongie Lab
Join the conversation
go.bsky.app
No one can explain stochastic gradient descent better than this panda.
a panda bear is rolling around in the grass in a zoo enclosure .
Alt: a panda bear is rolling around in the grass in a zoo enclosure .
media.tenor.com
๐คDo Vision and Language Models Share Concepts? ๐ We present an empirical evaluation and find that language models partially converge towards representations isomorphic to those of vision models. #EMNLP ๐ direct.mit.edu/tacl/article...
I'm recruiting 1-2 PhD students to work with me at the University of Colorado Boulder! Looking for creative students with interests in #NLP and #CulturalAnalytics. Boulder is a lovely college town 30 minutes from Denver and 1 hour from Rocky Mountain National Park ๐ Apply by December 15th!
Logging on! ๐งโ๐ป๐ฆ We're the Belongie Lab led by @sergebelongie.bsky.social. We study Computer Vision and Machine Learning, located at the University of Copenhagen and Pioneer Centre for AI. Follow along to hear about our research past and present! www.belongielab.org
Belongie Lab - Home
Belongie Lab -- Home.
belongielab.org
A new approach to training models in memory-constrained settings, LoQT allows for the pre-training of a 13B LLM on a 24GB GPU without model parallelism, checkpointing, or offloading strategies during training Code: github.com/sebulo/LoQT
GitHub - sebulo/LoQT
Contribute to sebulo/LoQT development by creating an account on GitHub.
github.com