📅 Join us next week for the 2nd VAR Workshop at CVPR 2026: June 3rd, 2026 from 8:30am to 1pm. 🎯 We are hosting an exciting line-line of speakers: Katerina Fragkiadaki, Wenhu Chen, Michael S. Ryoo, Ziwei Liu, Yao Qin, @vicenteor.bsky.social 👉Schedule/Papers: varworkshop.github.io/schedule/
Apratim Bhattacharyya
@apratimbh.bsky.social
ML Researcher, Qualcomm AI Research | Postdoc, University of Tübingen | PhD, Max Planck Institute for Informatics | AI Assistants, Multi-modal LLMs, Autonomous Driving
Open-source fueled the LLM revolution, but Physical AI hasn't fully benefited from this flywheel yet. Today, we're launching kesai.eu, our mission to democratize robotics research! First milestone: training a frontier-level self-driving policy using significantly less data than typically required.
KE:SAI — Open Science Autonomy Lab
KE:SAI is a Franco-German non-profit open science lab for scalable autonomous intelligence.
kesai.eu
🚨🚨🚨Take part in the AI Coach: Fitness challenge and the Low Power Computer Vision Challenge @ CVPR 2026 🎯Both challenges use the Qualcomm Exercise Video Dataset (QEVD) dataset. 👉Quick start guides and sample solutions: apratimbh.github.io/whatandwhen/ @cvprconference.bsky.social
📣📣📣Our team at Qualcomm AI Research is hiring Research Interns for Summer 2026 in Toronto to work on multi-modal LLMs and embodied AI. 👉Apply here: 1) Embodied AI: qualcomm.wd12.myworkdayjobs.com/External/job... 2) Multi-modal LLMs:
FY26 Intern - Deep Learning Research Internship - Embodied AI - Canada (4 months)
Company: Qualcomm Canada ULC Job Area: Interns Group, Interns Group > Interim Engineering Intern - SW Qualcomm Overview: Qualcomm is a company of inventors that unlocked 5G ushering in an age of ra...
qualcomm.wd12.myworkdayjobs.com
🚨Submit by 1st May @cvprconference.bsky.social: extended abstracts on streaming vision-language models, real-time activity understanding, grounding, ego-centric video understanding, language and robot learning. Contributions are encouraged to include a demo! 👉Details: varworkshop.github.io/calls/
🚨🚨🚨 We are now accepting submissions!
Call for Participation @cvprconference.bsky.social: Multi-Modal LLMs - prepare to engage in a dynamic, face-to-face conversation with a real human user! Details: varworkshop.github.io/challenges/ 🚨🚨🚨The winning teams will receive a prize and a contributed talk. P.S. GPT-4o does not do too well.
Call for Participation @cvprconference.bsky.social: Multi-Modal LLMs - prepare to engage in a dynamic, face-to-face conversation with a real human user! Details: varworkshop.github.io/challenges/ 🚨🚨🚨The winning teams will receive a prize and a contributed talk. P.S. GPT-4o does not do too well.
Call for Participation: We're excited to announce a challenge focused on developing AI assistants that can guide users through workout sessions with intelligent feedback! 🚨The winning teams will receive a prize along with a contributed talk. 🚨 Website: varworkshop.github.io/challenges/
🚨Submission are now open!
Call for Papers and Demos @cvprconference.bsky.social: on topics such as streaming vision-language models, real-time activity understanding, grounding, ego-centric video understanding, language and robot learning. Contributions are encouraged to include a demo! Link: varworkshop.github.io/calls/
Call for Papers and Demos @cvprconference.bsky.social: on topics such as streaming vision-language models, real-time activity understanding, grounding, ego-centric video understanding, language and robot learning. Contributions are encouraged to include a demo! Link: varworkshop.github.io/calls/
Join us at the @cvprconference.bsky.social Workshop on Vision-based Assistants in the Real-world (VAR) and tackle one of AI's biggest challenges: building systems that can comprehend and reason about dynamic, real-world scenes. Workshop Page: varworkshop.github.io
Vision-based Assistants in the Real-World
VAR Workshop @ CVPR 2025.
varworkshop.github.io
The list of #CVPR2025 workshops is now available. List: cvpr.thecvf.com/Conferences/...
By popular demand, we are extending #CVPR2025 coverage to Bluesky. Stay tuned!
🚨We present in "Enhancing Hallucination Detection through Noise Injection" [https://arxiv.org/pdf/2502.03799], an efficient approach to detect hallucinations in LLMs, within a Bayesian framework. TL; DR - We use noise injection to capture both epistemic and aleatoric uncertainty!
Join us at the CVPR 2025 Workshop on Vision-based Assistants in the Real-world (VAR) and tackle one of AI's biggest challenges: building systems that can comprehend and reason about dynamic, real-world scenes. Workshop Page: varworkshop.github.io
🚨Check out our new work on distilling reasoning skills from LLMs into efficient driving policies, to deal with critical "long-tail" scenarios. arXiv: arxiv.org/abs/2501.09757
🚨 The code for our NeurIPS 2024 (D&B track) paper: ClevrSkills: Compositional Language And Visual Understanding in Robotics (arxiv.org/abs/2411.09052), is now available. GitHub Repo: github.com/Qualcomm-AI-... Dataset Page: www.qualcomm.com/developer/so...
ClevrSkills AI Dataset
qualcomm.com