📢 New paper: You look at a New Yorker cartoon. In a split second, you see the scene, spot the absurdity, and get the joke. For a multimodal LLM, that split second is a wall. Not because it can't see the image, but because it doesn't know how to reason from what it sees to what's funny.#MultimodalAI
Aykut Erdem
@aykuterdem.bsky.social
Associate Professor of Computer Science at @kocuniversity; A Researcher at @KuisAICenter; Researcher in #ComputerVision #ML #AI; Geek; Boardgamer. WWW: https://aykuterdem.github.io
Basit and Canberk are presenting our work at #SIGGRAPHAsia2024, Tokyo.
GANs like StyleGAN generate highly realistic images. But adapting them to new domains or tasks like text-guided editing or reference-guided synthesis with limited data is challenging! 🖼️✨ Our #SIGGRAPHAsia 2024 paper, HyperGAN-CLIP, tackles this: youtu.be/X0VOYFhPWxQ (1/n)
What a great way to kick off my Bluesky journey! Excited to see you all in Tokyo this December at #SIGGRAPHAsia 2024.
GANs like StyleGAN generate highly realistic images. But adapting them to new domains or tasks like text-guided editing or reference-guided synthesis with limited data is challenging! 🖼️✨ Our #SIGGRAPHAsia 2024 paper, HyperGAN-CLIP, tackles this: youtu.be/X0VOYFhPWxQ (1/n)