Jiatao Gu

@jgu32.bsky.social

Machine Learning Researcher @Apple MLR Incoming Assistant Professor @Penn CIS See more details https://jiataogu.me

🤔Image-to-3D, monocular depth estimation, camera pose estimation, …, can we achieve all of this with just ONE model easily? 🚀Our answer is Yes -- Excited to introduce our latest work: World-consistent Video Diffusion (WVD) with Explicit 3D Modeling! arxiv.org/abs/2412.01821

WVD Pipeline