Recent advancements in video autoencoders (Video AEs) have significantly improved the quality and efficiency of video generation. In this paper, we propose a novel and compact video autoencoder, VidTwin, that decouples video into two distinct latent spaces: Structure latent vectors, which capture...
![[CVPR 2025] VidTwin: Video VAE with Decoupled Structure and Dynamics | Yuchi Wang (王宇驰)](https://wangyuchi369.github.io/publication/2024-vidtwin/featured.png)