7 Related Work7相关工作7.1 Video Generation Models7.1 视频生成模型Diffusion-based Video Generation. The success of Diffusion Models (DMs) [1o] in image synthesis has catalyzed their extension to the video domain. Early pioneers like Video Diffusion Models (VDM) [11] first extended the standard 2D U-Net to a 3D structure [156] by replacing 2D convolutions with space-time factorized convolutions. To alleviate the heavy computational burden of 3D operators, many subsequent works [12-15, 155] adopted a “spatial-then-temporal” paradigm, inserting iD temporal attention layers after 2D spatial blocks to capture dynamic dependencies. A significant architectural shift occurred with the introduction of Diffusion Transformers (DiT) [6
Kairos03:面向 Physical AI 的具备后悔感知能力的原生世界-动作模型栈
7 Related Work7相关工作7.1 Video Generation Models7.1 视频生成模型Diffusion-based Video Generation. The success of Diffusion Models (DMs) [1o] in image synthesis has catalyzed their extension to the video domain. Early pioneers like Video Diffusion Models (VDM) [11] first extended the standard 2D U-Net to a 3D structure [156] by replacing 2D convolutions with space-time factorized convolutions. To alleviate the heavy computational burden of 3D operators, many subsequent works [12-15, 155] adopted a “spatial-then-temporal” paradigm, inserting iD temporal attention layers after 2D spatial blocks to capture dynamic dependencies. A significant architectural shift occurred with the introduction of Diffusion Transformers (DiT) [6