< Back to Main README
370+ papers · 25+ datasets · Covers T2V, I2V, and Video Editing
Foundation video generation from GANs to diffusion transformers, image animation, and temporal dynamics
| Sub-domain | Focus | Key Models |
|---|
| T2V Foundation Models | Core text-to-video synthesis | Sora, Wan 2.1, Veo 2, CogVideoX, LaVIE |
| Controllable Synthesis | Controllable and efficient video generation | AnimateDiff, ControlVideo, VideoComposer |
| T2V Benchmarks | Video generation benchmarks and datasets | VBench, EvalCrafter, FETV |
| I2V Animation & Portraits | Image animation, talking heads, character-driven video | LivePortrait, Animate Anyone, MagicAnimate |
| I2V Video Editing | Video editing, motion transfer, enhancement | TokenFlow, Rerender-A-Video, CoDeF |
| Model | Title | Year |
|---|
| Sora | Video Generation Models as World Simulators | 2024 |
| Wan 2.1 | Scalable Diffusion Transformer for Video Generation | 2025 |
| Veo 2 | Photorealistic Video Generation from Google DeepMind | 2025 |
| CogVideoX | Text-to-Video Diffusion Models with An Expert Transformer | 2024 |
| Kling | High-Quality Video Generation from Kuaishou | 2024 |
| Align your Latents | High-Resolution Video Synthesis with Latent Diffusion | 2023 |
| LaVIE | High-Quality Video Generation with Cascaded Latent Diffusion | 2023 |
| Make-A-Video | Text-to-Video Generation without Text-Video Data | 2022 |
| Imagen Video | High Definition Video Generation with Diffusion Models | 2022 |
| CogVideo | Large-scale Pretraining for T2V via Transformers | 2022 |
| Model | Title | Year |
|---|
| LivePortrait | Efficient Portrait Animation with Stitching and Retargeting | 2024 |
| Animate Anyone | Consistent and Controllable Image-to-Video Synthesis | 2024 |
| MagicAnimate | Temporally Consistent Human Image Animation | 2024 |
| AnimateDiff | Animate Your Personalized T2I Diffusion Models | 2023 |
| DreamPose | Fashion Image-to-Video Synthesis via Stable Diffusion | 2023 |
| Model | Title | Venue | Year |
|---|
| Sora | Video Generation Models as World Simulators | OpenAI | 2024 |
| AnimateDiff | Animate Your Personalized T2I Diffusion Models | arXiv | 2023 |
| InstructPix2Pix | Learning to Follow Image Editing Instructions | CVPR | 2023 |
| Text2LIVE | Text-Driven Layered Image and Video Editing | arXiv | 2022 |
| DreamFusion | Text-to-3D using 2D Diffusion | ICLR | 2023 |
| ProlificDreamer | High-Fidelity Text-to-3D with Variational Score Distillation | arXiv | 2023 |
| Magic3D | High-Resolution Text-to-3D Content Creation | arXiv | 2022 |
| T2M-GPT | Generating Human Motion from Textual Descriptions | arXiv | 2023 |