Motion Prompting: Controlling Video Generation with Motion Trajectories
TLDR
Introduces motion prompts for controlling video generation via sparse/dense motion trajectories, enabling diverse applications and emergent physics.
Reasoning
Strengths include flexible motion conditioning, diverse applications (camera/object control, motion transfer), and quantitative/human evaluation. Weaknesses: core contribution is video generation control, not world modeling; claims about future world models are speculative and unsupported.
Read-first score
Read-first score 35.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 5.
Field roles
Rank sensitivity
Stability: volatile; rank range: 90.
Keyword Scores
Deep Analysis
Innovations
- Introduces motion prompts as a flexible conditioning representation for video generation that can encode spatio-temporally sparse or dense trajectories, object-specific or global scene motion, and temporally sparse motion.
- Proposes motion prompt expansion to translate high-level user requests into detailed, semi-dense motion prompts.
- Demonstrates versatility through applications including camera and object motion control, interacting with an image, motion transfer, and image editing.
Methodology
The authors train a video generation model conditioned on spatio-temporally sparse or dense motion trajectories. The conditioning representation, termed motion prompts, can encode any number of trajectories, object-specific or global scene motion, and temporally sparse motion. They also introduce motion prompt expansion to convert high-level user requests into detailed, semi-dense motion prompts.
Key Results
The model shows versatility across multiple applications (camera/object motion control, image interaction, motion transfer, image editing) and exhibits emergent behaviors such as realistic physics. Quantitative evaluation and a human study confirm strong performance.