DisCo: World Models with Discrete Camera Motion Control
TLDR
DisCo uses discrete camera motion primitives to improve action controllability in video world models, addressing representation entanglement.
Reasoning
The paper identifies a key bottleneck (action representation entanglement) and proposes a discrete action conditioning method, supported by a new benchmark. Strengths include clear problem identification and empirical validation; weaknesses are not explicitly stated in the abstract but may include limited discussion of failure cases or comparisons.
Read-first score
Read-first score 61.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 485.
Keyword Scores
Deep Analysis
Innovations
- Identifying action representation entanglement as a key bottleneck in controllable video generation
- Proposing discrete action primitives for camera motion control to improve action separability
- Introducing DisCoBench, a comprehensive benchmark for evaluating short-term, long-horizon, and highly dynamic exploration scenarios
Methodology
DisCo conditions video generation on a compact set of discrete action primitives to improve action separability, addressing the issue of action representation entanglement found in continuous camera representations. The model is evaluated on the proposed DisCoBench benchmark across short-term, long-horizon, and highly dynamic exploration scenarios.
Key Results
DisCo achieves significantly more reliable action following while preserving visual quality compared to existing approaches.