Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DisCo: World Models with Discrete Camera Motion Control

arXiv 2026 61.6 method

TLDR

DisCo uses discrete camera motion primitives to improve action controllability in video world models, addressing representation entanglement.

Reasoning

The paper identifies a key bottleneck (action representation entanglement) and proposes a discrete action conditioning method, supported by a new benchmark. Strengths include clear problem identification and empirical validation; weaknesses are not explicitly stated in the abstract but may include limited discussion of failure cases or comparisons.

Read-first score

Read-first score 61.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
95

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94975132

Topical relevance 29%
78.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,experiment

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 485.

Keyword Scores

world model
10
video world model
10
interactive world model
9
world dynamics prediction
9
generative world model
8
world simulator
7
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Identifying action representation entanglement as a key bottleneck in controllable video generation
  • Proposing discrete action primitives for camera motion control to improve action separability
  • Introducing DisCoBench, a comprehensive benchmark for evaluating short-term, long-horizon, and highly dynamic exploration scenarios

Methodology

DisCo conditions video generation on a compact set of discrete action primitives to improve action separability, addressing the issue of action representation entanglement found in continuous camera representations. The model is evaluated on the proposed DisCoBench benchmark across short-term, long-horizon, and highly dynamic exploration scenarios.

Key Results

DisCo achieves significantly more reliable action following while preserving visual quality compared to existing approaches.

Tags

video generationworld modelscamera controldiscrete actionsaction representationcontrollable generationCV