1275 papers
method
New world model architectures, training objectives, inference methods, or planning/control algorithms.
Literature review synthesis
Research Lines
Converts observed trajectories into compact latent dynamics and imagined rollouts for sample-efficient policy learning.
Open edge: Whether context and dynamics identifiability assumptions transfer beyond two CARL tasks or Atari-style visual control remains unverified.Learns high-dimensional visual future prediction with mask reconstruction, depth, or keypoint dynamics to improve fidelity and physical plausibility.
Open edge: Cross-domain and zero-shot generalization evidence remains limited; long-horizon physical consistency beyond reported driving and robot datasets is unresolved.Couples action and video generation, or factorizes language into primitive video plans, enabling world models to serve as policies or planners.
Open edge: How far composition generalizes to novel primitives, unseen action spaces, or real-world execution beyond simulation and demonstration data is unclear.Provides evaluation protocols for spatial consistency, memory consistency, and action responsiveness rather than frame realism alone.
Open edge: Benchmark results and baselines are only partially reported; open-domain long-term memory and cross-action-space generalization remain unresolved.Shared Direction
- World-model learning is moving from latent recurrent state-space dynamics toward transformer and diffusion architectures for high-dimensional visual prediction.
- Auxiliary self-supervised tasks are used to stabilize generation and encode geometry or physics, including mask reconstruction, depth prediction, and keypoint dynamics.
- Conditioning on context, language, goals, or actions is treated as essential for control, planning, and zero-shot generalization.
- Evaluation is expanding from visual quality to closed-loop consistency, action controllability, and spatial or temporal coherence.
Key Differences
- Representation differs between compact stochastic latent states and explicit video/diffusion generation.
- Supervision differs: some methods rely on reconstruction/generation, while others add physics-informed or action-coupling objectives.
- Evaluation targets differ across Atari/CARL benchmarks, driving video datasets, robot manipulation datasets, and open-domain first/third-person videos.
- Interaction mode differs: imagined rollouts for policy learning, decomposed language-goal video plans, unified action-video diffusion, and benchmark action-control loops.
Open Questions
- Which auxiliary objectives transfer beyond the reported task families and datasets?
- How much of reported policy or inference gains come from the learned dynamics rather than the planner, controller, or task prior?
- Can these models maintain long-term memory and spatial consistency across unseen viewpoints, scenes, and action spaces?
- Do current benchmarks isolate method contributions from surrounding system choices and dataset-specific metrics?
- Several works lack visible code, limitation statements, or benchmark results, leaving reproducibility and failure boundaries as verification gaps.