Awesome World Model Hub Papers · Datasets · Projects

Aggregate analysis

Research Analysis

Cross-paper synthesis of shared research patterns, differences, mainstream directions, and trend evolution.

1275 papers

method

New world model architectures, training objectives, inference methods, or planning/control algorithms.

21 Citation impact 1 Citation velocity 63 Methodology quality 90 Recency

Literature review synthesis

Research Lines

Stochastic latent state-space world models for model-based reinforcement learning

Converts observed trajectories into compact latent dynamics and imagined rollouts for sample-efficient policy learning.

Open edge: Whether context and dynamics identifiability assumptions transfer beyond two CARL tasks or Atari-style visual control remains unverified.
Diffusion Transformer video world models with reconstruction or physics auxiliary objectives

Learns high-dimensional visual future prediction with mask reconstruction, depth, or keypoint dynamics to improve fidelity and physical plausibility.

Open edge: Cross-domain and zero-shot generalization evidence remains limited; long-horizon physical consistency beyond reported driving and robot datasets is unresolved.
Unified or compositional action-video interfaces

Couples action and video generation, or factorizes language into primitive video plans, enabling world models to serve as policies or planners.

Open edge: How far composition generalizes to novel primitives, unseen action spaces, or real-world execution beyond simulation and demonstration data is unclear.
Closed-loop consistency and action-control benchmark design

Provides evaluation protocols for spatial consistency, memory consistency, and action responsiveness rather than frame realism alone.

Open edge: Benchmark results and baselines are only partially reported; open-domain long-term memory and cross-action-space generalization remain unresolved.

Shared Direction

  • World-model learning is moving from latent recurrent state-space dynamics toward transformer and diffusion architectures for high-dimensional visual prediction.
  • Auxiliary self-supervised tasks are used to stabilize generation and encode geometry or physics, including mask reconstruction, depth prediction, and keypoint dynamics.
  • Conditioning on context, language, goals, or actions is treated as essential for control, planning, and zero-shot generalization.
  • Evaluation is expanding from visual quality to closed-loop consistency, action controllability, and spatial or temporal coherence.

Key Differences

  • Representation differs between compact stochastic latent states and explicit video/diffusion generation.
  • Supervision differs: some methods rely on reconstruction/generation, while others add physics-informed or action-coupling objectives.
  • Evaluation targets differ across Atari/CARL benchmarks, driving video datasets, robot manipulation datasets, and open-domain first/third-person videos.
  • Interaction mode differs: imagined rollouts for policy learning, decomposed language-goal video plans, unified action-video diffusion, and benchmark action-control loops.

Open Questions

  • Which auxiliary objectives transfer beyond the reported task families and datasets?
  • How much of reported policy or inference gains come from the learned dynamics rather than the planner, controller, or task prior?
  • Can these models maintain long-term memory and spatial consistency across unseen viewpoints, scenes, and action spaces?
  • Do current benchmarks isolate method contributions from surrounding system choices and dataset-specific metrics?
  • Several works lack visible code, limitation statements, or benchmark results, leaving reproducibility and failure boundaries as verification gaps.
CVAIworld modelsROworld modelLGvideo generationreinforcement learning

1250 papers

application

Domain applications of world models, such as robotics, driving, games, healthcare, or scientific simulation.

21 Citation impact 0 Citation velocity 64 Methodology quality 90 Recency

Literature review synthesis

Research Lines

Embodied and domain task transfer

Adapting world models to robot imagination, contextual control, and policy learning by conditioning video or dynamics generation on language, actions, or physical context; examples include RoboDreamer, UWM, cRSSM, and RoboScape.

Open edge: Cross-embodiment, cross-dataset, and real-world transfer remain under-evidenced; several evaluations are simulation-only or limited to a small number of tasks, and unseen primitives or partially observed context are unresolved.
Autonomous driving and simulation use cases

Building generative driving world models for long-horizon and multi-view scene prediction, including zero-shot evaluation on driving datasets; MaskGWM exemplifies this direction.

Open edge: The packet lacks evidence connecting simulator or dataset gains to real-world driving safety and operational outcomes, and the representative paper provides no visible limitation statement.
Interactive agent and game environments

Using model-based reinforcement learning to learn efficient action-conditioned world models that improve policy performance and sample efficiency in interactive tasks; STORM is a clear example.

Open edge: Long-term memory consistency, cross-action-space generalization, and robustness to rule or goal changes are not fully benchmarked, and limitation details are missing from the evidence.
World-model evaluation benchmarks

Creating datasets and metrics to quantify spatial consistency, memory consistency, and action control that generic video-generation metrics do not capture; LoopNav and MIND represent this line.

Open edge: These benchmarks identify current challenges but do not yet establish standardized links to downstream task success, real-world deployment, or broad baseline comparability.

Shared Direction

  • Application-oriented world models increasingly share a video generation or dynamics learning objective trained on visual-action sequences to support prediction or policy learning.
  • Controllability is a common focus, with conditioning on actions, language, context, or physics to ground generated futures in task-relevant behavior.
  • Dedicated evaluation is emerging as a contribution in its own right, targeting spatial coherence, memory, and action control rather than relying only on visual quality.

Key Differences

  • Representation differs: stochastic state-space models emphasize sample-efficient policy learning, while diffusion transformers emphasize high-dimensional visual generation.
  • Supervision differs: approaches use context-conditioned dynamics, compositional language-primitive generation, physics-informed auxiliary tasks, or multimodal video-action diffusion.
  • Evaluation target differs: some work optimizes Atari human-normalized performance, others driving video rollout metrics, robot simulation execution, or memory and action benchmark scores.
  • Deployment setting differs: offline video prediction, closed-loop interactive control, and simulated policy training make different assumptions about observability and environment feedback.

Open Questions

  • Does improved simulation or benchmark performance transfer to real robot and driving deployments?
  • Can world models generalize to unseen physical contexts, action spaces, or object-action compositions beyond training distributions?
  • How should long-horizon spatial consistency, memory consistency, and action controllability be jointly improved and measured?
  • What failure modes and boundary conditions remain hidden when application papers do not report limitation statements?
CVAIworld modelsROworld modelLGvideo generationreinforcement learning

895 papers

benchmark

Datasets, evaluation suites, metrics, stress tests, leaderboards, or benchmark studies.

22 Citation impact 0 Citation velocity 67 Methodology quality 91 Recency

Literature review synthesis

Research Lines

Memory consistency and action-control benchmark design

Creates open-domain closed-loop datasets and metrics for evaluating temporal stability, spatial consistency, and action fidelity in world models.

Open edge: Coverage of long-horizon memory, diverse action spaces, and alignment with downstream task utility remains uncertain.
Task-suite evaluation for model-based RL

Uses standardized control and game tasks to compare world-model-trained agents and zero-shot generalization.

Open edge: Evaluations are limited to few tasks or observable context assumptions; cross-implementation stability is not fully verified.
Visual and physical fidelity stress tests

Probes geometry, keypoint dynamics, and visual detail to distinguish plausible generation from surface-level realism.

Open edge: Proxy fidelity metrics are not yet tied to downstream policy or robot decision utility.
Unified multimodal world-model baselines

Provides reference architectures that couple video, action, or physics signals to support reproducible benchmark comparisons.

Open edge: Many entries report limited limitations, and robustness across datasets, compute budgets, and implementation details is under-evidenced.

Shared Direction

  • The field is converging on benchmarks that require more than aggregate video similarity, emphasizing memory, action control, and long-horizon behavior.
  • Open datasets with action annotations and open code are used to make world-model evaluation more reproducible.
  • Standard suites such as Atari 100k and CARL remain common for measuring agent performance, while newer driving, navigation, and robot datasets add task diversity.
  • Zero-shot or long-horizon generalization is a shared evaluation pressure, especially across unseen scenes, contexts, or action spaces.

Key Differences

  • Some work evaluates closed-loop memory faithfulness in open-domain videos, while other work evaluates task-level human-normalized scores or policy performance.
  • Evaluation signals differ: scene-graph spatial consistency, pixel-level generation quality, keypoint/depth plausibility, and RL reward are not interchangeable.
  • Interaction and deployment settings differ: first-person and third-person video control, action-free video pretraining, driving scenes, games, and robot tasks create different assumptions.
  • The role of observable context differs: some methods assume explicit context values, while others rely on visual or action-only data.

Open Questions

  • How should consistency and fidelity metrics be weighted against downstream decision utility, and what evidence would validate that alignment?
  • Do benchmark rankings remain stable across different implementations, compute budgets, and baseline reproductions?
  • How can memory and action-control benchmarks cover longer horizons and broader action spaces without overfitting to a fixed set of scenes?
  • What happens to zero-shot generalization claims when context values or environment properties are unobservable, and how should benchmarks report this boundary?
  • Several top papers lack visible limitation or cross-benchmark stability details, leaving failure modes and robustness as verification gaps.
CVworld modelsROAIworld modelLGvideo generationautonomous driving

450 papers

system

Runnable systems, platforms, simulators, frameworks, toolkits, or deployed pipelines.

22 Citation impact 1 Citation velocity 66 Methodology quality 91 Recency

Literature review synthesis

Research Lines

Action-conditioned world-model simulation stacks

Turns video or latent generative models into controllable simulators that produce future states or frames from action, trajectory, context, or physics cues for downstream policy use.

Open edge: Generalization to unseen environments, sensor setups, and long rollouts beyond the training datasets; evidence is often limited to nuScenes or two CARL tasks.
Real-time and interactive closed-loop serving

Addresses latency and interactivity constraints for online simulation, including sub-second single-GPU inference for closed-loop driving generation.

Open edge: Multi-GPU or complex-scene scalability, long-sequence consistency, and trade-offs between speed, fidelity, and controllability are not fully established.
Benchmark artifacts and reproducible world-model evaluation

Provides datasets, scores, or evaluation frameworks to make spatial consistency and action fidelity inspectable and comparable.

Open edge: Independent reproduction across teams and infrastructure, correlation with real downstream utility, and missing code or artifact details in several systems remain verification gaps.

Shared Direction

  • Action-conditioned generation is treated as the core mechanism for making world models usable as simulators, whether via latent dynamics or pixel-space video generation.
  • Benchmarks and datasets are shifting evaluation from frame quality alone toward controllability, consistency, and downstream task value.
  • Physics, context, temporal structure, or trajectory priors are increasingly added to improve plausibility and generalization.

Key Differences

  • Representation choices differ: some use recurrent state-space latent models, others use diffusion or frame-level autoregressive generation, and some jointly learn video plus depth or keypoint physics.
  • Supervision and conditioning vary across explicit context values, future trajectories, keypoint dynamics, depth, and high-level or low-level driving controls.
  • Evaluation domains and targets differ: spatial loop consistency in Minecraft, zero-shot policy transfer on CARL, robotic video fidelity, and closed-loop driving on nuScenes are not directly comparable.
  • Deployment settings differ: some systems prioritize batch benchmarking or policy training, while others emphasize interactive low-latency serving.

Open Questions

  • Does improved action fidelity or spatial consistency in benchmarks translate to better downstream policy learning and real deployment?
  • How do these world-model systems generalize beyond nuScenes or two CARL tasks to new sensors, dynamics, and interaction modalities?
  • Can closed-loop simulators preserve long-horizon geometric and temporal consistency under iterative self-conditioning without sacrificing sub-second latency?
  • What is the trade-off between dynamic consistency and fine-grained visual detail when lightweight control modules are added to pre-trained video generators?
  • Which missing code, dataset, or evaluation details are necessary for independent reproduction across teams and infrastructure?
CVAIworld modelsROworld modelLGvideo generationbenchmark

241 papers

theory

Theoretical analysis, formalization, guarantees, scaling laws, or conceptual foundations.

25 Citation impact 0 Citation velocity 67 Methodology quality 92 Recency

Literature review synthesis

Research Lines

Action-conditioned video dynamics for robot manipulation

Learning predictive action-to-video models that can serve as policies, forward dynamics, inverse dynamics, or planners on large robot datasets.

Open edge: Action-frame alignment, cross-embodiment generalization, adaptation to closed-source video diffusion, and whether generated dynamics match formal world-model assumptions in realistic settings.
Physics-aware alignment for embodied video world models

Post-training generated rollouts toward physical realism, task logic, embodiment plausibility, and visual quality through preference or reward modeling.

Open edge: Evaluation is concentrated on contact-rich manipulation, reward signals depend on subjective human preferences, and transfer to other domains remains undemonstrated.
Conceptual taxonomy and functional evaluation of world models

Clarifying world models through state construction, dynamics modeling, test-time verification, and training-independent benchmarks beyond visual fidelity.

Open edge: Whether taxonomies and verifiers change model design or evaluation, and whether current imagined views are calibrated enough to improve fine-grained reasoning.

Shared Direction

  • Learned action-conditioned video generation is treated as a promising substrate for embodied world models.
  • Large-scale robot-video pretraining or small-domain adaptation is a recurring ingredient across recent work.
  • Evaluation increasingly includes physical plausibility, action controllability, or downstream decision value rather than visual quality alone.
  • There is shared recognition that current models face persistence, causality, or information-bottleneck limitations.

Key Differences

  • Generative formulation differs: unified video-action diffusion, heterogeneous masked autoregression, adapters on fixed pretrained video diffusion, and flow-based reward post-training coexist.
  • Evaluation target differs across visual fidelity and controllability, physical realism and task logic, verifier-selected spatial reasoning, or policy-simulator correlation.
  • Action interaction differs through diffusion timesteps, frame-level action modules, learned intermediate masks, parallel spatial context blocks, or trajectory selection.
  • Data assumptions differ: some adapt closed-source video models with small action-labeled datasets, while others require millions of curated robot clips or zero-shot protocols.

Open Questions

  • Can formal definitions of state, dynamics, and persistence be linked to measurable behavior of current video generation models?
  • Will physics-aligned embodied world models generalize beyond contact-rich manipulation and the reported benchmarks?
  • Are test-time verifiers or learned reward models calibrated enough to provide reliable signals for spatial reasoning and policy improvement?
  • How should future benchmarks separate physical realism, action-conditioned controllability, and downstream decision usefulness?
AIworld modelsCVLGworld modelROreinforcement learningvideo generation

144 papers

survey

Surveys, taxonomies, tutorials, position papers, or roadmap papers.

22 Citation impact 1 Citation velocity 63 Methodology quality 91 Recency

Literature review synthesis

Research Lines

World-model taxonomy and roadmap synthesis

Maps fragmented generative world-model work into application scenarios, technical challenges, and future directions.

Open edge: Whether the synthesis is backed by systematic inclusion and evidence extraction rather than selective examples; coverage is explicitly incomplete in fast-moving areas.
Comparative evaluation synthesis

Summarizes how benchmarks, metrics, and evaluation choices such as closed-loop control, controllability, and long-horizon consistency shape conclusions across world-model papers.

Open edge: Whether results compared across simulation, physical robots, open-loop generation, and interactive domains are actually compatible under different supervision and metrics.
Open problem and gap mapping

Turns broad world-model literature coverage into concrete challenges such as out-of-sight dynamics, action-conditioned controllability, real-time interaction, and physical robot learning gaps.

Open edge: Prioritization of which gaps matter most for progress, and whether proposed taxonomies or evaluation prescriptions are empirically validated rather than conceptually proposed.

Shared Direction

  • Recent work treats world models as generative simulators that connect state representation, dynamics, and interaction, with emphasis on long-horizon consistency, controllability, and responsiveness beyond short visual realism.
  • Surveys and taxonomies converge on separating global or latent state representations from generation and control, while method papers instantiate this with compositional or persistent state designs.
  • Evaluation is recognized as a central bottleneck; the evidence points to domain-specific benchmarks, closed-loop metrics, and open-source artifacts as important levers for progress.
  • Physical robot learning and video-based simulation share a common world-model foundation of learning dynamics from interaction or video, despite differing supervision and deployment settings.

Key Differences

  • State representation differs: compositional language primitives and latent spaces versus persistent 3D/dynamic entity states versus multimodal sparse features conditioned on ego-trajectory and human pose.
  • Supervision and interaction signals differ: language instructions, goal images, ego-trajectories, human poses, sparse robot rewards, and unlabeled video training create distinct generalization and controllability assumptions.
  • Evaluation targets differ: surveys compare benchmark landscapes across application domains, while method papers emphasize generation quality, simulated robot execution, physical robot learning speed, or interactive latency without a shared metric.
  • Deployment settings differ: some work targets physical robots without simulators, while other work targets open-source real-time simulators or offline video planning, changing the available verification evidence.

Open Questions

  • How systematically were the survey taxonomies constructed, and what inclusion criteria or evidence extraction methods would make their gap claims verifiable?
  • Are benchmark comparisons across domains comparable, or do differences in evaluation settings and supervision signals prevent direct aggregation?
  • Can models trained on combinations of seen primitives generalize to novel primitives without external parsing or pretrained primitive models?
  • Which open challenges should be prioritized first: out-of-sight dynamics, physical robot adaptation, real-time interactivity, or evaluation standardization?
AIworld modelsCVLGROworld modelembodied AIvideo generation