Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Beyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and Planning

arXiv 25.11 2025 60.8 method

TLDR

Survey of world models for clinical prediction, counterfactuals, and planning, introducing a capability rubric and identifying gaps in reliability and validation.

Reasoning

Strengths: comprehensive survey across healthcare domains, clear capability rubric (L1-L4), and identification of critical gaps. Weaknesses: review paper without new experiments, limited novelty, and reliance on existing works for validation.

Read-first score

Read-first score 60.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,validation

Topical relevance 42%
68.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 263.

Keyword Scores

world model
10
world dynamics prediction
9
world simulator
7
model-based reinforcement learning world model
7
generative world model
6
interactive world model
6
video world model
3

Deep Analysis

Innovations

  • Introduction of a capability rubric (L1-L4) for world models in healthcare: L1 temporal prediction, L2 action-conditioned prediction, L3 counterfactual rollouts, L4 planning/control
  • Identification of cross-cutting gaps limiting clinical reliability: under-specified action spaces and safety constraints, weak interventional validation, incomplete multimodal state construction, limited trajectory-level uncertainty calibration
  • Outline of a research agenda for clinically robust prediction-first world models integrating generative backbones with causal/mechanical foundations

Methodology

This paper conducts a literature review of world models in healthcare, surveying recent work across three domains: medical imaging and diagnostics, disease progression modeling from electronic health records, and robotic surgery and surgical planning. It introduces a capability rubric (L1-L4) to categorize systems and identifies cross-cutting gaps.

Key Results

Most reviewed systems achieve L1 (temporal prediction) and L2 (action-conditioned prediction), with fewer instances of L3 (counterfactual rollouts) and rare L4 (planning/control).

Limitations

  • Under-specified action spaces and safety constraints
  • Weak interventional validation
  • Incomplete multimodal state construction
  • Limited trajectory-level uncertainty calibration

Tags