Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Xray2Xray: World Model from Chest X-rays with Volumetric Context

arXiv 25.6 2025 49.5 method, application

TLDR

Xray2Xray learns a world model from chest X-rays to encode 3D volumetric context via transition dynamics, improving diagnosis and risk prediction.

Reasoning

The paper introduces a novel world model for chest X-rays that captures 3D structural information from 2D projections, with strong empirical results on risk prediction and disease diagnosis. However, it lacks interactive or reinforcement learning components, and the evaluation is limited to specific medical tasks.

Read-first score

Read-first score 49.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 29.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,metric,result

Topical relevance 42%
41.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 389.

Keyword Scores

world model
9
world dynamics prediction
8
generative world model
6
world simulator
2
video world model
2
interactive world model
1
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Introduces Xray2Xray, a World Model that learns latent representations encoding 3D structural information from 2D chest X-rays
  • Models transition dynamics of X-ray projections across different angular positions using a vision model and a transition model
  • Demonstrates that latent representations from the World Model outperform supervised and self-supervised methods for cardiovascular disease risk estimation and achieve competitive performance in multi-pathology classification

Methodology

Xray2Xray employs a vision model and a transition model to capture latent representations of the chest volume by modeling the transition dynamics of X-ray projections across different angular positions. These latent representations are then used for downstream tasks including cardiovascular disease risk estimation and classification of five pathologies. The quality of the latent representations is further assessed through synthesis tasks that reconstruct volumetric context.

Key Results

Xray2Xray outperformed both supervised methods and self-supervised pretraining methods for cardiovascular disease risk estimation, and achieved competitive performance in classifying five pathologies in chest X-rays. Additionally, the latent representations were shown to be capable of reconstructing volumetric context.

Tags