Xray2Xray: World Model from Chest X-rays with Volumetric Context
TLDR
Xray2Xray learns a world model from chest X-rays to encode 3D volumetric context via transition dynamics, improving diagnosis and risk prediction.
Reasoning
The paper introduces a novel world model for chest X-rays that captures 3D structural information from 2D projections, with strong empirical results on risk prediction and disease diagnosis. However, it lacks interactive or reinforcement learning components, and the evaluation is limited to specific medical tasks.
Read-first score
Read-first score 49.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 29.
Field roles
Rank sensitivity
Stability: volatile; rank range: 389.
Keyword Scores
Deep Analysis
Innovations
- Introduces Xray2Xray, a World Model that learns latent representations encoding 3D structural information from 2D chest X-rays
- Models transition dynamics of X-ray projections across different angular positions using a vision model and a transition model
- Demonstrates that latent representations from the World Model outperform supervised and self-supervised methods for cardiovascular disease risk estimation and achieve competitive performance in multi-pathology classification
Methodology
Xray2Xray employs a vision model and a transition model to capture latent representations of the chest volume by modeling the transition dynamics of X-ray projections across different angular positions. These latent representations are then used for downstream tasks including cardiovascular disease risk estimation and classification of five pathologies. The quality of the latent representations is further assessed through synthesis tasks that reconstruct volumetric context.
Key Results
Xray2Xray outperformed both supervised methods and self-supervised pretraining methods for cardiovascular disease risk estimation, and achieved competitive performance in classifying five pathologies in chest X-rays. Additionally, the latent representations were shown to be capable of reconstructing volumetric context.