Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

arXiv 2026 55.8 method

TLDR

Proposes a driving world model predicting latent BEV features for long-horizon future state modeling, achieving SOTA on Bench2drive.

Reasoning

The paper introduces a novel approach for long-horizon world modeling in autonomous driving by predicting latent semantic features in BEV space, which is a clear strength. However, the evaluation is limited to a single closed-loop benchmark (Bench2drive) and lacks real-world driving tests, and the method's reliance on BEV space may limit generalizability.

Read-first score

Read-first score 55.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 28.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Reproducibility 18%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Citation impact 18%
75.4

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.75435134

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,result

Topical relevance 29%
40

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 474.

Keyword Scores

world model
9
world dynamics prediction
8
generative world model
3
model-based reinforcement learning world model
3
world simulator
2
video world model
2
interactive world model
1

Deep Analysis

Innovations

  • Parallel prediction of latent semantic features for consecutive future frames in bird's-eye-view (BEV) space for long-horizon world modeling
  • Efficient and adaptive text reasoning mechanism that utilizes social knowledge and reasoning capabilities to improve driving performance in long-tail scenarios

Methodology

The paper proposes a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in BEV space, enabling long-horizon modeling of future world states. It also introduces an efficient and adaptive text reasoning mechanism that leverages additional social knowledge and reasoning capabilities. The approach is evaluated on the closed-loop Bench2drive benchmark, achieving state-of-the-art results.

Key Results

The proposed method achieves state-of-the-art results on the closed-loop Bench2drive benchmark.

Tags

autonomous drivingworld modelbird's-eye-viewlatent state predictionvision-language modelend-to-end drivingCVRO