Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

InfinityDrive: Breaking Time Limits in Driving World Models

arXiv 24.12 2024 59.5 method, application

TLDR

InfinityDrive is a driving world model that generates minute-scale, high-fidelity video with consistent spatial-temporal coherence.

Reasoning

The paper introduces a novel driving world model capable of generating over 2 minutes of high-resolution video, addressing the key limitation of short time windows in existing models. Strengths include the spatio-temporal co-modeling and memory mechanisms for long-term consistency, but the abstract lacks detailed quantitative comparisons and does not discuss interactive or reinforcement learning aspects.

Read-first score

Read-first score 59.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Topical relevance 42%
64.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,experiment

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset,github

Field roles

Candidate

Rank sensitivity

Stability: volatile; rank range: 161.

Keyword Scores

world model
10
video world model
10
generative world model
9
world dynamics prediction
8
world simulator
5
interactive world model
2
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • First driving world model with minute-scale video generation (over 1500 frames, more than 2 minutes)
  • Efficient spatio-temporal co-modeling module for high-resolution (576×1024) video generation with consistent spatial and temporal coherence
  • Extended temporal training strategy to enable long video generation
  • Memory injection and retention mechanisms to minimize cumulative errors
  • Adaptive memory curve loss for consistent long-term video generation

Methodology

InfinityDrive introduces an efficient spatio-temporal co-modeling module paired with an extended temporal training strategy to generate high-resolution driving videos. It incorporates memory injection and retention mechanisms along with an adaptive memory curve loss to reduce cumulative errors, enabling consistent video generation over 1500 frames. The model is evaluated on multiple datasets for fidelity, consistency, and diversity.

Key Results

InfinityDrive achieves state-of-the-art performance in high fidelity, consistency, and diversity, generating minute-scale driving videos (over 1500 frames, more than 2 minutes) at 576×1024 resolution. Comprehensive experiments on multiple datasets validate its ability to generate complex and varied scenarios.

Tags