Orbis 2: A Hierarchical World Model for Driving
TLDR
A hierarchical driving world model with two-level prediction and diffusion forcing pretraining achieves state-of-the-art results on benchmarks.
Reasoning
The paper introduces a novel hierarchical world model for driving that factorizes prediction across two abstraction levels, and a two-stage training paradigm combining diffusion and teacher forcing. Strengths include clear methodology, strong empirical results on standard benchmarks, and improved internal representations. Weaknesses are domain specificity to driving and lack of explicit discussion of limitations or generalizability.
Read-first score
Read-first score 47.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 499.
Keyword Scores
Deep Analysis
Innovations
- Hierarchical world model with two levels: high-level predictor for coarse scene structure over long horizons, and low-level generator for detailed predictions conditioned on high-level output.
- Diffusion forcing objective for pretraining, yielding richer internal representations compared to teacher forcing.
- Two-stage training paradigm: pretrain with diffusion forcing, fine-tune with teacher forcing to combine representational benefits and rollout stability.
- State-of-the-art performance on driving world model benchmarks including long-horizon generation, steering responsiveness, and representation quality.
Methodology
A hierarchical world model factorizes future prediction into high-level (coarse, long-horizon) and low-level (detailed, conditioned) components. Training uses a two-stage approach: diffusion forcing pretraining followed by teacher forcing fine-tuning. The model is evaluated on standard driving world model benchmarks for generation fidelity, steering responsiveness, and internal representation quality.
Key Results
The method achieves state-of-the-art results on long-horizon generation fidelity, steering responsiveness in counterfactual scenarios, and internal representation quality across established driving world model benchmarks.