OccSora: 4D Occupancy Generation Models as World Simulators for Autonomous Driving
TLDR
OccSora is a diffusion-based 4D occupancy generation model that simulates driving scenes as a world simulator for autonomous driving.
Reasoning
The paper introduces a novel diffusion-based approach for long-term 4D occupancy generation, addressing inefficiencies of autoregressive models. Strengths include trajectory-conditioned generation and temporal consistency, but it is limited to occupancy representation and evaluation on a single dataset (nuScenes).
Read-first score
Read-first score 72.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.
Field roles
Rank sensitivity
Stability: volatile; rank range: 84.
Keyword Scores
Deep Analysis
Innovations
- Diffusion-based 4D occupancy generation model for autonomous driving world simulation
- 4D scene tokenizer for compact discrete spatial-temporal representations of long-sequence occupancy videos
- Trajectory-prompt conditioned generation enabling controllable 4D occupancy video synthesis
- Ability to generate 16-second videos with authentic 3D layout and temporal consistency
Methodology
OccSora employs a 4D scene tokenizer to compress 4D occupancy input into compact discrete spatial-temporal tokens, enabling high-quality reconstruction of long-sequence occupancy videos. A diffusion transformer is then trained on these tokens to generate 4D occupancy conditioned on a trajectory prompt. The model is evaluated on the nuScenes dataset with Occ3D occupancy annotations.
Key Results
OccSora generates 16-second occupancy videos with authentic 3D layout and temporal consistency, demonstrating its ability to understand spatial and temporal distributions of driving scenes.