OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction
TLDR
OccTENS proposes a generative occupancy world model using temporal next-scale prediction for efficient, controllable long-term 3D scene generation.
Reasoning
The paper introduces a novel reformulation of occupancy world modeling as temporal next-scale prediction, addressing inefficiency, temporal degradation, and controllability. Strengths include a clear methodology (TensFormer, pose aggregation) and empirical outperformance over SOTA. Weaknesses: abstract lacks explicit mention of real-world datasets or benchmarks, though experiments imply standard evaluation.
Read-first score
Read-first score 48.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 36.
Field roles
Rank sensitivity
Stability: volatile; rank range: 446.
Keyword Scores
Deep Analysis
Innovations
- Reformulating occupancy world model as a temporal next-scale prediction (TENS) task, decomposing temporal sequence modeling into spatial scale-by-scale generation and temporal scene-by-scene prediction.
- TensFormer architecture that effectively manages temporal causality and spatial relationships of occupancy sequences in a flexible and scalable way.
- Holistic pose aggregation strategy for unified sequence modeling of occupancy and ego-motion, enhancing pose controllability.
Methodology
OccTENS reformulates the occupancy world model as a temporal next-scale prediction (TENS) task, decomposing temporal sequence modeling into spatial scale-by-scale generation and temporal scene-by-scene prediction using a TensFormer architecture. It also incorporates a holistic pose aggregation strategy for unified sequence modeling of occupancy and ego-motion to enhance controllability.
Key Results
OccTENS outperforms the state-of-the-art method with both higher occupancy quality and faster inference time.