LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
TLDR
LiDARCrafter generates and edits dynamic 4D LiDAR sequences from natural language using a tri-branch diffusion network, achieving state-of-the-art performance on nuScenes.
Reasoning
The paper presents a novel framework for LiDAR-based world modeling with strong controllability and temporal coherence, supported by a comprehensive benchmark. However, it is limited to LiDAR modality and evaluated only on nuScenes, lacking comparison to video-based methods.
Read-first score
Read-first score 63.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.
Field roles
Rank sensitivity
Stability: volatile; rank range: 333.
Keyword Scores
Deep Analysis
Innovations
- Unified framework for 4D LiDAR generation and editing from free-form natural language inputs
- Parsing instructions into ego-centric scene graphs to condition a tri-branch diffusion network
- Tri-branch diffusion network generating object structures, motion trajectories, and geometry
- Autoregressive module for temporally coherent 4D LiDAR sequences with smooth transitions
- Comprehensive benchmark with diverse metrics spanning scene-, object-, and sequence-level aspects
Methodology
LiDARCrafter uses a tri-branch diffusion network conditioned on ego-centric scene graphs parsed from natural language instructions to generate object structures, motion trajectories, and geometry. An autoregressive module ensures temporal coherence across 4D LiDAR sequences. The model is evaluated on the nuScenes dataset using a newly established benchmark with scene-, object-, and sequence-level metrics.
Key Results
LiDARCrafter achieves state-of-the-art performance in fidelity, controllability, and temporal consistency across all levels on the nuScenes dataset.