Towards foundational LiDAR world models with efficient latent flow matching
TLDR
First systematic study of LiDAR world model transferability across domains, proposing a latent flow matching framework that achieves SOTA with higher compression and less data.
Reasoning
Strengths include novel domain transfer study and efficient latent CFM framework with strong empirical results. Weaknesses: limited to LiDAR data, no interactive or RL aspects, and abstract lacks details on limitations or failure cases.
Read-first score
Read-first score 43.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 28.
Field roles
Rank sensitivity
Stability: volatile; rank range: 380.
Keyword Scores
Deep Analysis
Innovations
- First systematic domain transfer study for LiDAR world models across outdoor-to-indoor, sparse-beam to dense-beam, and non-semantic to semantic transfer.
- Proposes a latent conditional flow matching (CFM)-based framework that achieves state-of-the-art reconstruction accuracy with half the training data and 6x higher compression ratio than prior methods.
- Demonstrates strong transferability: a single pre-trained model achieves up to 11% absolute improvement (83% relative) over training from scratch, outperforming in 30/36 comparisons.
- Outperforms previous semantic occupancy forecasting models with only 5% of the labeled training data required by prior models.
- Achieves state-of-the-art on future-trajectory-conditioned semantic occupancy forecasting with 23x computational efficiency (28x FPS speedup) and on semantic occupancy forecasting with 2x efficiency (1.1x FPS speedup).
Methodology
The paper proposes a latent conditional flow matching (CFM)-based framework for LiDAR world models. It conducts the first systematic domain transfer study across three scenarios: outdoor-to-indoor generalization, sparse-beam to dense-beam adaptation, and non-semantic to semantic transfer. The model is pre-trained and then fine-tuned with varying amounts of data, evaluated on reconstruction accuracy, compression ratio, and semantic occupancy forecasting tasks.
Key Results
A single pre-trained model achieves up to 11% absolute improvement (83% relative) over training from scratch and outperforms in 30/36 comparisons. The method achieves state-of-the-art performance on future-trajectory-conditioned semantic occupancy forecasting with 23x computational efficiency (28x FPS speedup) and on semantic occupancy forecasting with 2x efficiency (1.1x FPS speedup).