Sekai: A Video Dataset towards World Exploration
TLDR
A large-scale first-person video dataset (5000+ hours, 100+ countries) with rich annotations for training video generation models for world exploration.
Reasoning
Strengths: massive scale, geographic diversity, and comprehensive annotations addressing key limitations of existing datasets. Weaknesses: no direct evaluation on world model tasks or interactive exploration; limited to first-person view and static video generation.
Read-first score
Read-first score 40.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 12.
Field roles
Rank sensitivity
Stability: volatile; rank range: 313.
Keyword Scores
Deep Analysis
Innovations
- Introduction of Sekai, a large-scale first-person view worldwide video dataset with over 5,000 hours of walking and drone footage from 100+ countries and 750 cities.
- Rich annotations including location, scene, weather, crowd density, captions, and camera trajectories, specifically designed for world exploration training.
- Development of an efficient and effective toolbox for collecting, preprocessing, and annotating videos at scale.
Methodology
The dataset comprises over 5,000 hours of first-person view (FPV) and unmanned aerial vehicle (UVA) videos sourced from over 100 countries and 750 cities. An efficient toolbox was developed to collect, preprocess, and annotate the videos with metadata such as location, scene, weather, crowd density, captions, and camera trajectories. The dataset's quality and utility are evaluated through comprehensive analyses and experiments on video generation models.
Key Results
Comprehensive analyses and experiments demonstrate the dataset's scale, diversity, annotation quality, and effectiveness for training video generation models, showing its potential to advance world exploration applications.