Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Sekai: A Video Dataset towards World Exploration

arXiv 25.7 2025 40.8 benchmark

TLDR

A large-scale first-person video dataset (5000+ hours, 100+ countries) with rich annotations for training video generation models for world exploration.

Reasoning

Strengths: massive scale, geographic diversity, and comprehensive annotations addressing key limitations of existing datasets. Weaknesses: no direct evaluation on world model tasks or interactive exploration; limited to first-person view and static video generation.

Read-first score

Read-first score 40.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 12.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,experiment

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset,github

Topical relevance 42%
17.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 313.

Keyword Scores

video world model
4
world model
2
generative world model
2
world simulator
1
interactive world model
1
world dynamics prediction
1
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Introduction of Sekai, a large-scale first-person view worldwide video dataset with over 5,000 hours of walking and drone footage from 100+ countries and 750 cities.
  • Rich annotations including location, scene, weather, crowd density, captions, and camera trajectories, specifically designed for world exploration training.
  • Development of an efficient and effective toolbox for collecting, preprocessing, and annotating videos at scale.

Methodology

The dataset comprises over 5,000 hours of first-person view (FPV) and unmanned aerial vehicle (UVA) videos sourced from over 100 countries and 750 cities. An efficient toolbox was developed to collect, preprocess, and annotate the videos with metadata such as location, scene, weather, crowd density, captions, and camera trajectories. The dataset's quality and utility are evaluated through comprehensive analyses and experiments on video generation models.

Key Results

Comprehensive analyses and experiments demonstrate the dataset's scale, diversity, annotation quality, and effectiveness for training video generation models, showing its potential to advance world exploration applications.

Tags