Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

arXiv 25.11 2025 60.7 method, system

TLDR

GigaWorld-0 is a unified world model framework that generates diverse, physically realistic embodied data to train Vision-Language-Action models, achieving strong real-world robot performance.

Reasoning

Strengths include a novel integration of video and 3D generation for scalable data synthesis, with real-world validation on physical robots. Weaknesses are limited discussion of limitations and potential computational costs despite efficiency claims.

Read-first score

Read-first score 60.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
74.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,metric

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 370.

Keyword Scores

world model
10
generative world model
9
video world model
9
world simulator
8
world dynamics prediction
8
interactive world model
5
model-based reinforcement learning world model
3

Deep Analysis

Innovations

  • Unified world model framework designed as a data engine for Vision-Language-Action (VLA) learning
  • Integration of GigaWorld-0-Video for large-scale video generation with fine-grained control of appearance, camera viewpoint, and action semantics
  • Integration of GigaWorld-0-3D combining 3D generative modeling, 3D Gaussian Splatting reconstruction, physically differentiable system identification, and executable motion planning for geometric consistency and physical realism
  • Joint optimization of video and 3D components to synthesize embodied interaction data that is visually compelling, spatially coherent, physically plausible, and instruction-aligned
  • Efficient GigaTrain framework exploiting FP8-precision and sparse attention to reduce memory and compute requirements for training at scale

Methodology

GigaWorld-0 consists of two synergistic components: GigaWorld-0-Video for generating diverse, texture-rich, temporally coherent embodied sequences under fine-grained control, and GigaWorld-0-3D for ensuring geometric consistency and physical realism via 3D generative modeling, Gaussian Splatting, differentiable system identification, and motion planning. Their joint optimization enables scalable synthesis of embodied interaction data, and training is made feasible through the GigaTrain framework using FP8-precision and sparse attention.

Key Results

GigaWorld-0 generates high-quality, diverse, and controllable data across multiple dimensions. A VLA model (GigaBrain-0) trained solely on GigaWorld-0-generated data achieves strong real-world performance, significantly improving generalization and task success on physical robots without any real-world interaction during training.

Tags