Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

arXiv 2026 63.3 method, system, application

TLDR

Embody4D transforms monocular robot videos into novel-view videos for embodied 4D world modeling via generative video-to-video world model.

Reasoning

Strengths include a novel data synthesis pipeline and geometric stability techniques; weaknesses are limited to visual benchmarks and unclear real-world impact beyond abstract mention.

Read-first score

Read-first score 63.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,experiment,metric

Topical relevance 29%
77.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation impact 18%
69.1

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.69139553

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 372.

Keyword Scores

world model
10
generative world model
10
video world model
10
interactive world model
8
world dynamics prediction
7
world simulator
6
model-based reinforcement learning world model
3

Deep Analysis

Innovations

  • 3D-aware compositional synthesis pipeline to curate a heterogeneous dataset compositing cross-embodiment robotic arms with diverse backgrounds
  • Latent confidence-aware expert modulation strategy that estimates reliability of warped latent priors and adaptively routes regions to copy, repair, or inpaint experts for spatiotemporally consistent 4D generation
  • Interaction-aware attention mechanism that explicitly attends to robotic interaction regions to enhance manipulation fidelity

Methodology

Embody4D is a dedicated video-to-video world model for embodied scenarios that transforms a monocular robot video into novel-view videos from flexible target camera viewpoints. It employs a 3D-aware compositional synthesis pipeline to generate training data, a latent confidence-aware expert modulation strategy for geometric stability, and an interaction-aware attention mechanism to improve manipulation fidelity.

Key Results

Embody4D achieves state-of-the-art performance on visual evaluation benchmarks, and both simulated and real-world robotic experiments demonstrate its effectiveness as a robust data engine for synthesizing high-fidelity, view-consistent videos that empower downstream robotic planning and learning.

Tags

embodied AIworld modelnovel view synthesisvideo generation4D modelingdata engineCV