Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Prisma-World: Camera-Controllable Multi-Agent Video World Model

arXiv 2026 61.9 method

TLDR

Prisma-World is a camera-controllable multi-agent video world model that ensures cross-view consistency via geometry-aware denoising and a new synthetic dataset.

Reasoning

The paper introduces a novel approach for multi-agent video generation with cross-view consistency, leveraging geometry-aware attention and a curriculum training strategy. Its main weakness is reliance on a synthetic dataset (UE5) without real-world validation, and it does not address interactive agent actions beyond camera control.

Read-first score

Read-first score 61.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
95.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.95211968

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,evaluation,experiment

Topical relevance 29%
68.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 451.

Keyword Scores

world model
10
video world model
10
generative world model
9
world simulator
8
world dynamics prediction
7
interactive world model
3
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Formulates multi-agent generation as a joint geometry-aware denoising process for cross-view consistency
  • Processes all agent videos within one full-attention sequence
  • Multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates
  • Injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence
  • Overlap-decaying curriculum training paradigm
  • Minimap-conditioned structural guidance
  • PrismaDataset: large-scale UE5 dataset with panoramic acquisition, composable multi-agent view groups, and precise camera/action annotations

Methodology

Prisma-World processes all agent videos within one full-attention sequence, uses a multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates, and injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence. It also employs an overlap-decaying curriculum training paradigm and minimap-conditioned structural guidance to strengthen multi-view consistency and global spatial perception.

Key Results

A single Prisma-World model generates high-fidelity multi-agent videos with flexible agent numbers, camera controllability, improved cross-view consistency, and spatial grounding under minimap guidance.

Tags

video world modelsmulti-agentcamera controlcross-view consistencygeometry-aware denoisingCV