Prisma-World: Camera-Controllable Multi-Agent Video World Model
TLDR
Prisma-World is a camera-controllable multi-agent video world model that ensures cross-view consistency via geometry-aware denoising and a new synthetic dataset.
Reasoning
The paper introduces a novel approach for multi-agent video generation with cross-view consistency, leveraging geometry-aware attention and a curriculum training strategy. Its main weakness is reliance on a synthetic dataset (UE5) without real-world validation, and it does not address interactive agent actions beyond camera control.
Read-first score
Read-first score 61.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 451.
Keyword Scores
Deep Analysis
Innovations
- Formulates multi-agent generation as a joint geometry-aware denoising process for cross-view consistency
- Processes all agent videos within one full-attention sequence
- Multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates
- Injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence
- Overlap-decaying curriculum training paradigm
- Minimap-conditioned structural guidance
- PrismaDataset: large-scale UE5 dataset with panoramic acquisition, composable multi-agent view groups, and precise camera/action annotations
Methodology
Prisma-World processes all agent videos within one full-attention sequence, uses a multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates, and injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence. It also employs an overlap-decaying curriculum training paradigm and minimap-conditioned structural guidance to strengthen multi-view consistency and global spatial perception.
Key Results
A single Prisma-World model generates high-fidelity multi-agent videos with flexible agent numbers, camera controllability, improved cross-view consistency, and spatial grounding under minimap guidance.