ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
TLDR
ShareVerse enables multi-agent consistent video generation for shared world modeling using spatial concatenation and cross-agent attention on CARLA data.
Reasoning
The paper presents a novel framework for multi-agent consistent video generation, with strengths in spatial concatenation and cross-agent attention for shared world consistency. However, it relies solely on simulated CARLA data and lacks real-world validation, limiting its generalizability.
Read-first score
Read-first score 48.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 331.
Keyword Scores
Deep Analysis
Innovations
- A dataset for large-scale multi-agent interactive world modeling built on the CARLA simulation platform, featuring diverse scenes, weather conditions, interactive trajectories, and paired multi-view videos (front/rear/left/right views per agent) with camera data.
- A spatial concatenation strategy for four-view videos of independent agents to model a broader environment and ensure internal multi-view geometric consistency.
- Integration of cross-agent attention blocks into a pretrained video model, enabling interactive transmission of spatial-temporal information across agents for shared world consistency in overlapping regions and reasonable generation in non-overlapping regions.
Methodology
ShareVerse builds a large-scale multi-agent interactive world modeling dataset using the CARLA simulator, with diverse scenes, weather conditions, and interactive trajectories paired with multi-view videos (four views per agent). It proposes a spatial concatenation strategy for four-view videos to model a broader environment and ensure geometric consistency, and integrates cross-agent attention blocks into a pretrained video model to enable interactive spatial-temporal information transmission across agents. The model supports 49-frame large-scale video generation.
Key Results
ShareVerse supports 49-frame large-scale video generation, accurately perceives the position of dynamic agents, and achieves consistent shared world modeling.