Multiplayer Interactive World Models with Representation Autoencoders
TLDR
First multiplayer world model using latent diffusion, trained on 10k hours of Rocket League, generates stable 4-player matches in real time.
Reasoning
Strengths include novel multiplayer conditioning, large-scale training, and stable long-horizon rollouts. Weaknesses are domain specificity to Rocket League and lack of explicit model-based RL integration or generalization evidence.
Read-first score
Read-first score 49, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.
Field roles
Rank sensitivity
Stability: volatile; rank range: 534.
Keyword Scores
Deep Analysis
Innovations
- First multiplayer world model for highly dynamic environments with complex physical interactions
- Conditioning on action streams of multiple agents with attribution of scene changes to the correct player
- Stable long-horizon rollouts from short-clip training, maintaining distributional quality for minutes to hours
- Real-time generation of four-player matches using a 5B-parameter latent diffusion model
Methodology
A 5-billion-parameter latent diffusion model trained on 10,000 hours of Rocket League gameplay from publicly available bots, generating four-player matches at 20 fps on a single Nvidia B200 GPU. The study systematically investigates video codec, generative objective, multiplayer conditioning scheme, and scaling behavior.
Key Results
Rollouts stay stable far beyond the training horizon, with distributional quality holding steady for at least five minutes and observed rollouts continuing for hours without collapse. Targeted evaluations probe physical understanding beyond visual appearance.
Limitations
- Persistent failure modes remain despite scaling model and data
- Evaluation limited to Rocket League environment