Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
TLDR
Proposes Audio-Visual World Models (AVWM) integrating binaural audio and visual dynamics, with a benchmark and a diffusion transformer model for multimodal prediction.
Reasoning
The paper introduces a novel formulation and benchmark for audio-visual world modeling, addressing a gap in multisensory simulation. Strengths include a unified framework and a new dataset; weaknesses are that the abstract cuts off before full results and practical validation details are limited.
Read-first score
Read-first score 45.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 493.
Keyword Scores
Deep Analysis
Innovations
- Unified formulation of Audio-Visual World Models (AVWM) as a partially observable Markov decision process with synchronized audio-visual observations.
- Construction of AVW-4k, a controlled benchmark comprising 30 hours of binaural audio-visual trajectories with action annotations across 76 indoor environments.
- Proposal of AV-CDiT, an Audio-Visual Conditional Diffusion Transformer with a novel modality expert architecture and three-stage training strategy for effective multimodal integration.
Methodology
The paper presents a unified formulation for audio-visual world modeling under low-level action control, casting multimodal environment simulation as a POMDP with synchronized audio-visual observations. They construct the AVW-4k benchmark with 30 hours of binaural audio-visual trajectories and action annotations across 76 indoor environments. They propose AV-CDiT, a conditional diffusion transformer with a modality expert architecture, optimized through a three-stage training strategy for multimodal integration. Evaluation is performed on the benchmark and in embodied navigation tasks.
Key Results
AV-CDiT achieves high-fidelity multimodal prediction across visual and auditory modalities on the AVW-4k benchmark. It also improves a vision-language-model-guided agent in continuous audio-visual navigation.