Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
TLDR
A training-free inference acceleration framework for interactive video world models using adaptive context management and sparse attention.
Reasoning
The paper presents a practical solution to a key bottleneck in interactive video world models, with clear methodology and empirical speedup. However, evaluation is limited to two synthetic benchmarks, and real-world applicability is not demonstrated.
Read-first score
Read-first score 59.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 485.
Keyword Scores
Deep Analysis
Innovations
- Adaptive context management that discards spatial memory during novel exploration and adjusts temporal context based on local latent dynamics
- Denoising cache acceleration that reuses early-step model outputs when the camera revisits familiar regions
- Hardware-software co-designed 3D block sparse attention with fused Triton kernels
Methodology
Light Interaction is a training-free inference acceleration framework for interactive video world models. It combines adaptive context management, denoising cache acceleration, and 3D block sparse attention with fused Triton kernels. The framework is evaluated on HY-WorldPlay and Matrix-Game-3.0 datasets, measuring speedup and visual quality.
Key Results
Light Interaction achieves up to 2.59x speedup without model retraining while maintaining competitive visual quality on HY-WorldPlay and Matrix-Game-3.0.