MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
TLDR
MaineCoon is a 22B parameter real-time audio-visual autoregressive model for social world modeling, achieving 47.5 FPS on a single GPU.
Reasoning
The paper introduces a novel social world model with innovative training techniques and real-time interaction capabilities, addressing a gap in human-centric social dynamics. However, the abstract lacks explicit details on real-world evaluation or benchmarks, and the concept of social world models is not yet validated against existing standards.
Read-first score
Read-first score 54.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 484.
Keyword Scores
Deep Analysis
Innovations
- First real-time audio-visual autoregressive model with 22B parameters for social-interactive applications
- Self-resampling technique for efficient training
- Cross-modal representation alignment
- Domain-aware preference optimization
- Reinforced online-policy distillation (ROPD)
- Agentic streaming inference framework with agentic cache management and prompt planning for thousand-second-scale generation
Methodology
MaineCoon is a 22B-parameter real-time audio-visual autoregressive model designed for streaming generation and sub-second interaction. Training employs novel techniques including self-resampling, cross-modal representation alignment, domain-aware preference optimization, and reinforced online-policy distillation. Inference uses an agentic streaming framework with cache management and prompt planning to mitigate drift over long horizons.
Key Results
Achieves a record-breaking frame rate of up to 47.5 FPS on a single GPU, setting a new state-of-the-art benchmark for high-quality, low-latency, and long-horizon audio-visual autoregressive models.