CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
TLDR
End-to-end autonomous driving models trained with imitation learning (IL) often generalize poorly, particularly in long-tail scenarios where expert demonstrations are sparse.
Reasoning
Fallback reasoning generated from available title and abstract metadata: End-to-end autonomous driving models trained with imitation learning (IL) often generalize poorly, particularly in long-tail scenarios where expert demonstrations are sparse. Reinforcement learning (RL) can provide complementary task-level supervision, but applying RL to real-world autonomous driving is challenging...
Read-first score
Read-first score 52.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 485.
Deep Analysis
Innovations
- Competitive dual-policy framework integrating IL and RL under unified offline training
- Decoupling imitation and reward optimization into separate actors to alleviate objective conflicts
- Imagined future rollouts for long-horizon reward estimation in latent world models
- Competition mechanism that selectively transfers beneficial behaviors while keeping RL anchored to expert-like driving
Methodology
CoIRL-AD uses a competitive dual-policy framework with separate actors for imitation and reward optimization, operating in latent world models. It employs imagined future rollouts for long-horizon reward estimation and a competition mechanism to selectively transfer beneficial behaviors while anchoring RL to expert-like driving. The model is trained offline on the nuScenes dataset without interactive simulators.
Key Results
Experiments on nuScenes benchmark show consistent improvement in robustness over strong IL-based baselines, with especially large gains in cross-city generalization and long-tail scenarios.