Social World Model-Augmented Mechanism Design Policy Learning
TLDR
Introduces SWM-AP, a social world model that infers agent traits and predicts responses to enhance mechanism design policy learning with improved sample efficiency.
Reasoning
The paper presents a novel method combining world models with mechanism design, demonstrating strong empirical results across diverse settings. However, the abstract lacks explicit discussion of limitations or comparison to state-of-the-art world model methods.
Read-first score
Read-first score 56.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.
Field roles
Rank sensitivity
Stability: volatile; rank range: 358.
Keyword Scores
Deep Analysis
Innovations
- Hierarchical modeling of agents' behavior by inferring persistent latent traits (e.g., skills, preferences) from interaction trajectories
- Trait-based social world model that predicts agents' responses to deployed mechanisms
- Online trait inference during real-world interactions to boost policy learning efficiency
Methodology
SWM-AP learns a social world model that hierarchically infers agents' latent traits from their interaction trajectories and uses a trait-based model to predict agents' responses to mechanisms. The mechanism design policy collects extensive training trajectories by interacting with the social world model, while concurrently inferring agents' traits online during real-world interactions to further enhance sample efficiency.
Key Results
In experiments on tax policy design, team coordination, and facility location, SWM-AP outperforms established model-based and model-free RL baselines in both cumulative rewards and sample efficiency.