Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Social World Model-Augmented Mechanism Design Policy Learning

arXiv 25.10 2025 56.5 method

TLDR

Introduces SWM-AP, a social world model that infers agent traits and predicts responses to enhance mechanism design policy learning with improved sample efficiency.

Reasoning

The paper presents a novel method combining world models with mechanism design, demonstrating strong empirical results across diverse settings. However, the abstract lacks explicit discussion of limitations or comparison to state-of-the-art world model methods.

Read-first score

Read-first score 56.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
64.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,experiment

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 358.

Keyword Scores

world model
10
model-based reinforcement learning world model
9
world dynamics prediction
8
world simulator
7
interactive world model
6
generative world model
5
video world model
0

Deep Analysis

Innovations

  • Hierarchical modeling of agents' behavior by inferring persistent latent traits (e.g., skills, preferences) from interaction trajectories
  • Trait-based social world model that predicts agents' responses to deployed mechanisms
  • Online trait inference during real-world interactions to boost policy learning efficiency

Methodology

SWM-AP learns a social world model that hierarchically infers agents' latent traits from their interaction trajectories and uses a trait-based model to predict agents' responses to mechanisms. The mechanism design policy collects extensive training trajectories by interacting with the social world model, while concurrently inferring agents' traits online during real-world interactions to further enhance sample efficiency.

Key Results

In experiments on tax policy design, team coordination, and facility location, SWM-AP outperforms established model-based and model-free RL baselines in both cumulative rewards and sample efficiency.

Tags