Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
TLDR
LLM-based embodied agents use dialogue to align world models, reducing conflicts but harming task success; new metrics measure alignment gaps.
Reasoning
The paper introduces a novel framework for measuring world-model alignment via dialogue in multi-agent coordination, with empirical results across three LLMs. Strengths include clear metrics and real-world benchmark (PARTNR), but task success degradation and limited scope (household robotics) are weaknesses.
Read-first score
Read-first score 51.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 24.
Field roles
Rank sensitivity
Stability: volatile; rank range: 305.
Keyword Scores
Deep Analysis
Innovations
- Extending the PARTNR benchmark with a natural-language dialogue channel for two agents with partial observability
- Proposing a framework for measuring world-model alignment using three metrics: observation convergence, information novelty, and belief-sensitive messaging
- Empirically demonstrating that dialogue reduces action conflicts (40–83 percentage points) but degrades task success, revealing a gap between superficial coordination and genuine world-model alignment
Methodology
The authors extend the PARTNR benchmark for collaborative household robotics by adding a natural-language dialogue channel that allows two partially observable agents to communicate during task execution. They evaluate three LLM-based agents using proposed metrics for world-model alignment (observation convergence, information novelty, belief-sensitive messaging) alongside standard measures of action conflicts and task success.
Key Results
Dialogue reduces action conflicts by 40 to 83 percentage points compared to silent coordination, but degrades task success across all three LLMs tested.
Limitations
- Dialogue degrades task success relative to silent coordination, indicating current models fail to leverage communication effectively
- The study only evaluates three LLMs, limiting generalizability of findings
- The benchmark is restricted to household robotics tasks with two agents and partial observability, so results may not extend to other domains or larger teams
- The proposed metrics reveal a gap between superficial coordination and genuine world-model alignment, but the paper does not provide a solution to close this gap