HarmonyDream: Task Harmonization Inside World Models
TLDR
HarmonyDream balances observation and reward modeling in world models for sample-efficient MBRL, achieving state-of-the-art on Atari 100K.
Reasoning
The paper provides a clear empirical investigation and insight into task imbalance in world models, proposing a simple yet effective dynamic loss adjustment method. Strengths include strong experimental results on visual robotic tasks and Atari benchmarks, but the approach is limited to simulated environments and may not generalize to real-world settings.
Read-first score
Read-first score 51.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.
Field roles
Rank sensitivity
Stability: volatile; rank range: 675.
Keyword Scores
Deep Analysis
Innovations
- Uncovering the overlooked potential of sample-efficient MBRL by mitigating the domination of either observation or reward modeling in world models
- Proposing HarmonyDream, a method that automatically adjusts loss coefficients to maintain a dynamic equilibrium between observation and reward modeling tasks
- Achieving new state-of-the-art results on the Atari 100K benchmark with 10%-69% absolute performance boosts on visual robotic tasks
Methodology
The authors conduct a dedicated empirical investigation to understand the roles of observation and reward modeling in world models. Based on insights that observation modeling struggles with environment complexity and limited capacity, while reward modeling lacks rich learning signals, they propose HarmonyDream, which automatically adjusts loss coefficients to maintain task harmonization. The method is evaluated on visual robotic tasks and the Atari 100K benchmark using a base MBRL method.
Key Results
HarmonyDream yields 10%-69% absolute performance improvements on visual robotic tasks and sets a new state-of-the-art result on the Atari 100K benchmark.