World Models
TLDR
Generative neural network models of RL environments enable unsupervised learning and agent training in hallucinated dreams.
Reasoning
Strengths include unsupervised learning and dream training for efficient policy transfer. Weaknesses: limited to simulated environments, no real-world validation.
Read-first score
Read-first score 52.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 554.
Keyword Scores
Deep Analysis
Innovations
- Unsupervised learning of a compressed spatial and temporal representation of reinforcement learning environments using generative neural network models
- Training a compact and simple policy using features extracted from the world model
- Training an agent entirely inside its own hallucinated dream generated by the world model and transferring the policy back to the actual environment
Methodology
The paper proposes building generative neural network models of popular reinforcement learning environments. The world model is trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation. Features extracted from this world model are then used as inputs to train a compact and simple policy agent.
Key Results
The approach enables training a very compact and simple policy that can solve the required task. The agent can even be trained entirely inside its own hallucinated dream generated by the world model, and the learned policy successfully transfers back to the actual environment.