Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
TLDR
Foresight detects failures in long-horizon robotic manipulation using action-conditioned world model latents and conformal prediction, validated in simulation and real robots.
Reasoning
The paper addresses an underexplored problem with a novel framework that leverages world model embeddings for failure detection, requiring only final task labels. Strengths include real-world validation and adaptive threshold calibration; weaknesses are limited detail on limitations and comparison baselines in the abstract.
Read-first score
Read-first score 53.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 27.
Field roles
Rank sensitivity
Stability: volatile; rank range: 387.
Keyword Scores
Deep Analysis
Innovations
- Using action-conditioned world model latents as a unified representation for failure detection across different policies
- Training failure detection with only final task-level success/failure labels, eliminating need for dense temporal annotations
- Adaptive threshold calibration via functional conformal prediction (FCP) for robust detection
Methodology
Foresight trains an action-conditioned world model to produce latent embeddings from manipulation trajectories. These embeddings are used as input to a failure detector that is trained solely on final task-level success/failure labels. Detection thresholds are adaptively calibrated using functional conformal prediction to handle varying conditions.
Key Results
Foresight outperforms state-of-the-art failure detection methods across multiple simulation benchmarks (LIBERO-Long, ManiSkill-Long, BEHAVIOR-1K) and is validated on real robots (ReactorX-200 arm and Franka arm) on long-horizon tasks.
Limitations
- Requires training an action-conditioned world model, which may be computationally expensive and data-intensive
- Performance may depend on the quality and coverage of the world model embeddings
- Evaluation is limited to specific simulation environments and a small set of real-robot tasks