Conformal Orbit-Valid Trust Horizons for Equivariant World Models
TLDR
Certifies trust horizons for equivariant world models using conformal prediction, showing orbit-constant rollout errors and zero violations in audits.
Reasoning
The paper introduces a novel conformal calibration method for trust-horizon certification in equivariant world models, with strong theoretical results on orbit invariance. However, the empirical evaluation is limited to symmetric 2D and 3D substrates, and the practical applicability to complex real-world scenarios remains unclear.
Read-first score
Read-first score 56.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 26.
Field roles
Rank sensitivity
Stability: volatile; rank range: 389.
Keyword Scores
Deep Analysis
Innovations
- Conformal orbit-valid trust horizons for equivariant world models
- Split-conformal multiplicative factor calibration of raw horizon curves
- Structural result: exact equivariance transports calibrated trust-horizon curve over group orbit
- Certificate-level calibration-cost study revealing two complementary regimes
Methodology
The method forms a raw horizon curve from a one-step latent residual and a finite-time expansion estimate, then calibrates it using a split-conformal multiplicative factor. Evaluation is performed on reproducible audit sets and stable audits across symmetric 2D and 3D yaw environments, comparing equivariant, plain, and augmented models. Metrics include violation rates, orbit-transport residuals, and certified-to-measured horizon ratios.
Key Results
The conformal factor γ_α=1.0 on the reproducible audit set, with zero anti-conservative violations across 50 stable audits (exact-binomial 95% upper bound 5.8%). Orbit-transport residuals are small (median 1.1%, max 4.1% over 14 orbit audits), and the certificate is non-vacuous (median certified-to-measured horizon ratio 0.67). On a 3D yaw audit, the equivariant model achieves a one-sector safe and non-vacuous orbit-valid certificate, while non-equivariant baselines incur violation, slack, sharpness, or additional-sector costs.
Limitations
- The certificate is a conservative, distributional audit rather than a global reachability guarantee
- Certificate-guided subgoal spacing is not confirmed in the current 3D CEM-MPC behavior layer