SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation
TLDR
SWoMo is a neuro-symbolic world model for cataract surgery simulation that combines rule-based motion dynamics with diffusion-based visual realism.
Reasoning
Strengths include a novel decoupling of motion and visual generation, an inverse pairing strategy for sim-to-real translation, and demonstrated generalization and downstream task improvements. Weaknesses are the domain specificity to cataract surgery and potential scalability limits of the rule-based simulator.
Read-first score
Read-first score 60.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 476.
Keyword Scores
Deep Analysis
Innovations
- Decoupling motion generation from visual realism in surgical simulation
- Symbolic component with rule-based simulator and scene graph representations for motion dynamics and tool-tissue interactions
- Diffusion model for realistic visual appearance including textures and tissue deformations
- Inverse pairing strategy to reconstruct real surgical videos in the simulator for paired simulated-real data
- Sim-to-real translation via video diffusion model trained on paired data
Methodology
SWoMo employs a neuro-symbolic architecture: a symbolic rule-based simulator with scene graphs models motion dynamics and tool-tissue interactions, while a diffusion model generates realistic visual appearance. An inverse pairing strategy reconstructs real surgical videos in the simulator to create paired simulated and real videos, which are then used to train a video diffusion model for sim-to-real translation.
Key Results
SWoMo achieves qualitative and quantitative improvements over prior work, demonstrating generalization to unseen interaction geometries, improvements in downstream phase detection, and unsupervised video style transfer.