Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

arXiv 2026 60.3 method, application

TLDR

SWoMo is a neuro-symbolic world model for cataract surgery simulation that combines rule-based motion dynamics with diffusion-based visual realism.

Reasoning

Strengths include a novel decoupling of motion and visual generation, an inverse pairing strategy for sim-to-real translation, and demonstrated generalization and downstream task improvements. Weaknesses are the domain specificity to cataract surgery and potential scalability limits of the rule-based simulator.

Read-first score

Read-first score 60.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
81.6

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.81615198

Topical relevance 29%
78.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment

Reproducibility 18%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,github

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 476.

Keyword Scores

world model
10
world simulator
9
video world model
9
generative world model
8
world dynamics prediction
8
interactive world model
7
model-based reinforcement learning world model
4

Deep Analysis

Innovations

  • Decoupling motion generation from visual realism in surgical simulation
  • Symbolic component with rule-based simulator and scene graph representations for motion dynamics and tool-tissue interactions
  • Diffusion model for realistic visual appearance including textures and tissue deformations
  • Inverse pairing strategy to reconstruct real surgical videos in the simulator for paired simulated-real data
  • Sim-to-real translation via video diffusion model trained on paired data

Methodology

SWoMo employs a neuro-symbolic architecture: a symbolic rule-based simulator with scene graphs models motion dynamics and tool-tissue interactions, while a diffusion model generates realistic visual appearance. An inverse pairing strategy reconstructs real surgical videos in the simulator to create paired simulated and real videos, which are then used to train a video diffusion model for sim-to-real translation.

Key Results

SWoMo achieves qualitative and quantitative improvements over prior work, demonstrating generalization to unseen interaction geometries, improvements in downstream phase detection, and unsupervised video style transfer.

Tags

cataract surgerysurgical simulationworld modelneuro-symbolicscene graphcomputer visionCV