Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Visuomotor Grasping with World Models for Surgical Robots

arXiv 25.8 2025 50.1 method, application

TLDR

A visuomotor learning framework using world models for surgical robot grasping, enabling sim-to-real transfer and object-agnostic grasping with a single stereo camera.

Reasoning

The paper presents a clear contribution with a world-model-based architecture for surgical grasping, addressing sim-to-real transfer and generalization. Strengths include practical deployment on real robots in ex vivo settings. Weaknesses are the lack of quantitative results or comparisons in the abstract, and limited detail on the world model's specific role.

Read-first score

Read-first score 50.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 30.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,experiment

Topical relevance 42%
42.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 379.

Keyword Scores

world model
9
model-based reinforcement learning world model
6
world simulator
5
world dynamics prediction
4
interactive world model
3
generative world model
2
video world model
1

Deep Analysis

Innovations

  • GASv2: a visuomotor learning framework for surgical grasping that leverages world models
  • Addresses sim-to-real transfer of visuomotor policies to ex vivo surgical scenes using domain randomization
  • Enables object-agnostic grasping with a single policy that generalizes to diverse unseen surgical objects without retraining
  • Uses only a single stereo camera pair (standard RAS setup) for visuomotor learning, avoiding explicit pose tracking or handcrafted features
  • Hybrid control system for safe execution in surgical environments

Methodology

GASv2 employs a world-model-based architecture combined with a surgical perception pipeline for processing visual observations from a single stereo camera pair. The policy is trained in simulation using domain randomization to facilitate sim-to-real transfer, and a hybrid control system is used for safe execution on a real robot. Evaluation is conducted in both phantom and ex vivo surgical settings.

Key Results

The policy achieves a 65% success rate in both phantom and ex vivo settings, generalizes to unseen objects and grippers, and adapts to diverse disturbances, demonstrating strong performance, generality, and robustness.

Limitations

  • Success rate of 65% indicates significant room for improvement, especially for safety-critical surgical applications requiring higher reliability
  • Evaluation is limited to phantom and ex vivo settings; in vivo performance and generalization to real surgical environments remain unverified

Tags