Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models

arXiv 25.9 2025 47 method, system, application

TLDR

We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution.

Reasoning

Fallback reasoning generated from available title and abstract metadata: We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video demonstrations of candidate trajectories, executes them on the robot, and...

Read-first score

Read-first score 47, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,evaluation,experiment,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Topical relevance 42%
23.5

Matches configured research keywords against title, abstract, tags, and analysis text. matched=5

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 497.

Deep Analysis

Innovations

  • Integration of iterative reasoning, diffusion-based video generation, and closed-loop execution for robotic manipulation
  • Use of foundation world models to enable general-purpose cognitive robotics
  • Evaluation across multiple perceptual modalities (egocentric, third-person, simulated) and robotic embodiments (bimanual UR3, Unitree G1 humanoid, simulated GR1)

Methodology

PhysicalAgent is an agentic framework that generates short video demonstrations of candidate trajectories from textual instructions using diffusion-based video generation, executes them on a robot, and iteratively re-plans in response to failures. It is evaluated across multiple perceptual modalities and robotic embodiments, comparing against state-of-the-art task-specific baselines.

Key Results

The method achieves up to 83% success on human-familiar tasks. First-attempt success is limited to 20-30%, but iterative correction increases overall success to 80% across platforms.

Limitations

  • First-attempt success is limited (20-30%)
  • Reliance on iterative correction to achieve high overall success rates

Tags