Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

arXiv 2026 48.8 method, application

TLDR

UniT unifies human and humanoid policy learning and world modeling via a visual-anchored latent action tokenizer, enabling zero-shot transfer and efficient video generation.

Reasoning

Strengths: novel cross-embodiment transfer using visual anchoring, demonstrated on real-world and simulation. Weaknesses: limited details on world modeling evaluation; reliance on human data may have biases.

Read-first score

Read-first score 48.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 38.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
62.5

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.62533644

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation

Topical relevance 29%
54.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridge

Rank sensitivity

Stability: volatile; rank range: 242.

Keyword Scores

world model
8
video world model
8
generative world model
7
world dynamics prediction
6
model-based reinforcement learning world model
4
interactive world model
3
world simulator
2

Deep Analysis

Innovations

  • Unified physical language for human-to-humanoid transfer via visual anchoring
  • Tri-branch cross-reconstruction mechanism: actions predict vision, vision reconstructs actions, fusion branch
  • Shared discrete latent space of embodiment-agnostic physical intents
  • Two paradigms: VLA-UniT for policy learning and WM-UniT for world modeling

Methodology

UniT employs a tri-branch cross-reconstruction mechanism where actions predict vision to anchor kinematics to physical outcomes, vision reconstructs actions to filter out irrelevant visual confounders, and a fusion branch synergizes these purified modalities into a shared discrete latent space of embodiment-agnostic physical intents. The framework is validated across policy learning (VLA-UniT) and world modeling (WM-UniT) using massive egocentric human data, with evaluation on humanoid simulation benchmarks and real-world deployments.

Key Results

VLA-UniT achieves state-of-the-art data efficiency and robust out-of-distribution generalization, including zero-shot task transfer on humanoid simulation and real-world deployments. WM-UniT enables direct human-to-humanoid action transfer for video generation, and t-SNE visualizations confirm convergence of human and humanoid features into a shared manifold.

Tags

humanoid roboticscross-embodiment transfervisual anchoringlatent action tokenizerpolicy learningworld modelingROAI