Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

arXiv 2026 60.3 method, application

TLDR

WMGen-v1 uses a text-based world model with LVLM and LLM to generate physically grounded long-tail spatial data from a single image for perception tasks.

Reasoning

The paper introduces a novel framework combining world models with LLMs for synthetic data generation, addressing data scarcity in long-tail spatial perception. Strengths include explicit spatial grounding and physical plausibility; weaknesses are reliance on a single reference image and limited evaluation on internal datasets.

Read-first score

Read-first score 60.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.

Methodology quality 18%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,dataset,experiment,metric,result

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.4

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94402284

Topical relevance 29%
45.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 447.

Keyword Scores

world model
9
generative world model
8
world dynamics prediction
5
video world model
4
world simulator
3
interactive world model
2
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Agentic text-based world model for one-shot image generation from a single reference image
  • Use of LVLM for structured scene representation and LLM for physically plausible scene expansion
  • Generation of long-tail spatial data with explicit spatial grounding and structural constraints via diffusion model conditioned on semantic representations

Methodology

WMGen-v1 employs a Large Vision-Language Model (LVLM) to construct a structured scene representation from a single reference image. A Large Language Model (LLM) then performs guidance-based scene expansion under physical plausibility and commonsense constraints. Finally, a diffusion model generates diverse and physically grounded long-tail training data conditioned on the structured semantic representations.

Key Results

On internal industrial datasets, ROADWork, and LaRS benchmarks, WMGen-v1 outperforms baseline approaches. Detectors trained solely on WMGen-v1 synthetic data approach real-only performance on aggregate dataset-level metrics.

Limitations

  • Synthetic-only training still does not fully match real-only performance, as indicated by 'approach real-only performance'

Tags

world modelone-shot generationlong-tail distributionspatial perceptionautonomous drivingsynthetic dataCVAI