One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception
TLDR
WMGen-v1 uses a text-based world model with LVLM and LLM to generate physically grounded long-tail spatial data from a single image for perception tasks.
Reasoning
The paper introduces a novel framework combining world models with LLMs for synthetic data generation, addressing data scarcity in long-tail spatial perception. Strengths include explicit spatial grounding and physical plausibility; weaknesses are reliance on a single reference image and limited evaluation on internal datasets.
Read-first score
Read-first score 60.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.
Field roles
Rank sensitivity
Stability: volatile; rank range: 447.
Keyword Scores
Deep Analysis
Innovations
- Agentic text-based world model for one-shot image generation from a single reference image
- Use of LVLM for structured scene representation and LLM for physically plausible scene expansion
- Generation of long-tail spatial data with explicit spatial grounding and structural constraints via diffusion model conditioned on semantic representations
Methodology
WMGen-v1 employs a Large Vision-Language Model (LVLM) to construct a structured scene representation from a single reference image. A Large Language Model (LLM) then performs guidance-based scene expansion under physical plausibility and commonsense constraints. Finally, a diffusion model generates diverse and physically grounded long-tail training data conditioned on the structured semantic representations.
Key Results
On internal industrial datasets, ROADWork, and LaRS benchmarks, WMGen-v1 outperforms baseline approaches. Detectors trained solely on WMGen-v1 synthetic data approach real-only performance on aggregate dataset-level metrics.
Limitations
- Synthetic-only training still does not fully match real-only performance, as indicated by 'approach real-only performance'