Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature

arXiv 2026 56.3 method, application

TLDR

A world model predicting 3D isovist (navigable geometry) from past isovists and actions, revealing emergent cross-city spatial signatures.

Reasoning

Strengths: novel predictive target (3D isovist) avoids appearance bias and preserves 3D structure; uses depth residuals and self-rollout sampling. Weaknesses: limited to two cities; abstract ends abruptly, lacking explicit real-world evaluation or benchmark results.

Read-first score

Read-first score 56.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 29.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,baseline,dataset,metric

Citation impact 18%
80.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.80650114

Reproducibility 18%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Topical relevance 29%
41.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 433.

Keyword Scores

world model
9
world dynamics prediction
8
generative world model
5
interactive world model
3
model-based reinforcement learning world model
2
world simulator
1
video world model
1

Deep Analysis

Innovations

  • Modeling navigable geometry as a 3D isovist (spherical visibility-depth map) instead of appearance or 2D occupancy grids.
  • Depth residual formulation for prediction to preserve sharp building edges.
  • Self-rollout scheduled sampling to keep corrupted context on the geometry manifold.
  • Persistent latent bird's-eye-view spatial map for cross-path consistency.
  • Emergent cross-city spatial signature decodable from temporal latents, showing city identity is learned in dynamics rather than appearance.

Methodology

The paper introduces an embodied world model that predicts the next 3D isovist from a short history of past isovists and a movement action. The prediction is formulated as a depth residual, trained with self-rollout scheduled sampling, and equipped with a persistent latent bird's-eye-view spatial map. The model is trained on data from Manhattan and Paris.

Key Results

A single city-blind model trained on Manhattan and Paris develops a cross-city spatial signature, with city identity linearly decodable from its temporal latents far above single-frame baselines.

Limitations

  • Only evaluated on two cities (Manhattan and Paris), so cross-city generalization to other urban forms is unverified.
  • The isovist representation assumes static geometry, ignoring dynamic obstacles like vehicles or pedestrians.
  • The method requires accurate depth sensing to compute isovists, which may not be available in all real-world settings.
  • The persistent latent BEV map may introduce additional memory and computational overhead.

Tags

3D isovistworld modelnavigationgeometryurban environmentembodied agentsROLG