Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

CityBench: Evaluating the Capabilities of Large Language Model as World Model

arXiv 24.6 2024 62.8 benchmark, system

TLDR

CityBench is a systematic benchmark using an interactive simulator to evaluate LLMs and VLMs as world models for diverse urban tasks across 13 cities.

Reasoning

The paper introduces a novel benchmark with real-world urban data and tasks, but its scope is limited to urban domains and does not address generative or video world models. The evaluation focuses on perception and decision-making, lacking explicit reinforcement learning or dynamics prediction components.

Read-first score

Read-first score 62.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Reproducibility 25%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Topical relevance 42%
50

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Methodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 343.

Keyword Scores

world model
9
world simulator
8
interactive world model
7
world dynamics prediction
5
generative world model
3
model-based reinforcement learning world model
2
video world model
1

Deep Analysis

Innovations

  • First systematic benchmark for evaluating LLMs and VLMs in urban research
  • CityData: integration of diverse urban data
  • CitySimu: simulation of fine-grained urban dynamics
  • 8 representative urban tasks in 2 categories (perception-understanding and decision-making)

Methodology

CityBench is an interactive simulator-based evaluation platform. It builds CityData to integrate diverse urban data and CitySimu to simulate fine-grained urban dynamics. Based on these, 8 representative urban tasks in two categories (perception-understanding and decision-making) are designed, and 30 well-known LLMs and VLMs are evaluated across 13 cities worldwide.

Key Results

Advanced LLMs and VLMs achieve competitive performance on tasks requiring commonsense and semantic understanding (e.g., human dynamics, semantic inference of urban images), but fail on tasks requiring professional knowledge and high-level numerical abilities (e.g., geospatial prediction, traffic control).

Tags