Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy

arXiv 2025 52.4 method

TLDR

AstroVisBench evaluates LLMs on astronomy scientific computing and visualization, revealing significant performance gaps.

Reasoning

The paper introduces a novel benchmark for evaluating LLMs in astronomy-specific data processing and visualization, validated by professional astronomers. Its strength lies in addressing an underexplored evaluation area, but it is limited to a single domain and relies on an LLM-as-a-judge method that may introduce biases.

Read-first score

Read-first score 52.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 38.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,evaluation,experiment,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Topical relevance 42%
31.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 78.

Keyword Scores

AI scientist
8
AI for scientific research
7
research automation
5
automated research
4
automated scientific discovery
3
scientific discovery agent
3
autonomous research agent
2
automated experimentation
2
literature review agent
1
survey generation
1
experiment design agent
1
paper writing agent
1

Deep Analysis

Innovations

  • First benchmark for both scientific computing and visualization in astronomy (AstroVisBench)
  • Novel LLM-as-a-judge workflow for evaluating visualizations, validated against annotations by five professional astronomers

Methodology

AstroVisBench is a benchmark that tasks language models with creating astronomy-specific data processing and analysis workflows and visualizing results through complex plots. Evaluation of visualizations uses an LLM-as-a-judge approach validated by five professional astronomers. State-of-the-art language models are evaluated on this benchmark.

Key Results

Evaluation reveals a significant gap in state-of-the-art language models' ability to serve as useful assistants for astronomy research.

Tags

CLIMLG