AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
TLDR
AstroVisBench evaluates LLMs on astronomy scientific computing and visualization, revealing significant performance gaps.
Reasoning
The paper introduces a novel benchmark for evaluating LLMs in astronomy-specific data processing and visualization, validated by professional astronomers. Its strength lies in addressing an underexplored evaluation area, but it is limited to a single domain and relies on an LLM-as-a-judge method that may introduce biases.
Read-first score
Read-first score 52.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 38.
Field roles
Rank sensitivity
Stability: volatile; rank range: 78.
Keyword Scores
Deep Analysis
Innovations
- First benchmark for both scientific computing and visualization in astronomy (AstroVisBench)
- Novel LLM-as-a-judge workflow for evaluating visualizations, validated against annotations by five professional astronomers
Methodology
AstroVisBench is a benchmark that tasks language models with creating astronomy-specific data processing and analysis workflows and visualizing results through complex plots. Evaluation of visualizations uses an LLM-as-a-judge approach validated by five professional astronomers. State-of-the-art language models are evaluated on this benchmark.
Key Results
Evaluation reveals a significant gap in state-of-the-art language models' ability to serve as useful assistants for astronomy research.