Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

QMBench: A Research Level Benchmark for Quantum Materials Research

arXiv 2025 58.5 method

TLDR

Introduces QMBench, a benchmark for evaluating LLM agents in quantum materials research using condensed matter physics and DFT.

Reasoning

The paper addresses a clear need for standardized evaluation in quantum materials research, with a well-defined benchmark covering multiple domains. However, it lacks real-world experimental validation and focuses solely on computational methods, limiting its scope.

Read-first score

Read-first score 58.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,experiment,metric,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
45

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 43.

Keyword Scores

AI scientist
8
AI for scientific research
8
autonomous research agent
7
scientific discovery agent
7
automated scientific discovery
6
automated research
6
research automation
5
automated experimentation
2
experiment design agent
2
literature review agent
1
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • Introduction of QMBench, a comprehensive benchmark for evaluating large language model agents in quantum materials research.
  • Benchmark covers multiple domains: structural properties, electronic properties, thermodynamic and other properties, symmetry principle, and computational methodologies.
  • Aims to provide a standardized evaluation framework to accelerate development of AI scientists for creative contributions in quantum materials.

Methodology

QMBench is designed as a benchmark that assesses large language model agents on their ability to apply condensed matter physics knowledge and computational techniques such as density functional theory. It includes tasks spanning structural, electronic, thermodynamic, symmetry, and computational methodology domains to evaluate research problem-solving in quantum materials science.

Key Results

The paper introduces the QMBench benchmark but does not report any experimental results or model evaluation outcomes.

Limitations

  • No experimental results or baseline performance metrics are provided.
  • The benchmark is presented as an initial version expected to be developed and improved by the research community, indicating potential gaps in coverage or difficulty calibration.

Tags

mtrl-sciAI