ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
Read-first score
Read-first score 21.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 39.