PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
Read-first score
Read-first score 23.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 39.