The Automated LLM Speedrunning Benchmark
Dataset Analysis
Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce existing work. To evaluate the ability of AI agents to reproduce ...
Provenance
Collected from papers.
Derived from paper: The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements