Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
TLDR
Introduces Large Discovery Model, coupling generative model with Bayesian surrogate for open-ended search, achieving gains in neural-network, antibody, and molecular design.
Reasoning
The paper presents a novel architecture with strong empirical results across multiple scientific domains, which is a key strength. However, the abstract is truncated and lacks details on baselines and limitations, making full assessment difficult.
Read-first score
Read-first score 60.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 64.
Field roles
Rank sensitivity
Stability: volatile; rank range: 34.
Keyword Scores
Deep Analysis
Innovations
- Coupling a generative model with a Bayesian non-parametric reward surrogate to provide uncertainty-aware value for candidate guidance
- Empirically grounded recurrent architecture that refines candidates using the surrogate's uncertainty quantification
- Continual updating of discovery memory and surrogate model with each new experimental observation
Methodology
LDM combines a generative model (e.g., LLM) to propose and refine candidate designs with a Bayesian non-parametric surrogate that predicts performance and quantifies uncertainty. The surrogate's uncertainty-aware value guides candidate generation, refinement, and selection, and both the discovery memory and surrogate are updated continuously as new observations arrive. It is evaluated on neural network training, antibody design, and molecular optimisation against LLM-only reflection and traditional statistical search.
Key Results
LDM achieved a 2.4× greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and over 60% relative gains in molecular multi-objective performance compared to baselines.