Awesome Auto Research Hub Papers · Datasets · Projects

Aggregate analysis

Research Analysis

Cross-paper synthesis of shared research patterns, differences, mainstream directions, and trend evolution.

379 papers

AI Scientist Systems

End-to-end autonomous research agents that combine ideation, experimentation, analysis, and manuscript generation.

56 Methodology quality 94 Recency 41 Reproducibility 39 Topical relevance

Literature review synthesis

Research Lines

End-to-end autonomous research pipelines

Integrates idea generation, experiment execution, result analysis, and manuscript writing into a single closed loop.

Open edge: Publication-grade consistency beyond a single accepted workshop paper or a small set of generated papers; external validation of scientific contribution remains weak.
Domain-specific discovery agents with verification gates

Targets theory-driven or physical domains by embedding logical coherence checks, knowledge retrieval, or vision-language inspection of rendered simulations.

Open edge: Verification is incomplete: some silent failures are missed, and gates are tightly coupled to domain-specific signals.
Benchmarking and evaluation environments

Provides standardized tasks, datasets, and human baselines for comparing scientific discovery agents.

Open edge: Proxy metrics and multiple-choice or task-completion measures may not capture full open-ended discovery; domain coverage and real decision utility are limited.
Human-AI collaborative research pipelines

Combines agent autonomy with targeted human oversight, self-healing execution, and verifiable reporting to reduce fabricated or hallucinated outputs.

Open edge: Optimal human intervention points and cost-effectiveness across domains are not established; evidence is limited to experiment-stage benchmarks.

Shared Direction

  • Autonomous components alone are insufficient; systems add verification, review, or human checks to suppress unsound and fabricated outputs.
  • End-to-end workflows converge on a hypothesis-experiment-analysis-writing loop.
  • Benchmarks and reviewer-style evaluations are becoming central for comparing AI scientist systems.
  • Current systems show potential but fall short of reliable fully autonomous discovery across domains.

Key Differences

  • Evaluation target differs: some use simulated review or workshop acceptance, while others use domain numerical improvement, task completion, or human-baselined biological questions.
  • Verification strategy differs: internal logical consistency checks, vision-language physics gates, verifiable result reporting, or curated benchmark tasks.
  • Domain focus differs: general machine learning, applied mathematics reasoning, computational fluid dynamics, and omics biology shape the required infrastructure and evidence.
  • Interaction mode differs: fully autonomous pipelines contrast with targeted human collaboration at high-leverage decision points.

Open Questions

  • Do automated review and rubric scores align with real scientific novelty, reproducibility, and downstream utility?
  • Can domain-specific verification mechanisms generalize beyond the narrow physical and biological settings where they were demonstrated?
  • What is the right granularity of human collaboration, and how does it affect reliability and cost?
  • How can success inconsistency and missing code, dataset, metric, baseline, or limitation details be reduced?
AILGCLmtrl-sciMACYDLIR

312 papers

Evaluation, Reliability & Governance

Evaluation, reproducibility, safety, epistemic reliability, governance, and benchmark infrastructure for automated research.

56 Methodology quality 93 Recency 42 Reproducibility 37 Topical relevance

Literature review synthesis

Research Lines

End-to-end autonomous research agents

Turning idea generation, code execution, experiment running, and paper writing into a single automated workflow; The AI Scientist and AI Scientist-v2 are representative examples.

Open edge: Success is inconsistent: only a fraction of submissions pass peer review, and acceptance is demonstrated mainly at workshop level rather than full conference or broad-domain reliability.
Task-suite and environment design

Making discovery capability comparable through structured tasks, baselines, and metrics such as task completion, task-relevant actions, and explanatory knowledge; DiscoveryWorld and BAISBench illustrate this line.

Open edge: Scientific field coverage remains limited, and agreement between task metrics and real discovery value or downstream decision utility is not established.
Verification and self-correction

Reducing silent failures and hallucinated results through physics-aware visual gates, grounded reporting, or multi-hypothesis failure attribution; SAGE and AI CFD Scientist are representative.

Open edge: Verification gates still miss some planted failures, and self-correction gains do not yet guarantee fully autonomous or conference-ready scientific writing.
Multi-agent and interactive research platforms

Coordinating customizable roles, real-time monitoring, artifact inspection, rollback or resume control, and local code integration; Claw AI Lab is an example.

Open edge: Evaluation relies on small internal case studies, so standardized external validation across hardware, datasets, and user settings is still missing.

Shared Direction

  • Autonomous research systems are evaluated not only on idea quality but also on whether they produce runnable experiments, metrics-bearing outputs, and valid artifacts.
  • Current AI scientist systems remain unreliable enough that explicit verification, review, or failure-attribution stages are becoming common.
  • Evaluation evidence is shifting toward structured tasks, human or expert baselines, and blind or automated judging rather than isolated demonstrations.
  • Reproducibility and external validity are documented unevenly; open-source artifacts are common in visible examples but cross-setting validation remains limited.

Key Differences

  • What counts as verification differs: some systems inspect rendered flow fields for physical validity, while others simulate manuscript review or attribute failed experiment trajectories to root causes.
  • Evaluation targets differ: full paper acceptance or blind review scores versus task completion and explanatory knowledge versus domain-specific cell annotation and discovery questions.
  • Interaction and oversight differ: fully autonomous loops, self-correcting loops, task-suite environments, and interactive multi-agent dashboards imply different human roles and levels of trust.
  • Domain scope differs: some systems specialize in computational fluid dynamics or omics, while others target general machine learning topics, making generalization claims hard to compare.

Open Questions

  • How can verification mechanisms be made broad enough to detect silent failures beyond the specific physics or visual cases used in current ablations?
  • Do task completion, automated review scores, and artifact quality metrics align with real scientific contribution or downstream research utility?
  • Which reported failure modes and self-correction gains transfer to other domains, longer-horizon research, and full conference-level contributions?
  • What minimal external validation, data and code availability, and limitation reporting should be required before autonomous research claims are trusted?
AICLLGMAmtrl-sciIRDLCY

275 papers

Literature Intelligence

Automated literature search, paper reading, citation analysis, literature review, survey synthesis, and research gap discovery.

52 Methodology quality 92 Recency 41 Reproducibility 35 Topical relevance

Literature review synthesis

Research Lines

Rubric-guided survey generation systems, exemplified by ARISE

Converts topic prompts into structured survey manuscripts through role-specialized LLM agents and iterative reviewer feedback against behaviorally anchored rubrics.

Open edge: Whether rubric-aligned quality scores predict acceptance by human domain experts or factual citation reliability outside the evaluated corpus.
Verification-aware literature synthesis and paper generation systems, exemplified by ReasFlow

Integrates verification, knowledge retrieval, literature synthesis, theorem proving, experimentation, and manuscript preparation in one multi-agent workflow.

Open edge: Generalizability beyond applied mathematics examples, with no visible limitations section or external benchmark comparison in the provided evidence.
Benchmark and human-baseline evaluation for AI scientists, exemplified by BAISBench

Creates expert-labeled discovery tasks and graduate-level human baselines to compare AI scientist capabilities on real scientific data.

Open edge: Tasks are limited to omics biology discovery and do not isolate literature search, citation analysis, or survey synthesis as independent skills.
Domain-agnostic self-evolving agent frameworks, exemplified by EvoMaster

Provides scalable infrastructure for hypothesis refinement, self-critique, and knowledge accumulation across multiple scientific benchmarks.

Open edge: Only self-reported benchmark scores are visible in this packet; external validation and reproducible adaptation across arbitrary disciplines remain open.

Shared Direction

  • Multiple systems decompose literature and scholarly workflows into specialized LLM agents for retrieval, drafting, verification, and review.
  • Automated quality scoring depends on LLM-based rubrics or benchmark-derived metrics rather than established human peer review.
  • Human oversight remains present, either as a principal-investigator role or through human baseline comparisons.
  • Current evidence is mostly self-reported on curated tasks or small case studies; external validation and reproducible comparison are limited.

Key Differences

  • Evaluation target differs: one line optimizes rubric-aligned survey quality; another measures domain-specific discovery through expert-labeled multiple-choice and annotation tasks; another evaluates end-to-end paper completion or acceptance.
  • Interaction mode differs: some systems position a human as principal investigator who inspects outputs, while others emphasize continuous self-evolution and accumulated agent knowledge.
  • Domain deployment differs across applied mathematics, omics biology, machine learning papers, and general scientific benchmarks, implying different evidence requirements and task constraints.
  • Supervision signal differs: structured behaviorally anchored rubrics versus human expert baselines versus multi-AI or human review of generated papers.

Open Questions

  • Do LLM-rubric quality scores align with human expert judgment and real scientific utility outside the evaluated corpus?
  • Can these systems maintain citation correctness, evidence traceability, and gap discovery across long-horizon literature reviews?
  • Do verification loops and self-evolution mechanisms reduce documented failure modes beyond four-case or benchmark-specific evidence?
  • What benchmark design would isolate literature search, paper reading, citation analysis, and research gap discovery from broader AI-scientist abilities?
AICLLGIRDLMAmtrl-sciCY

275 papers

Writing & Communication

Automatic paper writing, manuscript drafting, scientific reporting, peer review assistance, and research communication.

47 Methodology quality 93 Recency 41 Reproducibility 34 Topical relevance

Literature review synthesis

Research Lines

End-to-end autonomous manuscript generation

Turns raw research goals into hypotheses, experiments, results, and complete papers with minimal human intervention.

Open edge: Whether generated papers are consistently accepted beyond workshop-level or selected domains; reported acceptance is mixed across submissions.
Verification and validity gating

Detects logical, physical, numeric, or citation errors before claims enter a manuscript.

Open edge: Domain-specific gates still miss planted silent failures and may not transfer across fields.
Human-AI collaborative research workflows

Places human intervention at high-leverage decision points rather than exhaustive oversight, improving benchmark performance.

Open edge: Which intervention points generalize and how much expert load remains is not established.
Self-correcting failure attribution

Attributes experiment failures to hypothesis, design, or implementation and routes corrections to the appropriate layer.

Open edge: Conference-ready writing and long-horizon scientific reasoning remain unsolved.

Shared Direction

  • Autonomous research systems converge on decomposing scientific discovery into modular stages: ideation, implementation, execution, analysis, verification, and paper drafting.
  • Raw LLM generation alone is considered insufficient; some validity gate, verification loop, or grounded reporting mechanism is needed to prevent fabricated numbers or invalid claims.
  • Automated review or LLM-based scoring can approximate human judgment but requires external peer review or domain-specific checks for stronger validity.
  • Recent systems increasingly treat failure as information and add self-correction or human collaboration at high-leverage points rather than relying on fully autonomous loops.

Key Differences

  • Interaction mode differs: fully autonomous end-to-end generation versus targeted human-in-the-loop collaboration.
  • Verification target differs: logical coherence and LLM review versus domain-specific physics or vision validity checks versus verifiable numerical reporting.
  • Evaluation evidence differs: LLM-rubric scores, workshop peer acceptance, benchmark success rate, blind review score, or failure-case analysis.
  • Domain scope differs: general machine learning pipelines versus applied mathematics, computational fluid dynamics, or specific benchmark topic suites.

Open Questions

  • Can automated review scores predict human conference or workshop acceptance across domains when reported AI-generated submissions had mixed peer-review outcomes?
  • How can domain-specific validity gates generalize without inheriting miss rates, such as missing two of sixteen planted silent failures?
  • Does attribution-guided self-correction improve long-horizon scientific writing, or does it mainly improve code and experiment execution outputs?
  • What is the minimum human intervention needed to preserve research quality while reducing expert load?
AILGCLMAmtrl-sciDLCYIR

248 papers

Hypothesis & Idea Generation

Systems that propose hypotheses, research ideas, experimental directions, or scientific questions.

55 Methodology quality 93 Recency 41 Reproducibility 37 Topical relevance

Literature review synthesis

Research Lines

Autonomous experiment generation and self-correction

Transforms hypotheses into runnable experiments and improves outcomes by attributing failures to hypothesis, experimental design, or implementation.

Open edge: Verification gates still miss silent failures; domain-specific validity outside the training simulator or data distribution is not established.
Formally grounded theory generation

Produces machine-checked theoretical artifacts for causal inference and reduces proof errors through formal verification.

Open edge: Formal proofs do not guarantee faithful capture of intended scientific claims; human inspection of artifacts remains required.
Human-AI collaborative research pipelines

Allocates human intervention to high-leverage decision points instead of requiring full autonomy or exhaustive step-by-step oversight.

Open edge: Optimal interaction modes and cross-discipline generalization remain unknown; evidence comes from limited case studies.
Benchmarking and evidence design for AI scientists

Provides comparable tasks and human or baseline references on real biological data and multi-domain experiment benchmarks.

Open edge: Proxy metrics may not reflect genuine discovery utility; domain coverage is narrow and generalization or contamination is uncertain.

Shared Direction

  • Autonomous scientific discovery systems commonly decompose work into hypothesis generation, experiment execution, validation or failure detection, and reporting.
  • Reliability is treated as a core concern: multiple systems add explicit verification, grounding, or proof-checking because end-to-end generation alone can produce unsupported claims.
  • Human oversight remains valuable, especially at targeted intervention points, rather than being entirely removed or applied at every step.
  • Evaluation centers on benchmark scores, artifact quality, and comparison with a reference autonomous research baseline.

Key Differences

  • Verification signal differs: some rely on visual physics inspection of rendered outputs, others on formal proof assistants, executable result grounding, or evidence-grounded failure attribution.
  • Desired autonomy differs: open-ended autonomous search versus selective human collaboration versus human-reviewed formal artifacts.
  • Evaluation target differs: domain-specific CFD or omics discovery tasks versus broad multi-domain science benchmarks versus internal expert preference judgments.
  • Failure correction granularity differs: some route to hypothesis, experimental design, or implementation level, while others apply global reflection or self-evolution.

Open Questions

  • How can autonomous discovery systems establish external validity beyond their training simulators, datasets, or formal assumptions?
  • Which verification mechanism best catches silent scientific failures without blocking valid but unusual results?
  • What amount and timing of human interaction maximize novelty and reliability without requiring exhaustive oversight?
  • Do current benchmark scores predict meaningful scientific discovery or only improvements in coding, execution, and report formatting?
  • How can formally machine-checked claims be audited for semantic alignment with the intended non-formal scientific contribution?
AICLLGMAIRmtrl-sciDLCE

178 papers

Experiment Automation

Agents that design, execute, monitor, analyze, or iterate experiments, including code-based and lab-based experimentation.

62 Methodology quality 93 Recency 41 Reproducibility 43 Topical relevance

Literature review synthesis

Research Lines

End-to-end autonomous experiment workflows

turns a research prompt or minimal hypothesis into runnable code, experiment execution, result analysis, and manuscript generation

Open edge: reliability beyond selected code-centric domains; only one of three submissions passed peer review in the workshop-level evidence
Experiment planning and ablation benchmarks

makes automated proposal and review of experiment perturbations comparable through curated author/reviewer tasks and LLM-based judges

Open edge: best LM recall remains around 45% and author/reviewer performance trends invert, so grounding assumptions are unresolved
Simulated discovery evaluation environments

provides controlled task suites and metrics for complete hypothesis-experiment-conclusion cycles

Open edge: whether task completion and explanatory-knowledge metrics predict real scientific discovery remains unproven
Failure-aware experience memory

converts failed experiment attempts into retrievable records that downstream agents can adopt or reject, improving retry and cross-task performance with lower token use

Open edge: evidence is limited to ScienceAgentBench and nonlinear PDE tasks; transfer to broader experimental domains is unverified

Shared Direction

  • LLM agents can complete hypothesis proposal, code implementation, experiment execution, result analysis, and paper writing in a closed loop, at least in constrained code-centric domains.
  • Automated evaluation is widely used as a scalable proxy signal, while peer review or human expert judgment remains the reference.
  • Current systems are stronger on code-centric tasks and have weaker evidence for physical or more open-ended experimentation.
  • Feedback mechanisms such as verification, retrieval, simulated review, or failure memory improve later experiment iterations.

Key Differences

  • Autonomy level differs: some systems pursue full automation, while others keep human inspection, monitoring, and rollback.
  • Evaluation target differs: complete discovery tasks, paper acceptance, manuscript quality, ablation recall, or token efficiency are used as separate success criteria.
  • Orchestration differs: single-agent pipelines, tree search, multi-agent roles, failure-curated shared memory, and dashboard-controlled teams coexist.
  • Evidence design differs: simulated environments assess task completion, while empirical research benchmarks assess missing-ablation identification; these settings are not yet unified.

Open Questions

  • Do simulated task completion and LLM-based review scores align with real scientific reproducibility or impact?
  • Can failure-aware memory transfer from ScienceAgentBench and nonlinear PDE tasks to other experiment types?
  • Why does chain-of-thought prompting outperform the agent-based setup on ablation planning, and why do author and reviewer task performances show inverse trends?
  • Can fully automated workshop-level paper acceptance scale to conference-level, multidisciplinary, or physical lab experiments?
  • The visible packet lacks code, dataset, and limitation details for several high-ranked works, leaving reproducibility boundaries unverified.
AILGCLmtrl-sciMACYCEIR

41 papers

Survey Generation

Papers and resources related to Survey Generation.

52 Methodology quality 88 Recency 42 Reproducibility 33 Topical relevance

Literature review synthesis

Research Lines

Agentic survey drafting and iterative refinement

Turns retrieved literature into structured drafts through modular agents and repeated rubric-based or review-based revision.

Open edge: Quality gains are often measured by LLM rubrics rather than external expert agreement; limited evidence on failure cases and domain transfer.
Retrieval, evidence grounding, and citation reliability

Improves selection and attribution of source papers through full-text analysis, paper cards, citation-graph expansion, and evidence-constrained citation.

Open edge: Available evidence reports low expert-cited recall and uneven precision; missing standardized evaluation across corpora and disciplines.
Taxonomy and hierarchical organization generation

Learns or generates expert-like category structures that organize papers into coherent trees and sections.

Open edge: Generated taxonomies show sibling overlap, MECE violations, structural imbalance, and weak semantic alignment with human experts.
Evaluation and benchmarking for survey systems

Creates benchmarks, metrics, and paired comparison settings to assess retrieval, organization, content coverage, structure, and citation quality.

Open edge: Reference-dependent alignment may penalize valid alternative taxonomies; proxy metrics have not been shown to predict real scholarly utility.

Shared Direction

  • Iterative multi-step agent workflows outperform simpler single-pass or retrieval-only baselines on quality, structure, and citation-related metrics.
  • Retrieval coverage and citation faithfulness are central bottlenecks; systems increasingly use paper-level summaries or evidence constraints to ground claims.
  • Automated evaluation is necessary because expert review is costly, but current evaluation remains partial and often reference-dependent.
  • Taxonomy/tree organization is treated as a distinct capability from fluent drafting, with separate metrics for hierarchy and semantic alignment.

Key Differences

  • Representation differs: some methods use free-text iterative drafts and rubrics, others use paper cards, structured keynotes, or taxonomy trees as intermediate artifacts.
  • Supervision/evaluation target differs: rubric-guided scoring, multi-LLM pairwise or 12-dimension evaluation, expert taxonomy alignment, and real-world biomedical validation create inconsistent comparison targets.
  • Interaction mode differs: several systems are fully automated multi-agent pipelines, while governed research systems include scientist-in-the-loop oversight and institutional data controls.
  • Depth-grounding differs: methods vary from abstract/keyword retrieval to full-text analysis, code-repository inspection, and citation-graph expansion.

Open Questions

  • Do rubric-based or multi-LLM quality scores agree with domain expert judgments when survey claims are nuanced or contested?
  • Why does the best evaluated deep research agent retrieve only about 20.92% of expert-cited papers, and what retrieval or indexing changes would close that gap?
  • Can generated taxonomies be both reference-free and expert-aligned enough to reduce MECE violations and structural imbalance without forcing a single human taxonomy?
  • How well do survey-generation systems transfer from LLM research topics to non-CS fields, and which evidence constraints remain necessary for reliable citation?
CLAIIRDLHCLG

38 papers

Hypothesis Generation

Papers and resources related to Hypothesis Generation.

66 Methodology quality 94 Recency 38 Reproducibility 47 Topical relevance

Literature review synthesis

Research Lines

Autonomous multi-agent hypothesis generation and iterative refinement workflows

turning literature and data into hypotheses that are updated through experimental feedback, code execution, and result interpretation

Open edge: report statement accuracy, weak-direction rejection, and scaling beyond limited cycles or demonstrated domains remain uncertain
Human-in-the-loop and mixed-initiative research oversight

balancing autonomy with human judgment and domain-conditioned control over research workflows

Open edge: optimal intervention modes and accountable closure in embodied, delayed, heterogeneous, or ethical settings are not established
Structured search and ideation mechanisms for research trajectories

maintaining controlled exploration, branching, backtracking, and long-horizon coherence in hypothesis search

Open edge: how much benefit comes from search orchestration, base model capability, or task-specific prior is unclear
Benchmarks and evaluation frameworks for discovery capabilities

making claims comparable through tasks, baselines, metrics, and difficulty measurement

Open edge: alignment between proxy task accuracy and real scientific utility, and broader domain coverage, still lack evidence

Shared Direction

  • Most systems treat hypothesis generation as part of an iterative loop linking literature or data analysis with experimental feedback or code execution.
  • Multi-agent decomposition is common, with separate roles for literature search, data analysis, ideation, execution, and verification.
  • Human oversight is often most useful at targeted high-leverage decision points rather than exhaustive step-by-step supervision.
  • Traceable reporting and benchmark tasks are emerging as core mechanisms for evaluating autonomous research systems.
  • Current systems are considered more credible in structured, executable, and rapidly verifiable settings than in less controllable environments.

Key Differences

  • Human role differs: some systems target full autonomy, others emphasize multiple human intervention modes, and surveys frame autonomy as domain-conditioned.
  • Exploration mechanisms differ: structured multi-agent debate, evolutionary-systematic research trees, code-interpreter-based iterative optimization, and reinforcement learning are used as distinct coordination strategies.
  • Evaluation targets differ: biology task benchmarks, experiment-stage research benchmarks, equation discovery datasets, wet-lab validation, and report statement accuracy coexist without a unified standard.
  • Supervision and optimization signals differ: self-healing feedback, evolutionary search, structured world-model information sharing, and end-to-end reinforcement learning represent different routes to improve hypothesis quality.

Open Questions

  • How can systems reduce inaccurate or unsupported statements while preserving traceability and verifiable reporting?
  • Does linear scaling of valuable findings continue beyond the tested cycle counts and across open-ended scientific domains?
  • Which evaluation metrics predict real-world scientific utility, especially in embodied, delayed, heterogeneous, and ethically constrained settings?
  • Which structured exploration or human intervention mode generalizes across domains, and what evidence would demonstrate that transfer?
  • Can long-horizon autonomous research systems reject weak directions early and maintain evidence provenance without human cleanup?
  • What verification gaps remain because many source records do not state limitations or release reproducible artifacts?
AICLLGGNMADLCENE

37 papers

Scientific Discovery Agents

Papers and resources related to Scientific Discovery Agents.

30 Methodology quality 92 Recency 48 Reproducibility 22 Topical relevance

Literature review synthesis

Research Lines

Science discovery environments and benchmarks

Provides standardized tasks, interaction protocols, and metrics to evaluate AI agents on end-to-end scientific discovery, from hypothesis formation to experiment and conclusion.

Open edge: Task coverage is limited to certain domains; LLM-based judges show only moderate agreement with human experts; outcome metrics cannot detect incorrect reasoning that yields correct results.
Agent infrastructure and tool augmentation

Automates parts of the discovery pipeline—code generation, wet-lab execution, training, safety checks—to improve agent efficiency, reproducibility, and task completion rate.

Open edge: Generalization across scientific domains is unverified; most systems are evaluated in isolation or with simulated tasks, not in real-world lab settings; long-term autonomy and reliability remain open.
Faithful reasoning and safety mechanisms

Integrates explicit checks, structured certificates, or safety loops to ensure agents' reasoning is consistent with evidence and free from dangerous actions or tool-chain escapes.

Open edge: Current safety reasoning is evaluated on limited tasks; detecting all compositional risks without stifling exploration is hard; mechanism-fidelity checks often require known ground-truth, limiting use in open-ended discovery.

Shared Direction

  • All works adopt a multi-stage experiment loop (hypothesis, experiment, analysis) as the agent's decision cycle.
  • Current LLM-based agents are insufficient for robust scientific discovery and require additional scaffolding, such as tools, memory, or safety modules.
  • Evaluation should go beyond task completion and include process quality, safety, or reasoning faithfulness.

Key Differences

  • Environment type: some use fully text-based virtual labs, others rely on code execution with real software constraints, and a few simulate physical wet-lab operations.
  • Safety integration: one approach embeds safety reasoning into every agent reasoning step, while another applies external safety gates after trajectory completion.
  • Question formation: one line advocates structured, auditable research question certificates; others keep question generation as a free-form prompt step without explicit traceability.
  • Evaluation judges: one study reports moderate agreement between LLM-as-a-judge and domain experts, while others use LLM judges without human calibration, raising concerns about proxy reliability.

Open Questions

  • How can we build a single agent infrastructure that generalizes across diverse scientific fields without per-domain retuning?
  • What automatic methods can detect 'correct answer, wrong mechanism' failures when the ground-truth mechanism is unknown?
  • How can safety mechanisms be designed to prevent compositional tool-chain risks without over-restricting the agent's ability to explore novel hypotheses?
  • What is the minimal set of reality-grounded tasks and metrics needed to validate that a simulated discovery agent will succeed in a real laboratory?
AILGCLcomp-phmtrl-scichem-ph

25 papers

Research Agent Benchmarks

Papers and resources related to Research Agent Benchmarks.

8 Methodology quality 92 Recency 41 Reproducibility 6 Topical relevance

Literature review synthesis

Research Lines

Action-conditioned control and planning

Enables LLM agents to execute multi-step research actions—querying literature, generating hypotheses, running experiments, and self-correcting—through closed-loop decision-making and feedback.

Open edge: Robustness of agent decisions in open-ended, long-horizon tasks; whether improvements come from better planning, larger models, or domain-specific heuristics; and how to prevent error cascades.
Evaluation and evidence design

Standardizes comparison of research agents through curated benchmarks, metrics (e.g., accuracy, MAE improvement), and controlled task designs (coding problems, survey generation, held-out validation).

Open edge: Coverage of diverse scientific domains, alignment between proxy metrics and real-world research utility, and the lack of consensus on what constitutes a valid test of autonomous research capability.
Systems and reproducible infrastructure

Packages agent workflows into auditable, executable pipelines that produce reusable code and artifacts, enabling verification across different environments.

Open edge: Standardized external validation across hardware, data distributions, and user settings; most systems lack comprehensive reproducibility artifacts beyond the reported experiments.

Shared Direction

  • LLMs serve as the core reasoning and coding engine for research agents.
  • The research process is broken into discrete steps (search, plan, code, verify) with feedback signals to guide the agent.
  • Evaluation should go beyond toy tasks and move toward real scientific problem-solving, though implementations differ.
  • Auditability and reproducibility are emerging as desired properties of research agent systems.

Key Differences

  • Evaluation approach: static benchmarks with pre-defined metrics versus dynamic held-out task transfer that measures generalization to unseen data.
  • Domain scope: some works target specific scientific fields (materials, mathematics), while others aim for general-purpose research agents.
  • Agent architecture: recursive self-improvement via reinforcement learning against predefined search space decomposition with inner-fold validation.
  • Reproducibility granularity: some systems provide full code and auditable logs, while others only describe the methodology without released artifacts.

Open Questions

  • Do current benchmarks (e.g., academic survey tasks, coding problems) genuinely measure progress toward autonomous research, or do they reward narrow capabilities?
  • How can agent decision quality be validated in open-ended research where ground truth is unknown or evolves over time?
  • What is the minimum set of evidence (code, data, logs) required to trust a research agent's findings, and how can this be standardized across domains?

23 papers

Automated Experimentation

Papers and resources related to Automated Experimentation.

42 Methodology quality 86 Recency 40 Reproducibility 28 Topical relevance

Literature review synthesis

Research Lines

End-to-end autonomous research labs

Automating the complete research pipeline, from ideation and literature analysis to experimental design, execution, and manuscript drafting, reducing human involvement.

Open edge: Lack of implemented prototypes with empirical validation; dependence on community consensus and resource investment; generalization beyond recommender systems or specific molecular dynamics tasks.
Digital-twin-driven autonomous decision-making

Predicting experimental outcomes, uncertainties, and risks to guide autonomous operations in microscopy and self-driving labs, enabling open decision-making and collaborative optimization.

Open edge: Validation limited to specific microscopy modalities or simple color-mixing tasks; noise dynamics and missing physical mechanisms still cause prediction residuals; scalability to complex scientific problems not demonstrated.
LLM-based reasoning and multi-agent workflows

Using LLM agents to propose hypotheses, edit code, assess novelty, and reason about experimental data, operationalizing serendipity and improving sampling efficiency.

Open edge: Trustworthiness of LLM reasoning and physical admissibility under distribution shift; overhead of multi-agent coordination; lack of standardized benchmarks for scientific reasoning.
Data quality and experimental infrastructure

Ensuring data reliability and reproducibility in self-driving labs through noise detection and correction (kNN) and FAIR data pipelines with automatic indexing.

Open edge: Methods are sensitive to feature distribution and noise type; not validated across diverse instrument modalities; reliance on specific infrastructure (nanoHUB) limits adoption.

Shared Direction

  • Integrating physical constraints or domain knowledge is essential for reliable autonomous experimentation.
  • LLMs and foundation models are becoming central components in automating scientific reasoning and experiment planning.
  • Data quality, FAIR principles, and reproducibility are critical bottlenecks for deploying self-driving labs in practice.
  • Standardized benchmarks and evaluation protocols are needed to compare different autonomous experimentation approaches.

Key Differences

  • Scope of automation: some works target full end-to-end autonomy, while others keep a human-in-the-loop for critical decisions or serendipity.
  • Decision mechanism: explicit world models (digital twins) vs. implicit LLM reasoning without explicit state representation.
  • Application domain: materials science imposes physical constraints and multi-modal data, whereas recommender systems face fairness and governance challenges.
  • Role of human guidance: SciLink actively integrates real-time expert feedback, whereas MLIPilot and AutoRecLab aim to minimize human intervention.

Open Questions

  • How can a unified evaluation framework balance scientific validity, computational cost, and reproducibility across diverse autonomous experimentation systems?
  • Are LLM reasoning trajectories reliable when encountering novel physical phenomena outside the training distribution?
  • What mechanisms ensure that autonomous labs produce repeatable and verifiable scientific results?
  • How should data ownership, privacy, and governance be managed in distributed collaborative self-driving laboratories?
  • Can noise detection and imputation methods generalize to highly heterogeneous and out-of-distribution experimental data?
LGAImtrl-sciCLdata-anchem-phCEIR

19 papers

Literature Review Agents

Papers and resources related to Literature Review Agents.

15 Methodology quality 89 Recency 36 Reproducibility 13 Topical relevance

Literature review synthesis

Research Lines

Literature review generation workflows

Automates organizing, writing, editing, and reviewing a literature review through specialized or iterative agent roles, aiming for readable and citation-accurate output.

Open edge: Generalization beyond evaluated datasets and baselines is not evidenced; several systems lack reported limitation details or code/dataset artifacts in the packet, leaving cross-system reproducibility a verification gap.
Literature database knowledge layers

Upgrades scientific literature databases into agent-accessible retrieval and evidence-extraction layers, enabling natural-language queries and evidence-supported responses.

Open edge: Documented gains are tied to Europe PMC or specific biology/QA settings; coverage across other domains and standardized cross-domain validation remain unshown.
Integrated scientific agent orchestration

Frames automated scientific discovery as an integrated loop spanning hypothesis formation, experimentation, interpretation, and literature evidence, with component-level automation treated as the near-term target.

Open edge: Full component integration remains unsolved, and the packet provides no quantitative discovery-level metric; the strongest claim is a 2050 goal rather than an achieved result.
Reliability and adversarial evaluation

Assesses vulnerability or robustness of automated reviewing and nearby agent pipelines under adversarial text attacks or benchmark stress.

Open edge: Adversarial robustness evidence is not yet connected to real-world review quality, expert acceptance, or decision utility; standardized evaluation across attacks and review outcomes is missing.

Shared Direction

  • LLM-orchestrated agents are used to plan complementary subqueries, extract evidence, and synthesize outputs rather than relying on a single retrieval step.
  • Citation quality is a recurring evaluation target, measured through citation F1 or agreement with expert consensus.
  • Literature review automation is treated as an integration problem spanning retrieval, grounding, writing, and review, with component integration remaining a central challenge.

Key Differences

  • System topology differs: single-LLM orchestration versus multi-agent role decomposition versus iterative or hierarchical organization.
  • Output target differs: a database knowledge layer for downstream agents versus a complete review manuscript versus peer-review robustness.
  • Evaluation settings differ: literature synthesis/claim verification/QA benchmarks, similarity to human-written reviews, and adversarial textual attacks are not directly comparable.

Open Questions

  • How do citation-precision improvements measured on specific benchmarks transfer to less curated or non-biomedical literature?
  • Which workflow choices beyond agent count and iteration depth contribute most to factual accuracy and readability?
  • What standardized protocol can compare evidence-grounded review generation with adversarial robustness and expert acceptance?
CL

13 papers

Paper Writing Agents

Papers and resources related to Paper Writing Agents.

1 Methodology quality 90 Recency 47 Reproducibility 6 Topical relevance

Literature review synthesis

Research Lines

Automated paper drafting and generation

Producing a complete research paper from a high-level idea or outline, reducing manual writing effort.

Open edge: Quality, originality, and factual reliability of generated content; lack of standardized evaluation against human-written papers.
Interactive writing assistance and tutoring

Providing real-time feedback, editing suggestions, and writing guidance within an editor to improve draft quality.

Open edge: Adaptation to diverse writing styles and disciplines; measurable impact on final paper acceptance or readability.
Paper review and critique

Generating or evaluating critiques of scientific papers, aiming to support or automate peer review.

Open edge: Groundedness of critiques beyond surface claims; coverage of nuanced methodological flaws and integration into real review workflows.
Multimodal paper transformation

Converting paper content into other formats such as posters, slides, or visual summaries.

Open edge: Design quality, content selection accuracy, and user customization; evaluation beyond aesthetic appeal.

Shared Direction

  • All works leverage large language models as the core reasoning engine for paper-related tasks.
  • Multi-agent architectures are widely adopted to decompose complex writing or analysis workflows.
  • The research is concentrated in 2025–2026, indicating a nascent but rapidly growing area.
  • There is a shared emphasis on integrating tools into existing academic environments (e.g., Overleaf) or providing reproducible artifacts.

Key Differences

  • Task scope: some systems target end-to-end paper generation, while others focus on editing, review, poster creation, or experiment reproduction.
  • Autonomy level: fully automated generation versus human-in-the-loop tutoring or assistance.
  • Interaction mode: in-editor plugins, standalone frameworks, or benchmark suites with no direct user interaction.
  • Evaluation target: output text quality, critique groundedness, reproduction success rate, or search recall/precision.

Open Questions

  • How can we evaluate the novelty and scientific contribution of automatically generated papers beyond surface-level metrics?
  • What mechanisms ensure factual accuracy, proper citation, and avoidance of plagiarism in generated content?
  • How do these agents perform on papers requiring deep domain expertise or novel experimental designs?
  • What are the ethical implications and potential for misuse in academic publishing, and how can they be mitigated?