RATIO
2026 Datasets
Dataset Analysis
Introduces RATIO, a benchmark for retrieving scientific literature by ideation operations (Address, Broaden, Specify), built via distant supervision and human/LLM vetting.
Provenance
Collected from papers.
Derived from paper: RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature
Related papers
RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable1GC-7RC: One Graphic Card -- Seven Research Challenges! How Good Are AI Agents at Doing Your Job?AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI AgentsAI Can Learn Scientific TasteAI Scientist via Synthetic Task ScalingAI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset ValuationAI Scientists as Engines of Discovery: A Case for Development within Reformed InstitutionsAI scientists produce results without reasoning scientificallyAI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World ModelARIS: Autonomous Research via Adversarial Multi-Agent CollaborationAct As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research LifecycleAgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement LearningAgentic AI Scientists Are Not Built For Autonomous Scientific DiscoveryAgentic Exploration of PDE Spaces using Latent Foundation Models for Parameterized SimulationsAgentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation AgentsAn Axiomatic Benchmark for Evaluation of Scientific Novelty MetricsAn Empirical Study of Multi-Agent Collaboration for Automated ResearchAuthor-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer ReviewAutoFigure-Edit: Generating Editable Scientific IllustrationAutoResearch AI: Towards AI-Powered Research Automation for Scientific DiscoveryAutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI CollaborationAutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model DiscoveryAutoTrainess: Teaching Language Models to Improve Language Models AutonomouslyAutonomous Scientific Discovery via Iterative Meta-ReflectionBeyond AI as Assistants: Toward Autonomous Discovery in CosmologyBioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge DiscoveryBloClaw: An Omniscient, Multi-Modal Agentic Workspace for Next-Generation Scientific DiscoveryCan Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert TaxonomiesCausalEvolve: Towards Open-Ended Discovery with Causal ScratchpadClarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific CollaborationClaw AI Lab: An Autonomous Multi-Agent Research TeamCompeting with AI Scientists: Agent-Driven Approach to Astrophysics ResearchDOVA: Deliberation-First Multi-Agent Orchestration for Autonomous Research AutomationDeciphering Scientific Reasoning Steps from Outcome Data for Molecule OptimizationDeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific SurveyDeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey GenerationDeterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics ArchitectureEarly Discoveries of Algorithmist I: Promise of Provable Algorithm Synthesis at ScaleEpistemic Uncertainty for Test-Time DiscoveryEurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific DiscoveryEvoMaster: A Foundational Evolving Agent Framework for Agentic Science at ScaleEvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific DiscoveryFlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature ExplorationFrom AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM AgentsFrom Passive Generation to Investigation: A Proactive Scientific Peer Review AgentGIANTS: Generative Insight Anticipation from Scientific LiteratureGenerating Literature-Driven Scientific Theories at ScaleGoodPoint: Learning Constructive Scientific Paper Feedback from Author ResponsesHLER: Human-in-the-Loop Economic Research via Multi-Agent Pipelines for Empirical DiscoveryHalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI ScientistsHarnessing AtomisticSkills for Agentic Atomistic ResearchInspectable AI for Science: A Research Object Approach to Generative AI GovernanceIntern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI ScientistsJoint discovery of governing partial differential equations from multi-source datasets by competitive optimizationLABBench2: An Improved Benchmark for AI Systems Performing Biology ResearchLECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction GenerationMIND: AI Co-Scientist for Material ResearchMVSS: A Unified Framework for Multi-View Structured Survey GenerationMany Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent SystemMeasuring Biological Capabilities and Risks of AI AgentsNORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data ScienceNanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research AutomationOR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm DiscoveryPRISM: Protocol Refinement through Intelligent Simulation ModelingPRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics ResearchPaperBanana: Automating Academic Illustration for AI ScientistsProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information DisclosureQuantifying the Reconstructability of Astrophysical Methods with Large Language Models and Information Theory: A Case Study in Spectral ReconstructionRESCORE: LLM-Driven Simulation Recovery in Control Systems Research PapersRbtAct: Rebuttal as Supervision for Actionable Review Feedback GenerationResearchEVO: An End-to-End Framework for Automated Scientific Discovery and DocumentationRethinking Publication: A Certification Framework for AI-Enabled ResearchRethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific DiscoverySciAtlas: A Large-Scale Knowledge Graph for Automated Scientific ResearchSciCoQA: Quality Assurance for Scientific Paper--Code AlignmentSciFig: Towards Automating Scientific Figure GenerationSciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse ParsingSciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery AgentsSocratic agents for autonomous scientific discovery in high-dimensional physical systemsSoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?Sparking Scientific Creativity via LLM-Driven Interdisciplinary InspirationSurveyLens: A Discipline-Aware Benchmark for Automatic Survey GenerationThe Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed ClaimsTianJi-Environ: An Autonomous AI Scientist for Atmospheric Environmental ResearchTowards Automating Scientific Review with Google's Paper Assistant ToolTowards End-to-End Automation of AI ResearchTowards a Medical AI ScientistUnderstanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction DatasetUnlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific LiteratureA Survey of AI ScientistsA collaborative digital twin built on FAIR data and compute infrastructureAI Idea Bench 2025: AI Research Idea Generation BenchmarkAI Scientists Fail Without Strong Implementation CapabilityAI Urban Scientist: Multi-Agent Collaborative Automation for Urban ResearchAI for Scientific Discovery is a Social ProblemAI-Driven Automation Can Become the Foundation of Next-Era Science of Science ResearchAISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data ScienceARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper GenerationAccelerating Scientific Discovery with Autonomous Goal-evolving AgentsAgentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future DirectionsAgentic AutoSurvey: Let LLMs Survey LLMsAligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection SamplingAutoEmpirical: LLM-Based Automated Research for Empirical Software Fault AnalysisAutoSurvey2: Empowering Researchers with Next Level Automated Literature SurveysBayes-Entropy Collaborative Driven Agents for Research Hypotheses Generation and OptimizationBeyond Optimization: Exploring Novelty Discovery in Autonomous ExperimentsBohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at ScaleBoxingGym: Benchmarking Progress in Automated Experimental Design and Model DiscoveryBuild Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science AutomationCASCADE: Cumulative Agentic Skill Creation through Autonomous Development and EvolutionCan Large Language Models Adequately Perform Symbolic Reasoning Over Time Series?Citegeist: Automated Generation of Related Work Analysis on the arXiv CorpusDeep Literature Survey Automation with an Iterative WorkflowDeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyDynamic Knowledge Exchange and Dual-diversity Review: Concisely Unleashing the Potential of a Multi-Agent Research TeamEnabling AI Scientists to Recognize Innovation: A Domain-Agnostic Algorithm for Assessing NoveltyEvaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and EvaluatorExploring Flow-Lenia Universes with a Curiosity-driven AI Scientist: Discovering Diverse Ecosystem DynamicsExploring the Limitations of kNN Noisy Feature Detection and Recovery for Self-Driving LabsFrom AutoRecSys to AutoRecLab: A Call to Build, Evaluate, and Govern Autonomous Recommender-Systems Research LabsHybridQuestion: Human-AI Collaboration for Identifying High-Impact Research QuestionsImpact of a Deployed LLM Survey Creation Tool through the IS Success ModelIn-situ graph reasoning and knowledge expansion using Graph-PReFLexORInteractiveSurvey: An LLM-based Personalized and Interactive Survey Paper Generation SystemJr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline PaperKnowledge Integration for Physics-informed Symbolic Regression Using Pre-trained Large Language ModelsKosmos: An AI Scientist for Autonomous DiscoveryLLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent SystemLLaMEA-BO: A Large Language Model Evolutionary Algorithm for Automatically Generating Bayesian Optimization AlgorithmsLarge Language Models for Scientific Idea Generation: A Creativity-Centered SurveyLiRA: A Multi-Agent Framework for Reliable and Readable Literature Review GenerationMIR: Methodology Inspiration Retrieval for Scientific Research ProblemsMachine Learning - Driven Materials Discovery: Unlocking Next-Generation Functional Materials - A reviewMegaScience: Pushing the Frontiers of Post-Training Datasets for Science ReasoningMeow: End-to-End Outline Writing for Automatic Academic SurveyMirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human ScientistsNeural surrogates for designing gravitational wave detectorsOmniScientist: Toward a Co-evolving Ecosystem of Human and AI ScientistsOpenLens AI: Fully Autonomous Research Agent for Health InfomaticsOperationalizing Serendipity: Multi-Agent AI Workflows for Enhanced Materials Characterization with Theory-in-the-LoopPhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics ResearchPiFlow: Principle-Aware Scientific Discovery with Multi-Agent CollaborationPosition: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AIREMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement LearningResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task DecompositionRobin: A multi-agent system for automating scientific discoverySGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation SystemsSR-Scientist: Scientific Equation Discovery With Agentic AISafeScientist: Toward Risk-Aware Scientific Discoveries by LLM AgentsScaling Laws in Scientific Discovery with AI and Robot ScientistsSciSage: A Multi-Agent Framework for High-Quality Scientific Survey GenerationSlideGen: Collaborative Multimodal Agents for Scientific Slide GenerationSpark: A System for Scientifically Creative Idea GenerationSpec-Driven AI for Science: The ARIA Framework for Automated and Reproducible Data AnalysisStructural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional ArchitectureSurGE: A Benchmark and Evaluation Framework for Scientific Survey GenerationSurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic SurveysSurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey WritingSurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey GenerationSurveyGen-I: Consistent Scientific Survey Generation with Evolving Plans and Memory-Guided WritingSurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsSurveyX: Academic Survey Automation via Large Language ModelsTaxoAlign: Scholarly Taxonomy Generation Using Language ModelsThe AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree SearchThe More You Automate, the Less You See: Hidden Pitfalls of AI Scientist SystemsVirtuous Machines: Towards Artificial General ScienceaiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists"Turing Tests" For An AI ScientistAiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image ClassificationBioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical ScienceCycleResearcher: Improving Automated Research via Automated ReviewEmpowering Biomedical Discovery with AI AgentsHow Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?Instruct Large Language Models to Generate Scientific Literature Survey Step by StepIntegration of Scanning Probe Microscope with High-Performance Computing: fixed-policy and reward-driven workflows implementationLiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal ContextMARG: Multi-Agent Review Generation for Scientific PapersMLR-Copilot: Autonomous Machine Learning Research based on Large Language Models AgentsMatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine CollaborationMeasurements with Noise: Bayesian Optimization for Co-optimizing Noise and Property Discovery in Automated ExperimentsResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language ModelsReviewer2: Optimizing Review Generation Through Prompt GenerationRisks of AI Scientists: Prioritizing Safeguarding Over AutonomyToward a Team of AI-made Scientists for Scientific Discovery from Gene Expression DataAI empowering research: 10 ways how science can benefit from AIAutoNMT: A Framework to Streamline the Research of Seq2Seq ModelsCan ChatGPT be used to generate scientific hypotheses?Deep Learning for Automated Experimentation in Scanning Transmission Electron MicroscopyLarge Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP ConceptsSciMON: Scientific Inspiration Machines Optimized for NoveltyTowards Autonomous Hypothesis Verification via Language Models with Minimal GuidanceGraph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual RecombinationGrounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physicsResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference OutcomesResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and BlogRethinking Scientific Discovery in an Agentic EraLearning to Trigger: Reinforcement Learning at the Large Hadron ColliderRethinking Scientific Discovery in the Agentic EraFirstResearch: Auditable Question Formation for LLM Scientific Discovery AgentsFrom Closed-Loop Optimization to Open Decision Making: Coupled Digital Twins for Predictive and Autonomous MicroscopyBibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and PublishingCausalDS: Benchmarking Causal Reasoning in Data-Science AgentsIdeas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea GenerationEvolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery SystemsToward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM AgentsGAE: Graph-Augmented Evolution for Scientific Discovery via Reinforcement OptimizationAre LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI ScientistsNVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case StudyTowards Autonomous and Auditable Medical Imaging Model DevelopmentOpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory SynthesisGigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearchReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent SystemAuto Research for Materials: Auditable AI-Scientist Workflows with Held-Out TransferPEARL: Auditable Repair for Scientific Reasoning Graph ExtractionSciForma: Structure-Faithful Generation of Scientific DiagramsAutomated Synthesis and Adversarial Validation of Executable Causal Research PipelinesIDEAgent: Agentic Quality-Diversity Search for Research Idea GenerationA Vocabulary for Multi-Agent Automated Research SystemsCausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal InferenceOmniQEC: discovering practical quantum error-correcting codes by an AI scientistLabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab AgentsIs Deep Research Reliable? Misleading Knowledge Induces False ConclusionsScaling Scientific Discovery Environments for Turn-Level Agentic RLEMBL AI Librarian: Life-Sciences Knowledge Layer for AI AgentsRecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender SystemsDeepCode: Open Agentic CodingGaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection PressureAutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchSpark-to-Paper: End-to-End Research Paper Generation as a Composable SkillDSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research AgentsAn AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research LoopAdversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist CapabilitiesEviGraph: Evidence-Guided Autonomous Research AgentsCapability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment SelectionAgentic Auto-Research is Fuzz TestingMechanist: AI as a Scientific Instrument for Discovering the Mechanisms of IntelligenceOmniScientist: An Omni-Modal Omni-Discipline AI ScientistIntern-S2-Preview: Scientific Agentic Foundation ModelHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer ReviewHow Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research TasksLarge Discovery Models: Empirically-grounded Model-Based Open-Ended SearchThe Problem Is the Problem: Towards Scalable Mathematical DiscoveryASI-Bench: At the Dawn of Artificial SuperintelligenceAutoResearch: Insight In, Hallucination OutFrontierChallenge: Evaluating Scientific Workflow CompletionAutonomous Mathematical Discovery in an Open-World Multi-Agent EnvironmentThe Past and Future of AI ScientistsSGHA: Evidence-Grounded Research Problem Discovery with Local Language ModelsSymposium: Trust via Auditable Records for Communities of AI Scientist AgentsHypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill LearningResearch Design Tracking and Assessment for the Social SciencesLearning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research AgentsPaperGym: Rubric-Centered Evolution for Research-Plan GenerationAutomated Researchers Can Reliably Mitigate Alignment Failures