README.md
April 4, 2026 · View on GitHub
Awesome-AI-Research is a curated repository for the AI-for-Research ecosystem. It groups autonomous research systems, infrastructure, workflow modules, benchmarks, surveys, datasets, and meta-resources into one layered structure.
Website: Awesome-AI-Research
Contents
- How to Use This List
- Feedback & Community
- Tag System
- 🧠 Autonomous Research Systems
- 🏗 Research Infrastructure & Platforms
- 🔬 Workflow Modules
- 📚 Benchmarks, Surveys & Meta-Resources
- Contributing
- License
How to Use This List
- Start with
Autonomous Research Systemsif you want loop-spanning AI scientist systems. - Jump to
Research Infrastructure & Platformsfor orchestration, runtimes, sandboxes, and observability layers. - Use
Workflow Moduleswhen you need stage-specific tools for literature, ideation, coding, experiments, analysis, or writing. - Use
Benchmarks, Surveys & Meta-Resourceswhen you need evaluation suites, datasets, or landscape overviews. - This list is curated for signal, not exhaustiveness: representative projects beat near-duplicate wrappers.
Feedback & Community
Share how often you use these AI research tools, how useful they feel in real research workflows, and where they still fall short. We are collecting responses to understand actual usage patterns, identify high-value tools, and improve the curation priorities of this repository.
Feishu Form: AI Research Tools Usage Survey
Tag System
README entries use a compact five-tag ribbon:
Level · Stage · Loop · Domain · Openness
The full curation model still tracks Scope and Maturity; full definitions live in docs/tag-system.md.
🧠 Autonomous Research Systems
Systems that attempt to cover meaningful parts of the scientific loop, from ideation to execution, analysis, or writing.
End-to-End AI Scientist Systems
-
The AI Scientist - Open-source end-to-end system that turns a seed codebase into ideas, experiments, figures, reviews, and a draft paper.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source -
AI-Researcher - Open-source autonomous research system for end-to-end scientific innovation, covering idea generation, implementation, experimentation, and paper writing.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source -
AutoResearchClaw - Fully autonomous research pipeline that turns a single research idea into literature-grounded experiments, analysis, and a conference-ready paper with OpenClaw-compatible execution.
Code · Docs ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source -
NanoResearch - End-to-end autonomous research engine that turns a topic into planned experiments, executed jobs, grounded analysis, and paper drafts backed by real run outputs.
Code · Docs ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source -
DeepScientist - Local-first AI research studio for turning papers or research goals into persistent quest repositories that support reproduction, experimentation, analysis, and paper drafting with human takeover anytime.
Code · Docs · Paper · Homepage ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Auto-Research - Prototype autonomous generalist scientist framework spanning literature review, proposal generation, experimentation, writing, submission, and review workflows.
Code ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Agent Laboratory - End-to-end autonomous research workflow with specialized agents for literature review, experimentation, and report writing, with optional co-pilot mode and AgentRxiv support.
Code · Paper · Homepage ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
MLR-Copilot - Autonomous machine learning research framework that generates research ideas, implements experiments, and executes them with iterative debugging and human feedback.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: CS·Openness: Open-source -
Robin - Multi-agent system from FutureHouse for automating scientific discovery, including candidate generation, assay planning, and optional experimental data analysis.
Code ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Biology·Openness: Open-source
Self-Improving / Self-Evolving Research Systems
-
The AI Scientist-v2 - Agentic-tree-search successor designed for workshop-level automated scientific discovery and higher-quality research trajectories.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source -
EvoScientist - Self-evolving AI scientist framework for end-to-end scientific discovery with iterative improvement and human oversight.
Code · Docs ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Scientify - AI-native scientific research system built around OpenClaw for automated literature review, experimentation, writing, and iterative research workflows.
Code · Docs ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
AutoResearch-SibylSystem - Fully autonomous AI scientist with 20+ agents and a dual-loop architecture for end-to-end ML research, GPU experiment execution, paper writing, and cross-project self-evolution.
Code · Docs ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: CS·Openness: Open-source -
AGI - Distributed peer-to-peer research network where autonomous agents run experiments, share results, synthesize papers, and compound discoveries across shared leaderboards.
Code ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: Multi-domain·Openness: Open-source
Closed-Loop Discovery Systems
-
Coscientist - Chemistry agent that plans syntheses, searches documentation, controls instruments, and iterates through experimental workflows.
Paper
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: Chemistry·Openness: Paper-only -
PiFlow - Principle-aware multi-agent framework for iterative scientific discovery across nanomaterials, biomolecules, and superconductors.
Code · Paper ·
Level: System·Stage: Experiment·Loop: Closed-loop·Domain: Multi-domain·Openness: Open-source -
Curie - Research experimentation agent for automating hypothesis formulation, implementation, execution, analysis, and reproducible reporting with built-in rigor checks.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source
Domain-Specific Autonomous Discovery Systems
-
ChemCrow - Tool-augmented chemistry agent that combines LLM reasoning with scientific software for synthesis and discovery tasks.
Code · Paper ·
Level: System·Stage: Planning·Loop: Human-in-the-loop·Domain: Chemistry·Openness: Open-source -
Biomni - General-purpose biomedical AI agent designed to autonomously execute a wide range of research tasks across biomedical subfields.
Code ·
Level: System·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Biology·Openness: Open-source -
ML-Agent - 7B reinforcement-trained agent for end-to-end machine learning engineering that learns from interactive experimentation on MLAgentBench and MLE-bench tasks.
Code · Paper ·
Level: System·Stage: End-to-end·Loop: Closed-loop·Domain: CS·Openness: Partially Open
🏗 Research Infrastructure & Platforms
The runtime substrate for AI-native research: orchestration, execution, memory, observability, and collaboration.
Research Platforms
-
FutureHouse - Research-native platform focused on automating scientific discovery with agentic systems and domain-aware tooling.
Homepage
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Closed-source -
AgentRxiv - Collaborative autonomous research framework and preprint layer where agent laboratories can publish, retrieve, and build on each other's findings.
Homepage · Paper
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Paper-only -
Research-Claw - Local-first academic research workspace with a dashboard, literature management, writing, experiment workflows, and OpenClaw-based extensions for always-on research assistance.
Code · Docs ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Partially Open -
ResearchClaw - Local-first AI research assistant for literature discovery, notes, experiment tracking, and paper writing across the scientific workflow.
Code · Docs ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source
Agent Runtimes & Orchestration
-
LangGraph - Stateful graph runtime for long-running agent workflows with branching, memory, recovery, and explicit control flow.
Docs · Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
AutoGen - Multi-agent programming framework widely used to build research copilots, literature agents, and evaluation pipelines.
Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
AgentScope - Agent-oriented programming framework for composing multi-agent workflows with explicit roles, collaboration patterns, and tool integration.
Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
ClawTeam - Swarm-intelligence orchestration layer where leader agents spawn specialized workers across worktrees, tmux sessions, and GPUs for autonomous research and engineering workflows.
Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
LDP - Framework for modular interchange of language agents, environments, and optimizers.
Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source
Research Workflow Orchestration
-
ARIS: Auto-claude-code-research-in-sleep - Lightweight Markdown-only research workflow kit for idea discovery, cross-model review loops, experiment automation, and paper writing across Claude Code, Codex, OpenClaw, and similar agents.
Code · Docs ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
AI Research Skills Library - Open-source skills library plus autoresearch orchestration layer that gives coding agents reusable research skills spanning literature survey, experimentation, and paper writing.
Code ·
Level: Platform·Stage: End-to-end·Loop: Human-in-the-loop·Domain: CS·Openness: Open-source -
autocontext - Workflow harness for running scenarios, tasks, and missions while carrying forward validated playbooks, artifacts, and distilled knowledge across repeated agent runs.
Code ·
Level: Platform·Stage: End-to-end·Loop: Closed-loop·Domain: General·Openness: Open-source
Tool-Use & Execution Infrastructure
- E2B - Sandboxed execution layer for code, browser, and desktop-style tool use inside AI-driven research workflows.
Homepage · Docs · Code ·
Level: Platform·Stage: Data·Loop: Human-in-the-loop·Domain: General·Openness: Partially Open
Evaluation & Training Environments
-
Aviary - Language-agent gym with built-in scientific task environments, including scientific literature search and protein stability.
Code ·
Level: Platform·Stage: End-to-end·Loop: Open-loop·Domain: Multi-domain·Openness: Open-source -
DiscoveryWorld - Virtual environment for building and evaluating automated scientific discovery agents with interactive tasks, scorecards, and baseline agents.
Code · Paper · Homepage ·
Level: Platform·Stage: End-to-end·Loop: Open-loop·Domain: Multi-domain·Openness: Open-source -
MLGym - Gym environment and benchmark for open-ended AI research tasks that require ideation, coding, experimentation, analysis, and iteration.
Code · Paper · Homepage ·
Level: Platform·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Open-source
Memory / Observability / Collaboration Layers
- Weights & Biases - Experiment tracking and collaboration layer for instrumenting long-running research agents, ablations, and benchmark runs.
Homepage · Docs
Level: Platform·Stage: Analysis·Loop: Human-in-the-loop·Domain: General·Openness: Partially Open
🔬 Workflow Modules
Stage-specific building blocks for literature review, ideation, planning, coding, experimentation, analysis, and writing.
Literature Discovery & Review
-
Elicit - AI research assistant for literature search, evidence extraction, and structured review workflows.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
PaperQA2 - Open-source literature QA and evidence-synthesis stack optimized for scientific documents and citation-grounded answers.
Code · Paper ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
LatteReview - Multi-agent literature review framework for title and abstract screening, evidence abstraction, multimodal review, and customizable reviewer workflows.
Code · Paper · Docs ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
LitLLM - AI-powered literature review assistant that combines hybrid retrieval, re-ranking, and structured generation for related-work drafting.
Code · Paper · Homepage ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
GPT Researcher - Open deep research agent for web and local documents that uses planner and execution agents to produce citation-backed research reports.
Code · Docs · Homepage ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
OpenScholar - Retrieval-augmented language model for searching scientific literature and generating grounded synthesis answers from relevant papers.
Code · Paper · Homepage ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
OpenResearcher - Scientific research assistant with access to the arXiv corpus for answering research queries using retrieval, embeddings, and web search.
Code · Paper ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
PaSa - LLM paper-search agent that autonomously searches, reads, expands citations, and selects relevant scholarly references for complex academic queries.
Code · Paper · Homepage ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
AutoResearcher - Early-stage open-source Python package for automating scientific workflows, currently focused on literature reviews with a longer-term goal of autonomous discovery.
Code · Docs ·
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
ResearchRabbit - Visual citation-graph exploration tool for expanding seed papers into neighborhoods of related work.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
Litmaps - Literature discovery and monitoring tool built around citation-network exploration, visualization, and alerts.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
Connected Papers - Visual paper exploration tool for finding related academic work around a seed paper.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
Scite - Research discovery and evaluation platform centered on Smart Citations, showing whether studies support, contrast, or mention prior work.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
SciSpace Literature Review - AI literature review workspace for finding, analyzing, organizing, and comparing scientific papers.
Homepage
Level: Module·Stage: Literature·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source
Research Ideation & Hypothesis Generation
-
AI co-scientist - Google Research's multi-agent scientific collaborator for proposing, debating, ranking, and refining hypotheses with human oversight.
Homepage · Paper
Level: Module·Stage: Ideation·Loop: Human-in-the-loop·Domain: General·Openness: Paper-only -
ResearchAgent - Iterative research idea generation system that retrieves scientific literature and refines candidate problems, methods, and experiment designs with LLM reviewer feedback.
Code · Paper ·
Level: Module·Stage: Ideation·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Consensus - Scientific search engine geared toward claim-grounded answers, useful for scoping evidence and framing candidate hypotheses.
Homepage
Level: Module·Stage: Ideation·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source
Planning & Experimental Design
-
ChemCrow - Tool-augmented chemistry agent that combines LLM reasoning with scientific software for synthesis and discovery tasks.
Code · Paper ·
Level: Module·Stage: Planning·Loop: Human-in-the-loop·Domain: Chemistry·Openness: Open-source -
Coscientist - Chemistry agent that plans syntheses, searches documentation, controls instruments, and iterates through experimental workflows.
Paper
Level: Module·Stage: Planning·Loop: Closed-loop·Domain: Chemistry·Openness: Paper-only
Data, Environment & Tool Use
-
E2B - Sandboxed execution layer for code, browser, and desktop-style tool use inside AI-driven research workflows.
Homepage · Docs · Code ·
Level: Module·Stage: Data·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
GeneAgent - Domain-database agent for gene set analysis that shows how scientific tool use can be grounded in external biomedical resources.
Code ·
Level: Module·Stage: Data·Loop: Human-in-the-loop·Domain: Biology·Openness: Open-source
Method Development & Research Coding
-
AutoGen - Multi-agent programming framework widely used to build research copilots, literature agents, and evaluation pipelines.
Code ·
Level: Module·Stage: Coding·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
OpenHands - Agent runtime for repo-level coding, execution, and issue-driven engineering that adapts well to research coding workflows.
Code · Docs ·
Level: Module·Stage: Coding·Loop: Human-in-the-loop·Domain: General·Openness: Open-source
Experiment Execution & Optimization
-
Optuna - Open-source optimization framework for trial scheduling, hyperparameter search, and controlled experiment iteration.
Homepage · Code ·
Level: Module·Stage: Experiment·Loop: Open-loop·Domain: General·Openness: Open-source -
AIDE ML - Tree-search machine learning engineering agent that iteratively drafts, debugs, benchmarks, and improves code against a target metric.
Code · Paper ·
Level: Module·Stage: Experiment·Loop: Closed-loop·Domain: CS·Openness: Open-source -
PiFlow - Principle-aware multi-agent framework for iterative scientific discovery across nanomaterials, biomolecules, and superconductors.
Code · Paper ·
Level: Module·Stage: Experiment·Loop: Closed-loop·Domain: Multi-domain·Openness: Open-source
Analysis, Evaluation & Interpretation
-
PaperBench - Benchmark that evaluates whether agents can reproduce frontier AI research workflows from understanding claims to running experiments.
Homepage
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Partially Open -
Weights & Biases - Experiment tracking and collaboration layer for instrumenting long-running research agents, ablations, and benchmark runs.
Homepage · Docs
Level: Module·Stage: Analysis·Loop: Human-in-the-loop·Domain: General·Openness: Partially Open
Writing, Publication & Communication
-
Overleaf AI - AI-assisted writing and editing features inside a collaborative LaTeX environment used heavily in academic publication workflows.
Homepage
Level: Module·Stage: Writing·Loop: Human-in-the-loop·Domain: General·Openness: Closed-source -
STORM - Open-source system for grounded long-form report generation with citation-backed outlining and drafting.
Code ·
Level: Module·Stage: Writing·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
SurveyX - Academic survey automation system for generating domain-specific survey papers from user-provided references and LLM-driven writing pipelines.
Code · Paper · Homepage ·
Level: Module·Stage: Writing·Loop: Human-in-the-loop·Domain: General·Openness: Partially Open -
AutoSurvey - Framework for automatically writing long-form literature surveys with structured outlines, retrieval, and citation-aware generation.
Code · Paper ·
Level: Module·Stage: Writing·Loop: Human-in-the-loop·Domain: General·Openness: Open-source
📚 Benchmarks, Surveys & Meta-Resources
Benchmarks, surveys, datasets, and other reference layers that keep the ecosystem legible and comparable.
Surveys & Taxonomies
-
A Survey of AI Scientists - Survey focused on automatic scientists and end-to-end AI research pipelines.
Paper
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Paper-only -
AI4Research - Living survey site mapping AI for scientific research across domains, tasks, and papers.
Homepage
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Partially Open -
AI-for-Research - Repository accompanying a survey on AI support across the research lifecycle from hypothesis generation to publication.
Code · Paper ·
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Awesome Deep Research Agent - Curated map of deep research agent papers, architectures, tool-use methods, and benchmarks, anchored by a dedicated survey and roadmap.
Code · Paper ·
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
LLM-Agent-Optimization - Reading list for a survey on optimizing LLM-based agents, covering fine-tuning, reflection, tool use, datasets, and real-world applications.
Code ·
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Awesome LLM Scientific Discovery - Curated list and taxonomy for LLMs in scientific discovery, organized by autonomy level from tool to analyst to scientist.
Code · Paper ·
Level: Survey·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Open-source
Benchmarks & Evaluation Suites
-
Frontiers in Science - Benchmark suite for evaluating scientific reasoning across olympiad-style and research-style tasks.
Homepage
Level: Benchmark·Stage: Analysis·Loop: Open-loop·Domain: Multi-domain·Openness: Partially Open -
PaperBench - Benchmark that evaluates whether agents can reproduce frontier AI research workflows from understanding claims to running experiments.
Homepage
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Partially Open -
ScienceBench - Autonomous laboratory benchmark for end-to-end scientific operation and discovery with minimal human oversight.
Homepage
Level: Benchmark·Stage: End-to-end·Loop: Closed-loop·Domain: Multi-domain·Openness: Partially Open -
AIRS-Bench - Benchmark for quantifying the end-to-end AI research abilities of LLM agents on machine learning tasks drawn from recent papers.
Code ·
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Open-source -
MLAgentBench - Suite of end-to-end machine learning experimentation tasks where agents autonomously develop and improve models from dataset and task descriptions.
Code ·
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Open-source -
ML-Bench - Benchmark for large language models and agents on end-to-end machine learning workflows over repository-level code, including ML-LLM-Bench and ML-Agent-Bench tracks.
Code · Paper ·
Level: Benchmark·Stage: End-to-end·Loop: Human-in-the-loop·Domain: CS·Openness: Open-source -
MLE-bench - Benchmark for machine learning engineering agents, including task construction, evaluation logic, and reference agent implementations.
Code ·
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Open-source -
MLR-Bench - Benchmark for open-ended machine learning research with 201 tasks, together with agent and judge baselines for evaluation.
Code ·
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: CS·Openness: Open-source -
LAB-Bench - Biology benchmark for capabilities foundational to scientific research in biology.
Code ·
Level: Benchmark·Stage: End-to-end·Loop: Open-loop·Domain: Biology·Openness: Open-source -
AgentBench - General benchmark for evaluating LLMs as agents across diverse interactive environments such as operating systems, databases, web tasks, and knowledge graphs.
Code · Paper ·
Level: Benchmark·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
ScienceAgentBench - Benchmark for language agents on 102 expert-validated tasks from real scientific workflows across multiple disciplines.
Code · Paper · Homepage ·
Level: Benchmark·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Partially Open -
ScholarQABench - Holistic benchmark for testing scientific literature synthesis, citation accuracy, coverage, relevance, and organization in long-form scholarly answers.
Code · Paper ·
Level: Benchmark·Stage: Literature·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Open-source -
RE-Bench - Task suite for evaluating frontier AI R&D capabilities of language-model agents against human experts on realistic research engineering tasks.
Code · Paper ·
Level: Benchmark·Stage: End-to-end·Loop: Human-in-the-loop·Domain: CS·Openness: Partially Open -
EXP-Bench - Benchmark for assessing whether AI agents can conduct AI research experiments with rigorous implementation, execution, and analysis loops.
Paper · Docs
Level: Benchmark·Stage: Experiment·Loop: Human-in-the-loop·Domain: CS·Openness: Partially Open -
ResearchBench - Benchmark for scientific discovery via inspiration retrieval, hypothesis composition, and hypothesis ranking across multiple disciplines.
Paper
Level: Benchmark·Stage: Ideation·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Paper-only -
IdeaBench - Benchmarking framework for evaluating large language models on research idea generation quality, novelty, and relevance.
Paper
Level: Benchmark·Stage: Ideation·Loop: Human-in-the-loop·Domain: General·Openness: Paper-only -
LiveIdeaBench - Benchmark for evaluating scientific creativity and divergent thinking in LLM-generated ideas using minimal-context prompts.
Homepage · Paper
Level: Benchmark·Stage: Ideation·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Partially Open
Datasets
-
OpenAlex - Open index of works, authors, venues, institutions, and concepts that underpins many research-native retrieval systems.
Homepage · Docs
Level: Dataset·Stage: Literature·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Partially Open -
Semantic Scholar Academic Graph API - Structured paper metadata and graph endpoints for retrieval, paper linking, and citation analysis.
Docs
Level: Dataset·Stage: Literature·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Partially Open
Other Awesome Lists
-
awesome-research - Curated collection organized around research workflow tasks, useful as a complementary module-first view.
Code ·
Level: Meta·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Awesome-AI-Scientists - Curated list centered on AI Scientist systems, complementary to this repository's broader layered-map perspective.
Code ·
Level: Meta·Stage: End-to-end·Loop: Human-in-the-loop·Domain: General·Openness: Open-source -
Awesome Papers on Agents for Science - Curated bibliography of agents-for-science papers across domains, task types, and benchmarks.
Code ·
Level: Meta·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Open-source -
Awesome AI Scientist Papers - Curated bibliography of AI Scientist and Robot Scientist papers, surveys, benchmarks, timelines, and community resources.
Code ·
Level: Meta·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Open-source -
Awesome AutoResearch - Curated list of autoresearch use cases, implementations, and optimization traces for keep-or-revert research loops across domains.
Code ·
Level: Meta·Stage: End-to-end·Loop: Human-in-the-loop·Domain: Multi-domain·Openness: Open-source
Contributing
Please read CONTRIBUTING.md, docs/inclusion-criteria.md, and docs/tag-system.md before submitting new entries.
Recommended inclusion standards:
- The project should be clearly relevant to scientific research or AI-for-Research workflows.
- Prefer official project pages, official GitHub repositories, or papers over secondary summaries.
- Avoid adding generic AI tools unless they are widely used as research-native infrastructure or workflow components.
- Prefer representative systems over near-duplicate wrappers.