README.md

April 4, 2026 · View on GitHub

Awesome-AI-Research logo

Awesome-AI-Research

English · 中文

Awesome Website PRs Welcome License: MIT

A curated, layered map of AI-native research systems.

Awesome-AI-Research is a curated repository for the AI-for-Research ecosystem. It groups autonomous research systems, infrastructure, workflow modules, benchmarks, surveys, datasets, and meta-resources into one layered structure.

Website: Awesome-AI-Research

Contents

How to Use This List

  • Start with Autonomous Research Systems if you want loop-spanning AI scientist systems.
  • Jump to Research Infrastructure & Platforms for orchestration, runtimes, sandboxes, and observability layers.
  • Use Workflow Modules when you need stage-specific tools for literature, ideation, coding, experiments, analysis, or writing.
  • Use Benchmarks, Surveys & Meta-Resources when you need evaluation suites, datasets, or landscape overviews.
  • This list is curated for signal, not exhaustiveness: representative projects beat near-duplicate wrappers.

Feedback & Community

Share how often you use these AI research tools, how useful they feel in real research workflows, and where they still fall short. We are collecting responses to understand actual usage patterns, identify high-value tools, and improve the curation priorities of this repository.

Feishu Form: AI Research Tools Usage Survey

Tag System

README entries use a compact five-tag ribbon:

Level · Stage · Loop · Domain · Openness

The full curation model still tracks Scope and Maturity; full definitions live in docs/tag-system.md.

🧠 Autonomous Research Systems

Systems that attempt to cover meaningful parts of the scientific loop, from ideation to execution, analysis, or writing.

End-to-End AI Scientist Systems

  • The AI Scientist - Open-source end-to-end system that turns a seed codebase into ideas, experiments, figures, reviews, and a draft paper.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

  • AI-Researcher - Open-source autonomous research system for end-to-end scientific innovation, covering idea generation, implementation, experimentation, and paper writing.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

  • AutoResearchClaw - Fully autonomous research pipeline that turns a single research idea into literature-grounded experiments, analysis, and a conference-ready paper with OpenClaw-compatible execution.
    Code · Docs · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

  • NanoResearch - End-to-end autonomous research engine that turns a topic into planned experiments, executed jobs, grounded analysis, and paper drafts backed by real run outputs.
    Code · Docs · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

  • DeepScientist - Local-first AI research studio for turning papers or research goals into persistent quest repositories that support reproduction, experimentation, analysis, and paper drafting with human takeover anytime.
    Code · Docs · Paper · Homepage · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Auto-Research - Prototype autonomous generalist scientist framework spanning literature review, proposal generation, experimentation, writing, submission, and review workflows.
    Code · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Agent Laboratory - End-to-end autonomous research workflow with specialized agents for literature review, experimentation, and report writing, with optional co-pilot mode and AgentRxiv support.
    Code · Paper · Homepage · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • MLR-Copilot - Autonomous machine learning research framework that generates research ideas, implements experiments, and executes them with iterative debugging and human feedback.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: CS · Openness: Open-source

  • Robin - Multi-agent system from FutureHouse for automating scientific discovery, including candidate generation, assay planning, and optional experimental data analysis.
    Code · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Biology · Openness: Open-source

Self-Improving / Self-Evolving Research Systems

  • The AI Scientist-v2 - Agentic-tree-search successor designed for workshop-level automated scientific discovery and higher-quality research trajectories.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

  • EvoScientist - Self-evolving AI scientist framework for end-to-end scientific discovery with iterative improvement and human oversight.
    Code · Docs · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Scientify - AI-native scientific research system built around OpenClaw for automated literature review, experimentation, writing, and iterative research workflows.
    Code · Docs · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • AutoResearch-SibylSystem - Fully autonomous AI scientist with 20+ agents and a dual-loop architecture for end-to-end ML research, GPU experiment execution, paper writing, and cross-project self-evolution.
    Code · Docs · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: CS · Openness: Open-source

  • AGI - Distributed peer-to-peer research network where autonomous agents run experiments, share results, synthesize papers, and compound discoveries across shared leaderboards.
    Code · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: Multi-domain · Openness: Open-source

Closed-Loop Discovery Systems

  • Coscientist - Chemistry agent that plans syntheses, searches documentation, controls instruments, and iterates through experimental workflows.
    Paper
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: Chemistry · Openness: Paper-only

  • PiFlow - Principle-aware multi-agent framework for iterative scientific discovery across nanomaterials, biomolecules, and superconductors.
    Code · Paper · GitHub stars
    Level: System · Stage: Experiment · Loop: Closed-loop · Domain: Multi-domain · Openness: Open-source

  • Curie - Research experimentation agent for automating hypothesis formulation, implementation, execution, analysis, and reproducible reporting with built-in rigor checks.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

Domain-Specific Autonomous Discovery Systems

  • ChemCrow - Tool-augmented chemistry agent that combines LLM reasoning with scientific software for synthesis and discovery tasks.
    Code · Paper · GitHub stars
    Level: System · Stage: Planning · Loop: Human-in-the-loop · Domain: Chemistry · Openness: Open-source

  • Biomni - General-purpose biomedical AI agent designed to autonomously execute a wide range of research tasks across biomedical subfields.
    Code · GitHub stars
    Level: System · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Biology · Openness: Open-source

  • ML-Agent - 7B reinforcement-trained agent for end-to-end machine learning engineering that learns from interactive experimentation on MLAgentBench and MLE-bench tasks.
    Code · Paper · GitHub stars
    Level: System · Stage: End-to-end · Loop: Closed-loop · Domain: CS · Openness: Partially Open

🏗 Research Infrastructure & Platforms

The runtime substrate for AI-native research: orchestration, execution, memory, observability, and collaboration.

Research Platforms

  • FutureHouse - Research-native platform focused on automating scientific discovery with agentic systems and domain-aware tooling.
    Homepage
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Closed-source

  • AgentRxiv - Collaborative autonomous research framework and preprint layer where agent laboratories can publish, retrieve, and build on each other's findings.
    Homepage · Paper
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Paper-only

  • Research-Claw - Local-first academic research workspace with a dashboard, literature management, writing, experiment workflows, and OpenClaw-based extensions for always-on research assistance.
    Code · Docs · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Partially Open

  • ResearchClaw - Local-first AI research assistant for literature discovery, notes, experiment tracking, and paper writing across the scientific workflow.
    Code · Docs · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

Agent Runtimes & Orchestration

  • LangGraph - Stateful graph runtime for long-running agent workflows with branching, memory, recovery, and explicit control flow.
    Docs · Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • AutoGen - Multi-agent programming framework widely used to build research copilots, literature agents, and evaluation pipelines.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • AgentScope - Agent-oriented programming framework for composing multi-agent workflows with explicit roles, collaboration patterns, and tool integration.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • ClawTeam - Swarm-intelligence orchestration layer where leader agents spawn specialized workers across worktrees, tmux sessions, and GPUs for autonomous research and engineering workflows.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • LDP - Framework for modular interchange of language agents, environments, and optimizers.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

Research Workflow Orchestration

  • ARIS: Auto-claude-code-research-in-sleep - Lightweight Markdown-only research workflow kit for idea discovery, cross-model review loops, experiment automation, and paper writing across Claude Code, Codex, OpenClaw, and similar agents.
    Code · Docs · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • AI Research Skills Library - Open-source skills library plus autoresearch orchestration layer that gives coding agents reusable research skills spanning literature survey, experimentation, and paper writing.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Human-in-the-loop · Domain: CS · Openness: Open-source

  • autocontext - Workflow harness for running scenarios, tasks, and missions while carrying forward validated playbooks, artifacts, and distilled knowledge across repeated agent runs.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Closed-loop · Domain: General · Openness: Open-source

Tool-Use & Execution Infrastructure

  • E2B - Sandboxed execution layer for code, browser, and desktop-style tool use inside AI-driven research workflows.
    Homepage · Docs · Code · GitHub stars
    Level: Platform · Stage: Data · Loop: Human-in-the-loop · Domain: General · Openness: Partially Open

Evaluation & Training Environments

  • Aviary - Language-agent gym with built-in scientific task environments, including scientific literature search and protein stability.
    Code · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Open-loop · Domain: Multi-domain · Openness: Open-source

  • DiscoveryWorld - Virtual environment for building and evaluating automated scientific discovery agents with interactive tasks, scorecards, and baseline agents.
    Code · Paper · Homepage · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Open-loop · Domain: Multi-domain · Openness: Open-source

  • MLGym - Gym environment and benchmark for open-ended AI research tasks that require ideation, coding, experimentation, analysis, and iteration.
    Code · Paper · Homepage · GitHub stars
    Level: Platform · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Open-source

Memory / Observability / Collaboration Layers

  • Weights & Biases - Experiment tracking and collaboration layer for instrumenting long-running research agents, ablations, and benchmark runs.
    Homepage · Docs
    Level: Platform · Stage: Analysis · Loop: Human-in-the-loop · Domain: General · Openness: Partially Open

🔬 Workflow Modules

Stage-specific building blocks for literature review, ideation, planning, coding, experimentation, analysis, and writing.

Literature Discovery & Review

  • Elicit - AI research assistant for literature search, evidence extraction, and structured review workflows.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • PaperQA2 - Open-source literature QA and evidence-synthesis stack optimized for scientific documents and citation-grounded answers.
    Code · Paper · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • LatteReview - Multi-agent literature review framework for title and abstract screening, evidence abstraction, multimodal review, and customizable reviewer workflows.
    Code · Paper · Docs · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • LitLLM - AI-powered literature review assistant that combines hybrid retrieval, re-ranking, and structured generation for related-work drafting.
    Code · Paper · Homepage · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • GPT Researcher - Open deep research agent for web and local documents that uses planner and execution agents to produce citation-backed research reports.
    Code · Docs · Homepage · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • OpenScholar - Retrieval-augmented language model for searching scientific literature and generating grounded synthesis answers from relevant papers.
    Code · Paper · Homepage · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • OpenResearcher - Scientific research assistant with access to the arXiv corpus for answering research queries using retrieval, embeddings, and web search.
    Code · Paper · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • PaSa - LLM paper-search agent that autonomously searches, reads, expands citations, and selects relevant scholarly references for complex academic queries.
    Code · Paper · Homepage · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • AutoResearcher - Early-stage open-source Python package for automating scientific workflows, currently focused on literature reviews with a longer-term goal of autonomous discovery.
    Code · Docs · GitHub stars
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • ResearchRabbit - Visual citation-graph exploration tool for expanding seed papers into neighborhoods of related work.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • Litmaps - Literature discovery and monitoring tool built around citation-network exploration, visualization, and alerts.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • Connected Papers - Visual paper exploration tool for finding related academic work around a seed paper.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • Scite - Research discovery and evaluation platform centered on Smart Citations, showing whether studies support, contrast, or mention prior work.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • SciSpace Literature Review - AI literature review workspace for finding, analyzing, organizing, and comparing scientific papers.
    Homepage
    Level: Module · Stage: Literature · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

Research Ideation & Hypothesis Generation

  • AI co-scientist - Google Research's multi-agent scientific collaborator for proposing, debating, ranking, and refining hypotheses with human oversight.
    Homepage · Paper
    Level: Module · Stage: Ideation · Loop: Human-in-the-loop · Domain: General · Openness: Paper-only

  • ResearchAgent - Iterative research idea generation system that retrieves scientific literature and refines candidate problems, methods, and experiment designs with LLM reviewer feedback.
    Code · Paper · GitHub stars
    Level: Module · Stage: Ideation · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Consensus - Scientific search engine geared toward claim-grounded answers, useful for scoping evidence and framing candidate hypotheses.
    Homepage
    Level: Module · Stage: Ideation · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

Planning & Experimental Design

  • ChemCrow - Tool-augmented chemistry agent that combines LLM reasoning with scientific software for synthesis and discovery tasks.
    Code · Paper · GitHub stars
    Level: Module · Stage: Planning · Loop: Human-in-the-loop · Domain: Chemistry · Openness: Open-source

  • Coscientist - Chemistry agent that plans syntheses, searches documentation, controls instruments, and iterates through experimental workflows.
    Paper
    Level: Module · Stage: Planning · Loop: Closed-loop · Domain: Chemistry · Openness: Paper-only

Data, Environment & Tool Use

  • E2B - Sandboxed execution layer for code, browser, and desktop-style tool use inside AI-driven research workflows.
    Homepage · Docs · Code · GitHub stars
    Level: Module · Stage: Data · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • GeneAgent - Domain-database agent for gene set analysis that shows how scientific tool use can be grounded in external biomedical resources.
    Code · GitHub stars
    Level: Module · Stage: Data · Loop: Human-in-the-loop · Domain: Biology · Openness: Open-source

Method Development & Research Coding

  • AutoGen - Multi-agent programming framework widely used to build research copilots, literature agents, and evaluation pipelines.
    Code · GitHub stars
    Level: Module · Stage: Coding · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • OpenHands - Agent runtime for repo-level coding, execution, and issue-driven engineering that adapts well to research coding workflows.
    Code · Docs · GitHub stars
    Level: Module · Stage: Coding · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

Experiment Execution & Optimization

  • Optuna - Open-source optimization framework for trial scheduling, hyperparameter search, and controlled experiment iteration.
    Homepage · Code · GitHub stars
    Level: Module · Stage: Experiment · Loop: Open-loop · Domain: General · Openness: Open-source

  • AIDE ML - Tree-search machine learning engineering agent that iteratively drafts, debugs, benchmarks, and improves code against a target metric.
    Code · Paper · GitHub stars
    Level: Module · Stage: Experiment · Loop: Closed-loop · Domain: CS · Openness: Open-source

  • PiFlow - Principle-aware multi-agent framework for iterative scientific discovery across nanomaterials, biomolecules, and superconductors.
    Code · Paper · GitHub stars
    Level: Module · Stage: Experiment · Loop: Closed-loop · Domain: Multi-domain · Openness: Open-source

Analysis, Evaluation & Interpretation

  • PaperBench - Benchmark that evaluates whether agents can reproduce frontier AI research workflows from understanding claims to running experiments.
    Homepage
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Partially Open

  • Weights & Biases - Experiment tracking and collaboration layer for instrumenting long-running research agents, ablations, and benchmark runs.
    Homepage · Docs
    Level: Module · Stage: Analysis · Loop: Human-in-the-loop · Domain: General · Openness: Partially Open

Writing, Publication & Communication

  • Overleaf AI - AI-assisted writing and editing features inside a collaborative LaTeX environment used heavily in academic publication workflows.
    Homepage
    Level: Module · Stage: Writing · Loop: Human-in-the-loop · Domain: General · Openness: Closed-source

  • STORM - Open-source system for grounded long-form report generation with citation-backed outlining and drafting.
    Code · GitHub stars
    Level: Module · Stage: Writing · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • SurveyX - Academic survey automation system for generating domain-specific survey papers from user-provided references and LLM-driven writing pipelines.
    Code · Paper · Homepage · GitHub stars
    Level: Module · Stage: Writing · Loop: Human-in-the-loop · Domain: General · Openness: Partially Open

  • AutoSurvey - Framework for automatically writing long-form literature surveys with structured outlines, retrieval, and citation-aware generation.
    Code · Paper · GitHub stars
    Level: Module · Stage: Writing · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

📚 Benchmarks, Surveys & Meta-Resources

Benchmarks, surveys, datasets, and other reference layers that keep the ecosystem legible and comparable.

Surveys & Taxonomies

  • A Survey of AI Scientists - Survey focused on automatic scientists and end-to-end AI research pipelines.
    Paper
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Paper-only

  • AI4Research - Living survey site mapping AI for scientific research across domains, tasks, and papers.
    Homepage
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Partially Open

  • AI-for-Research - Repository accompanying a survey on AI support across the research lifecycle from hypothesis generation to publication.
    Code · Paper · GitHub stars
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Awesome Deep Research Agent - Curated map of deep research agent papers, architectures, tool-use methods, and benchmarks, anchored by a dedicated survey and roadmap.
    Code · Paper · GitHub stars
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • LLM-Agent-Optimization - Reading list for a survey on optimizing LLM-based agents, covering fine-tuning, reflection, tool use, datasets, and real-world applications.
    Code · GitHub stars
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Awesome LLM Scientific Discovery - Curated list and taxonomy for LLMs in scientific discovery, organized by autonomy level from tool to analyst to scientist.
    Code · Paper · GitHub stars
    Level: Survey · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Open-source

Benchmarks & Evaluation Suites

  • Frontiers in Science - Benchmark suite for evaluating scientific reasoning across olympiad-style and research-style tasks.
    Homepage
    Level: Benchmark · Stage: Analysis · Loop: Open-loop · Domain: Multi-domain · Openness: Partially Open

  • PaperBench - Benchmark that evaluates whether agents can reproduce frontier AI research workflows from understanding claims to running experiments.
    Homepage
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Partially Open

  • ScienceBench - Autonomous laboratory benchmark for end-to-end scientific operation and discovery with minimal human oversight.
    Homepage
    Level: Benchmark · Stage: End-to-end · Loop: Closed-loop · Domain: Multi-domain · Openness: Partially Open

  • AIRS-Bench - Benchmark for quantifying the end-to-end AI research abilities of LLM agents on machine learning tasks drawn from recent papers.
    Code · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Open-source

  • MLAgentBench - Suite of end-to-end machine learning experimentation tasks where agents autonomously develop and improve models from dataset and task descriptions.
    Code · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Open-source

  • ML-Bench - Benchmark for large language models and agents on end-to-end machine learning workflows over repository-level code, including ML-LLM-Bench and ML-Agent-Bench tracks.
    Code · Paper · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Human-in-the-loop · Domain: CS · Openness: Open-source

  • MLE-bench - Benchmark for machine learning engineering agents, including task construction, evaluation logic, and reference agent implementations.
    Code · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Open-source

  • MLR-Bench - Benchmark for open-ended machine learning research with 201 tasks, together with agent and judge baselines for evaluation.
    Code · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: CS · Openness: Open-source

  • LAB-Bench - Biology benchmark for capabilities foundational to scientific research in biology.
    Code · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Open-loop · Domain: Biology · Openness: Open-source

  • AgentBench - General benchmark for evaluating LLMs as agents across diverse interactive environments such as operating systems, databases, web tasks, and knowledge graphs.
    Code · Paper · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • ScienceAgentBench - Benchmark for language agents on 102 expert-validated tasks from real scientific workflows across multiple disciplines.
    Code · Paper · Homepage · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Partially Open

  • ScholarQABench - Holistic benchmark for testing scientific literature synthesis, citation accuracy, coverage, relevance, and organization in long-form scholarly answers.
    Code · Paper · GitHub stars
    Level: Benchmark · Stage: Literature · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Open-source

  • RE-Bench - Task suite for evaluating frontier AI R&D capabilities of language-model agents against human experts on realistic research engineering tasks.
    Code · Paper · GitHub stars
    Level: Benchmark · Stage: End-to-end · Loop: Human-in-the-loop · Domain: CS · Openness: Partially Open

  • EXP-Bench - Benchmark for assessing whether AI agents can conduct AI research experiments with rigorous implementation, execution, and analysis loops.
    Paper · Docs
    Level: Benchmark · Stage: Experiment · Loop: Human-in-the-loop · Domain: CS · Openness: Partially Open

  • ResearchBench - Benchmark for scientific discovery via inspiration retrieval, hypothesis composition, and hypothesis ranking across multiple disciplines.
    Paper
    Level: Benchmark · Stage: Ideation · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Paper-only

  • IdeaBench - Benchmarking framework for evaluating large language models on research idea generation quality, novelty, and relevance.
    Paper
    Level: Benchmark · Stage: Ideation · Loop: Human-in-the-loop · Domain: General · Openness: Paper-only

  • LiveIdeaBench - Benchmark for evaluating scientific creativity and divergent thinking in LLM-generated ideas using minimal-context prompts.
    Homepage · Paper
    Level: Benchmark · Stage: Ideation · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Partially Open

Datasets

  • OpenAlex - Open index of works, authors, venues, institutions, and concepts that underpins many research-native retrieval systems.
    Homepage · Docs
    Level: Dataset · Stage: Literature · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Partially Open

  • Semantic Scholar Academic Graph API - Structured paper metadata and graph endpoints for retrieval, paper linking, and citation analysis.
    Docs
    Level: Dataset · Stage: Literature · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Partially Open

Other Awesome Lists

  • awesome-research - Curated collection organized around research workflow tasks, useful as a complementary module-first view.
    Code · GitHub stars
    Level: Meta · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Awesome-AI-Scientists - Curated list centered on AI Scientist systems, complementary to this repository's broader layered-map perspective.
    Code · GitHub stars
    Level: Meta · Stage: End-to-end · Loop: Human-in-the-loop · Domain: General · Openness: Open-source

  • Awesome Papers on Agents for Science - Curated bibliography of agents-for-science papers across domains, task types, and benchmarks.
    Code · GitHub stars
    Level: Meta · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Open-source

  • Awesome AI Scientist Papers - Curated bibliography of AI Scientist and Robot Scientist papers, surveys, benchmarks, timelines, and community resources.
    Code · GitHub stars
    Level: Meta · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Open-source

  • Awesome AutoResearch - Curated list of autoresearch use cases, implementations, and optimization traces for keep-or-revert research loops across domains.
    Code · GitHub stars
    Level: Meta · Stage: End-to-end · Loop: Human-in-the-loop · Domain: Multi-domain · Openness: Open-source

Contributing

Please read CONTRIBUTING.md, docs/inclusion-criteria.md, and docs/tag-system.md before submitting new entries.

Recommended inclusion standards:

  • The project should be clearly relevant to scientific research or AI-for-Research workflows.
  • Prefer official project pages, official GitHub repositories, or papers over secondary summaries.
  • Avoid adding generic AI tools unless they are widely used as research-native infrastructure or workflow components.
  • Prefer representative systems over near-duplicate wrappers.

License

MIT