README.md

June 25, 2026 · View on GitHub

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

Zekun Zhou1, Xiaocheng Feng1†,2, Lei Huang1, Xiachong Feng3, Ziyun Song1, Ruihan Chen1, Liang Zhao1, Weitao Ma1, Yuxuan Gu1, Baoxin Wang4, Dayong Wu4, Guoping Hu4, Ting Liu1,2, Bing Qin1,2
1Harbin Institute of Technology, Harbin, China   2Peng Cheng Laboratory, Shenzhen, China
3The University of Hong Kong, China   4iFLYTEK Research, China

Paper Github License

This repository contains the resources for paper From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems, including related papers, benchmarks, and tools that can accelerate research.

overview taxonomy

For more details, please refer to the paper: From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems.

🎉 Updates

  • This paper is accepted to EMNLP 2025, camera ready version released.
  • 2025/03/03 Paper is available on arxiv.
  • 2025/03/02 We created this reading list repository.

🎁 Resources

Backgrounds

  • Could AI help you to write your next paper? Nature 2022, [paper]
  • Algorithmic ghost in the research shell: Large language models and academic knowledge creation in management research, arXiv.2303.07304, [paper]
  • Scientists' Perspectives on the Potential for Generative AI in their Fields, arXiv.2304.01420, [paper]
  • Friend or foe? Exploring the implications of large language models on the science system, arXiv.2306.09928, [paper]
  • An Interdisciplinary Outlook on Large Language Models for Scientific Research, arXiv.2311.04929 [paper]
  • Artificial intelligence and illusions of understanding in scientific research, Nature 2024, [paper]
  • Monitoring ai-modified content at scale: A case study on the impact of chatgpt on ai conference peer reviews, ICML 2024, [paper]
  • Mapping the increasing use of LLMs in scientific papers, arXiv.2404.01268, [paper]
  • LLMs as Research Tools: A Large Scale Survey of Researchers' Usage and Perceptions, arXiv.2411.05025, [paper]

Surveys

  • What's Missing in Autonomous Research? A Systematization of Systems, Benchmarks, and Verification, ResearchGate.2026, [paper]
  • LLM4SR: A Survey on Large Language Models for Scientific Research, arXiv.2501.04306, [paper]
  • Paper recommender systems: a literature survey, International Journal on Digital Libraries 2016, [paper]
  • Scientific paper recommendation: A survey, IEEE Access 2019, [paper]
  • A review on personalized academic paper recommendation, Comput. Inf. Sci 2019, [paper]
  • Insights into relevant knowledge extraction techniques: a comprehensive review, The Journal of Supercomputing 2019, [paper]
  • Scientific paper recommendation systems: a literature review of recent publications, International journal on digital libraries 2022, [paper]
  • An anatomization of research paper recommender system: Overview, approaches and challenges, Engineering Applications of Artificial Intelligence 2023, [paper]
  • Automatic summarization of scientific articles: A survey, Journal of King Saud University-Computer and Information Sciences 2022, [paper]
  • Artificial intelligence for literature reviews: Opportunities and challenges, Artificial Intelligence Review 2024, [paper]
  • Scientific fact-checking: A survey of resources and approaches, ACL 2023, [paper]
  • Automated justification production for claim veracity in fact checking: A survey on architectures and approaches, ACL 2024, [paper]
  • Claim verification in the age of large language models: A survey, arXiv.2408.14317, [paper]
  • A survey on deep learning for theorem proving, arXiv.2404.09939, [paper]
  • Related work and citation text generation: A survey, EMNLP 2024, [paper]
  • Automated scholarly paper review: Concepts, technologies, and challenges, Information fusion 2023, [paper]
  • Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review, arXiv.2411.09763, [paper]
  • What Can Natural Language Processing Do for Peer Review?, arXiv.2405.06563, [paper]
  • Artificial intelligence to support publishing and peer review: A summary and review, Learned Publishing 2024, [paper]
  • Large language models for scientific synthesis, inference and explanation, arXiv.2310.07984, [paper]
  • A comprehensive survey of scientific large language models and their applications in scientific discovery, EMNLP 2024, [paper]
  • Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions, arXiv.2503.08979, [paper]
  • A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models, arXiv.2504.05496, [paper]
  • Scientific Hypothesis Generation and Validation: Methods, Datasets, and Future Directions, arXiv.2505.04651, [paper]

Integration of Diverse Research Tasks

  • Scaling Laws of Scientific Discovery with AI and Robot Scientists, arXiv.2503.22444,[paper]

Hypothesis Formulation

HypothesisFormulation

Knowledge Synthesis

Research Paper Recommendation
  • From Who You Know to What You Read: Augmenting Scientific Recommendations with Implicit Social Networks, CHI 2022, [paper]
  • Comlittee: Literature discovery with personal elected author committees, CHI 2023, [paper]
  • ArZiGo: A recommendation system for scientific articles, Information Systems 2024, [paper]
  • An academic recommender system on large citation data based on clustering, graph modeling and deep learning, Knowledge and Information Systems 2024, [paper]
  • Paperweaver: Enriching topical paper alerts by contextualizing recommended papers with user-collected papers, CHI 2024, [paper]
  • Scholar Inbox: Personalized Paper Recommendations for Scientists, arXiv.2504.08385, [paper]
Systematic Literature Review
  • Multi-document scientific summarization from a knowledge graph-centric view, COLING 2022, [paper]
  • Hierarchical catalogue generation for literature review: a benchmark, EMNLP 2023, [paper]
  • Bio-sieve: exploring instruction tuning large language models for systematic review automation, arXiv.2308.06610, [paper]
  • Assisting in writing wikipedia-like articles from scratch with large language models, NAACL 2024, [paper]
  • Autosurvey: Large language models can automatically write surveys, NeurIPS 2024, [paper]
  • Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature, EMNLP 2024, [paper]
  • Chime: Llm-assisted hierarchical organization of scientific studies for literature review support, ACL 2024, [paper]
  • Instruct Large Language Models to Generate Scientific Literature Survey Step by Step, NLPCC 2024, [paper]
  • Hireview: Hierarchical taxonomy-driven automatic literature review generation, arXiv.2410.03761, [paper]
  • Language agents achieve superhuman synthesis of scientific knowledge, arXiv.2409.13740, [paper]
  • Litllm: A toolkit for scientific literature review, arXiv.2402.01788, [paper]
  • LLMs for Literature Review: Are we there yet?, arXiv.2412.15249, [paper]
  • Automating research synthesis with domain-specific large language model fine-tuning, arXiv.2404.08680, [paper]
  • OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs, arXiv.2411.14199, [paper]
  • Agent Laboratory: Using LLM Agents as Research Assistants, arXiv.2501.04227, [paper]
  • SurveyX: Academic Survey Automation via Large Language Models, arXiv.2502.14776,[paper]
  • LimTopic: LLM-based Topic Modeling and Text Summarization for Analyzing Scientific Articles limitations, JCDL 2024,[paper]
  • Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science, arXiv.2503.16515, [paper]
  • Generalization Bias in Large Language Model Summarization of Scientific Research, arXiv.2504.00025,[paper]
  • Identifying Aspects in Peer Reviews, arXiv.2504.06910, [paper]
  • Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol, arXiv.2504.10284, [paper]
  • Ai2 Scholar QA: Organized Literature Synthesis with Attribution, arXiv.2504.10861, [paper]
  • Deep literature reviews: an application of fine-tuned language models to migration research, arXiv.2504.13685, [paper]
  • Science Hierarchography: Hierarchical Organization of Science Literature, arXiv.2504.13834, [paper]
  • Towards Artificial Intelligence Research Assistant for Expert-Involved Learning, arXiv.2505.04638, [paper]
Other Works
  • What's In Your Field? Mapping Scientific Research with Knowledge Graphs and Large Language Models, arXiv.2503.09894, [paper]

Hypothesis Generation

  • Literature based discovery: models, methods, and trends, Journal of biomedical informatics 2017, [paper]
  • Predicting the Future of AI with AI: High-quality link prediction in an exponentially growing knowledge network, arXiv.2210.00881, [paper]
  • Ideas are dimes a dozen: Large language models for idea generation in innovation, The Wharton School Research Paper Forthcoming 2023, [paper]
  • Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers, arXiv.2409.04109, [paper]
  • Can Large Language Models Unlock Novel Scientific Research Ideas?, arXiv.2409.06185, [paper]
  • MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses, arXiv.2410.07076, [paper]
  • IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback, arXiv.2410.04025, [paper]
  • Data-driven discovery with large generative models, ICML 2024, [paper]
  • DiscoveryWorld: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents, NeurIPS 2024,[paper]
  • Literature meets data: A synergistic approach to hypothesis generation, arXiv.2410.17309, [paper]
  • SciMON: Scientific Inspiration Machines Optimized for Novelty, ACL 2024, [paper]
  • Chain of ideas: Revolutionizing research via novel idea development with llm agents, arXiv.2410.13185, [paper]
  • SciPIP: An LLM-based Scientific Paper Idea Proposer, arXiv.2410.23166, [paper]
  • Accelerating scientific discovery with generative knowledge extraction, graph-based representation, and multimodal intelligent graph reasoning, Machine Learning: Science and Technology 2024, [paper]
  • Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning, arXiv.2409.05556, [paper]
  • Generation and human-expert evaluation of interesting research ideas using knowledge graphs and large language models, arXiv.2405.17044, [paper]
  • Hypothesis generation with large language models, arXiv.2404.04326, [paper]
  • Large language models for automated open-domain scientific hypotheses discovery, ACL 2024, [paper]
  • Researchagent: Iterative research idea generation over scientific literature with large language models, arXiv.2404.07738, [paper]
  • LLM and Simulation as Bilevel Optimizers: A New Paradigm to Advance Physical Scientific Discovery, ICML 2024, [paper]
  • The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, arXiv.2408.06292, [paper]
  • Acceleron: A tool to accelerate research ideation, arXiv.2403.04382, [paper]
  • Sciagents: Automating scientific discovery through multi-agent intelligent graph reasoning, arXiv.2409.05556, [paper]
  • Two heads are better than one: A multi-agent system has the potential to improve scientific idea generation, arXiv.2410.09403, [paper]
  • Nova: An iterative planning and search approach to enhance novelty and diversity of llm generated ideas, arXiv.2410.14255, [paper]
  • Learning to Generate Research Idea with Dynamic Control, arXiv.2412.14626, [paper]
  • ResearchTown: Simulator of Human Research Community, arXiv.2412.17767,[paper]
  • Dolphin: Closed-loop Open-ended Auto-research through Thinking, Practice, and Feedback, arXiv.2501.03916, [paper]
  • Towards an AI co-scientist, arXiv.2502.18864,[paper]
  • Graph of AI Ideas: Leveraging Knowledge Graphs and LLMs for AI Research Idea Generation, arXiv.2503.08549, [paper]
  • AgentRxiv: Towards Collaborative Autonomous Research, arXiv.2503.18102,[paper]
  • Exploring Topic Trends in COVID-19 Research Literature using Non-Negative Matrix Factorization, arXiv.2503.18182,[paper]
  • Iterative Hypothesis Generation for Scientific Discovery with Monte Carlo Nash Equilibrium Self-Refining Trees, arXiv.2503.19309, [paper]
  • SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings, arXiv.2503.19257,[paper]
  • CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation, arXiv.2503.22708,[paper]
  • Advancing AI-Scientist Understanding: Making LLM Think Like a Physicist with Interpretable Reasoning, arXiv.2504.01911, [paper]
  • The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search, arXiv.2504.08066, [paper]
  • Reduction of Supervision for Biomedical Knowledge Discovery, arXiv.2504.09582, [paper]
  • Sparks of Science: Hypothesis Generation Using Structured Paper Data, arXiv.2504.12976, [paper]
  • IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery, arXiv.2504.16728, [paper]
  • FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights, arXiv.2505.04649, [paper]

Hypothesis Validation

HypothesisValidation

Scientific Claim Verification

  • Fact or fiction: Verifying scientific claims, EMNLP 2020, [paper]
  • Generating fact checking explanations, ACL 2020, [paper]
  • Missing counter-evidence renders NLP fact-checking unrealistic for misinformation, EMNLP 2022, [paper]
  • MultiVerS: Improving scientific claim verification with weak supervision and full-document context, NAACL 2022, [paper]
  • Proofver: Natural logic theorem proving for fact verification, TACL 2022, [paper]
  • Fact-checking complex claims with program-guided reasoning, ACL 2023, [paper]
  • FactKG: Fact verification via reasoning on knowledge graphs, ACL 2023, [paper]
  • Prompt to be Consistent is Better than Self-Consistent? Few-Shot and Zero-Shot Fact Verification with Pre-trained Language Models, ACL 2023, [paper]
  • Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method, IJCNLP 2023, [paper]
  • Investigating zero-and few-shot generalization in fact verification, IJCNLP 2023, [paper]
  • The state of human-centered NLP technology for fact-checking, Information processing & management 2023, [paper]
  • Characterizing and Verifying Scientific Claims: Qualitative Causal Structure is All You Need, EMNLP 2023, [paper]
  • aedFaCT: Scientific Fact-Checking Made Easier via Semi-Automatic Discovery of Relevant Expert Opinions, arXiv.2305.07796, [paper]
  • What Makes Medical Claims (Un)Verifiable? Analyzing Entity and Relation Properties for Fact Verification, ACL 2024, [paper]
  • Unsupervised Pretraining for Fact Verification by Language Model Distillation, ICLR 2024, [paper]
  • Comparing knowledge sources for open-domain scientific claim verification, EACL 2024, [paper]
  • Improving health question answering with reliable and time-aware evidence retrieval, NAACL 2024, [paper]
  • MAGIC: Multi-Argument Generation with Self-Refinement for Domain Generalization in Automatic Fact-Checking, COLING 2024, [paper]
  • Understanding Fine-grained Distortions in Reports of Scientific Findings, ACL 2024, [paper]
  • ClaimVer: Explainable claim-level verification and evidence attribution of text through knowledge graphs, EMNLP 2024, [paper]
  • Grounding fallacies misrepresenting scientific publications in evidence, arXiv.2408.12812, [paper]
  • Can Large Language Models Detect Misinformation in Scientific News Reporting?, arXiv.2402.14268, [paper]
  • Enhancing natural language inference performance with knowledge graph for COVID-19 automated fact-checking in Indonesian language, arXiv.2409.00061, [paper]
  • Augmenting the Veracity and Explanations of Complex Fact Checking via Iterative Self-Revision with LLMs, arXiv.2410.15135, [paper]
  • Optimizing Decomposition for Optimal Claim Verification, arXiv.2503.15354, [paper]
  • SciClaims: An End-to-End Generative System for Biomedical Claim Analysis, arXiv.2503.18526, [paper]
  • Can LLMs Automate Fact-Checking Article Writing?, arXiv.2503.17684 , [paper]
  • CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?, arXiv.2503.21717, [paper]
  • Measuring and Analyzing Subjective Uncertainty in Scientific Communications, arXiv.2503.21114, [paper]

Theorem Proving

  • Generative language modeling for automated theorem proving, arXiv.2009.03393, [paper]
  • HyperTree Proof Search for Neural Theorem Proving, NeurIPS 2022, [paper]
  • Thor: Wielding hammers to integrate language models and automated theorem provers, NeurIPS 2022, [paper]
  • Draft, sketch, and prove: Guiding formal theorem provers with informal proofs, ICLR 2023, [paper]
  • Dt-solver: Automated theorem proving with dynamic-tree sampling guided by proof-level value function, ACL 2023, [paper]
  • Baldur: Whole-proof generation and repair with large language models, ESEC/FSE 2023, [paper]
  • An in-context learning agent for formal theorem-proving, arXiv.2310.04353, [paper]
  • Decomposing the enigma: Subgoal-based demonstration learning for formal theorem proving, arXiv.2305.16366, [paper]
  • Proving theorems recursively, NeurIPS 2024, [paper]
  • Lean-star: Learning to interleave thinking and proving, arXiv.2407.10040, [paper]
  • Lego-prover: Neural theorem proving with growing libraries, ICLR 2024, [paper]
  • Mustard: Mastering uniform synthesis of theorem and proof data, ICLR 2024, [paper]
  • Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data, arXiv.2405.14333, [paper]
  • Towards large language models as copilots for theorem proving in lean, arXiv.2404.12534, [paper]
  • Automating Mathematical Proof Generation Using Large Language Model Agents and Knowledge Graphs, arXiv.2503.11657,[paper]
  • LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models, arXiv.2504.14089, [paper]
  • DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition, arXiv.2504.21801, [paper]
  • Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving, arXiv.2505.04528, [paper]

Experiment Validation

  • Autonomous chemical research with large language models, Nature 2023 [paper]
  • An autonomous laboratory for the accelerated synthesis of novel materials, Nature 2023, [paper]
  • Automl-gpt: Automatic machine learning with gpt, arXiv.2305.02499, [paper]
  • MechAgents: Large language model multi-agent collaborations can solve mechanics problems, generate new data, and integrate knowledge, arXiv.2311.08166, [paper]
  • Position: LLMs can't plan, but can help planning in LLM-modulo frameworks, ICML 2024, [paper]
  • An automatic end-to-end chemical synthesis development platform powered by large language models, Nature communications 2024, [paper]
  • Mlcopilot: Unleashing the power of large language models in solving machine learning tasks, EACL 2024, [paper]
  • Meta-Designing Quantum Experiments with Language Models, arXiv.2406.02470, [paper]
  • Augmenting large language models with chemistry tools, Nature Machine Intelligence 2024, [paper]
  • Crispr-gpt: An llm agent for automated design of gene-editing experiments, arXiv.2404.18021, [paper]
  • Chain of ideas: Revolutionizing research via novel idea development with llm agents, arXiv.2410.13185, [paper]
  • The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, arXiv.2408.06292, [paper]
  • Large language model agent for hyper-parameter optimization, arXiv.2402.01881, [paper]
  • MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration, arXiv.2411.08063, [paper]
  • Mlr-copilot: Autonomous machine learning research based on large language models agents, arXiv.2408.14033, [paper]
  • Researchagent: Iterative research idea generation over scientific literature with large language models, arXiv.2404.07738, [paper]
  • Training socially aligned language models on simulated social interactions, ICLR 2024, [paper]
  • Automated social science: Language models as scientist and subjects, arXiv.2404.11794, [paper]
  • AAAR-1.0: Assessing AI's Potential to Assist Research, arXiv.2410.22394, [paper]
  • Dolphin: Closed-loop Open-ended Auto-research through Thinking, Practice, and Feedback, arXiv.2501.03916, [paper]
  • Agent Laboratory: Using LLM Agents as Research Assistants, arXiv.2501.04227, [paper]
  • Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning, arXiv.2504.17192, [paper]
  • ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies, AAAI 2025, [paper]
  • Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents, arxiv.2502.16069, [paper]

Manuscript Publication

ManuscriptPublication

Manuscript Writing

  • Disencite: Graph-based disentangled representation learning for context-specific citation generation, AAAI 2022, [paper]
  • Intent-controllable citation text generation, Mathematics 2022, [paper]
  • Understanding Iterative Revision from Human-Written Text, ACL 2022, [paper]
  • Improving Iterative Text Revision by Learning Where to Edit from Other Revision Tasks, EMNLP 2022, [paper]
  • Controllable citation sentence generation with language models, arXiv.2211.07066, [paper]
  • Read, revise, repeat: A system demonstration for human-in-the-loop iterative text revision, arXiv.2204.03685, [paper]
  • Enabling large language models to generate text with citations, EMNLP 2023, [paper]
  • Towards a unified framework for reference retrieval and related work generation, EMNLP 2023, [paper]
  • SciLit: A platform for joint scientific literature discovery, summarization and citation generation, ACL 2023, [paper]
  • Exploring the boundaries of reality: investigating the phenomenon of artificial intelligence hallucination in scientific writing through ChatGPT references, Cureus 2023, [paper]
  • Can artificial intelligence help for scientific writing?, Critical care 2023, [paper]
  • Decoding the End-to-end Writing Trajectory in Scholarly Manuscripts, In2Writing @CHI 2023, [paper]
  • Text revision in scientific writing assistance: An overview, arXiv.2303.16726, [paper]
  • Text revision by on-the-fly representation optimization, arXiv.2303.16726, [paper]
  • A publishing infrastructure for Artificial Intelligence (AI)-assisted academic authoring, JAMIA 2024, [paper]
  • Automated focused feedback generation for scientific writing assistance, ACL 2024, [paper]
  • Aries: A corpus of scientific paper edits made in response to peer reviews, ACL 2024, [paper]
  • Reinforced Subject-Aware Graph Neural Network for Related Work Generation, KSEM 2024, [paper]
  • Toward Structured Related Work Generation with Novelty Statements, SDP 2024, [paper]
  • Instruct Large Language Models to Generate Scientific Literature Survey Step by Step, NLPCC 2024, [paper]
  • Techniques for supercharging academic writing with generative AI, Nature Biomedical Engineering 2024, [paper]
  • Reference hallucination score for medical artificial intelligence chatbots: development and usability study, JMIR Medical Informatics 2024, [paper]
  • Shallow synthesis of knowledge in gpt-generated texts: A case study in automatic related work composition, arXiv.2402.12255, [paper]
  • The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, arXiv.2408.06292, [paper]
  • Cocoa: Co-Planning and Co-Execution with AI Agents, arXiv.2412.10999, [paper]
  • ResearchTown: Simulator of Human Research Community, arXiv.2412.17767,[paper]
  • Autonomous LLM-Driven Research—from Data to Human-Verifiable Research Papers, arXiv.2404.17605, [paper]
  • Step-Back Profiling: Distilling User History for Personalized Scientific Writing, arXiv.2406.14275, [paper]
  • Agent Laboratory: Using LLM Agents as Research Assistants, arXiv.2501.04227, [paper]
  • CorpusStudio: Surfacing Emergent Patterns in a Corpus of Prior Work while Writing, ACM CHI 2025,[paper]
  • FutureGen: LLM-RAG Approach to Generate the Future Work of Scientific Article, arXiv.2503.16561, [paper]
  • ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations, arXiv.2504.00824,[paper]
  • Divergent LLM Adoption and Heterogeneous Convergence Paths in Research Writing, arXiv.2504.13629, [paper]
  • CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction, arXiv.2504.15629, [paper]

Peer Review

  • Exploiting Labeled and Unlabeled Data via Transformer Fine-tuning for Peer-Review Score Prediction, EMNLP 2022, [paper]
  • Can We Automate Scientific Reviewing?, JAIR 2022, [paper]
  • The Quality Assist: A Technology-Assisted Peer Review Based on Citation Functions to Predict the Paper Quality, IEEE Access 2022, [paper]
  • KID-Review: Knowledge-Guided Scientific Review Generation with Oracle Pre-training, AAAI 2022, [paper]
  • Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation, EMNLP 2023, [paper]
  • When Reviewers Lock Horn: Finding Disagreement in Scientific Peer Reviews, EMNLP 2023, [paper]
  • Gpt4 is slightly helpful for peer-review assistance: A pilot study, arXiv.2307.05492, [paper]
  • ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing, arXiv.2306.00622, [paper]
  • Unveiling the Sentinels: Assessing AI Performance in Cybersecurity Peer Review, arXiv.2309.05457, [paper]
  • Can large language models provide useful feedback on research papers, arXiv.2310.01783, [paper]
  • Scientific Opinion Summarization: Meta-review Generation with Checklist-guided Iterative Introspection, arXiv.2305.14647, [paper]
  • MetaWriter: Exploring the Potential and Perils of AI Writing Support in Scientific Peer Review, PACMHCI 2024, [paper]
  • GLIMPSE: Pragmatically Informative Multi-Document Summarization for Scholarly Reviews, ACL 2024, [paper]
  • Automated Focused Feedback Generation for Scientific Writing Assistance, ACL 2024, [paper]
  • AgentReview: Exploring Peer Review Dynamics with LLM Agents, EMNLP 2024, [paper]
  • A sentiment consolidation framework for meta-review generation, ACL 2024, [paper]
  • Human-in-the-loop AI reviewing: Feasibility, opportunities, and risks, JAIS 2024, [paper]
  • RelevAI-Reviewer: A Benchmark on AI Reviewers for Survey Paper Relevance, arXiv.2406.10294, [paper]
  • The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, arXiv.2408.06292, [paper]
  • MARG: Multi-Agent Review Generation for Scientific Papers, arXiv.2401.04259, [paper]
  • Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions, arXiv.2406.05688, [paper]
  • CycleResearcher: Improving Automated Research via Automated Review, arXiv.2411.00816, [paper]
  • OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews, arXiv.2412.11948, [paper]
  • Prompting LLMs to Compose Meta-Review Drafts from Peer-Review Narratives of Scholarly Manuscripts, arXiv.2402.15589, [paper]
  • PeerArg: Argumentative Peer Review with LLMs, arXiv.2409.16813, [paper]
  • ResearchTown: Simulator of Human Research Community, arXiv.2412.17767,[paper]
  • ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews, arXiv.2503.08506, [paper]
  • DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process, arXiv.2503.08569, [paper]
  • Enabling Inclusive Systematic Reviews: Incorporating Preprint Articles with Large Language Model-Driven Evaluations, arXiv.2503.13857,[paper]
  • Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025, arXiv.2504.09737, [paper]
  • LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews, arXiv.2504.11042, [paper]

💯 Benchmarks

benchmark

Research Paper Recommendation

  • Academic Paper Recommendation Method Combining Heterogeneous Network and Temporal Attributes, CCSCW 2020, [paper]
  • Paper recommend based on LDA and PageRank, ICAIS 2020, [paper]
  • A hybrid approach for paper recommendation, IEICE TRANS 2021, [paper]

Systematic Literature Review

  • Bringing structure into summaries: a faceted summarization dataset for long scientific documents, ACL/IJCNLP 2021, [paper]
  • Generating a structured summary of numerous academic papers: Dataset and method, IJCAI 2022, [paper]
  • SciReviewGen: a large-scale dataset for automatic literature review generation, ACL 2023, [paper]
  • Hierarchical catalogue generation for literature review: a benchmark, EMNLP 2023, [paper]
  • Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature, EMNLP 2024, [paper]
  • CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support, ACL 2024, [paper]
  • Overview of the NLPCC2024 Shared Task 6: Scientific Literature Survey Generation, NLPCC 2024, [paper]
  • OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs, arXiv.2411.14199, [paper]

Hypothesis Generation

  • SciMON: Scientific Inspiration Machines Optimized for Novelty, ACL 2024, [paper]
  • Large Language Models for Automated Open-domain Scientific Hypotheses Discovery, ACL 2024, [paper]
  • MASSW: A new dataset and benchmark tasks for ai-assisted scientific workflows, arXiv.2406.06357, [paper]
  • DiscoveryBench: Towards Data-Driven Discovery with Large Language Models, arXiv.2407.01725, [paper]
  • Can Large Language Models Unlock Novel Scientific Research Ideas?, arXiv.2409.06185, [paper]
  • IdeaBench: Benchmarking Large Language Models for Research Idea Generation, arXiv.2411.02429, [paper]
  • LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context, arXiv.2412.17596, [paper]
  • MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research, arXiv.2503.13399,[paper]
  • ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition, arXiv.2503.21248, [paper]
  • DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments, arXiv.2504.03160, [paper]
  • LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models, arXiv.2504.10415, [paper]
  • ArxivBench: Can LLMs Assist Researchers in Conducting Research?, arXiv.2504.10496, [paper]
  • HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation, arXiv.2504.11524, [paper]
  • AI Idea Bench 2025: AI Research Idea Generation Benchmark, arXiv.2504.14191, [paper]

Scientific Claim Verification

  • FEVER: a Large-scale Dataset for Fact Extraction and VERification, NAACL-HLT 2018, [paper]
  • Fact or fiction: Verifying scientific claims, EMNLP 2020, [paper]
  • Evidence-based Fact-Checking of Health-related Claims, EMNLP 2021, [paper]
  • SciFact-Open: Towards open-domain scientific claim verification, EMNLP 2022, [paper]
  • SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables, EMNLP 2023, [paper]
  • Check-COVID: Fact-Checking COVID-19 News Claims with Scientific Evidence, ACL 2023, [paper]
  • FactKG: Fact Verification via Reasoning on Knowledge Graphs, ACL 2023, [paper]
  • Missci: Reconstructing Fallacies in Misrepresented Science, ACL 2024, [paper]
  • MAGIC: Multi-Argument Generation with Self-Refinement for Domain Generalization in Automatic Fact-Checking, LREC/COLING 2024, [paper]
  • QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims, SIGIR 2024, [paper]
  • HealthFC: Verifying Health Claims with Evidence-Based Medical Fact-Checking, LREC/COLING 2024, [paper]
  • What Makes Medical Claims (Un)Verifiable? Analyzing Entity and Relation Properties for Fact Verification, EACL 2024, [paper]
  • Sciriff: A resource to enhance language model instruction-following over scientific literature, arXiv.2406.07835, [paper]
    • EvidenceBench: A Benchmark for Extracting Evidence from Biomedical Papers, arXiv.2504.18736,[paper]

Theorem Proving

  • Learning to Prove Theorems via Interacting with Proof Assistants, ICML 2019, [paper]
  • miniF2F: a cross-system benchmark for formal Olympiad-level mathematics, ICLR 2022, [paper]
  • LeanDojo: Theorem Proving with Retrieval-Augmented Language Models, NeurIPS 2023, [paper]
  • TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models, EMNLP 2023, [paper]
  • FIMO: A Challenge Formal Dataset for Automated Theorem Proving, arXiv.2309.04295, [paper]
  • LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover, arXiv.2407.17227, [paper]

Experiment Validation

  • MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation, ICML 2024, [paper]
  • TaskBench: Benchmarking Large Language Models for Task Automation, NeurIPS 2024, [paper]
  • Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?, NeurIPS 2024, [paper]
  • SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories, EMNLP 2024, [paper]
  • LAB-Bench: Measuring Capabilities of Language Models for Biology Research, arXiv.2407.10362, [paper]
  • CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark, arXiv.2409.11363, [paper]
  • AAAR-1.0: Assessing AI's Potential to Assist Research, arXiv.2410.22394, [paper]
  • MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research, arXiv.2503.13399,[paper]
  • SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers, arXiv.2504.00255, [paper]

Manuscript Writing

  • The ACL anthology network corpus, Language Resources and Evaluation 2013, [paper]
  • ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks, AAAI 2019, [paper]
  • DisenCite: Graph-Based Disentangled Representation Learning for Context-Specific Citation Generation, AAAI 2022, [paper]
  • arXivEdits: Understanding the Human Revision Process in Scientific Writing, EMNLP 2022, [paper]
  • SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning, AAAI 2023, [paper]
  • CiteBench: A Benchmark for Scientific Citation Text Generation, EMNLP 2023, [paper]
  • Enabling Large Language Models to Generate Text with Citations, EMNLP 2023, [paper]
  • CASIMIR: A Corpus of Scientific Articles Enhanced with Multiple Author-Integrated Revisions, LREC/COLING 2024, [paper]
  • ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction, WRAICOGS 2025, [paper]
  • SCHOLAWRITE: A Dataset of End-to-End Scholarly Writing Process, arXiv.2502.02904, [paper]

Peer Review

  • A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications, NAACL-HLT 2018, [paper]
  • Argument Mining for Understanding Peer Reviews, NAACL-HLT 2019, [paper]
  • ReAct: A Review Comment Dataset for Actionability (and more), WISE 2021, [paper]
  • MReD: A Meta-Review Dataset for Structure-Controllable Text Generation, ACL 2022, [paper]
  • Can We Automate Scientific Reviewing?, JAIR 2022, [paper]
  • NLPeer: A Unified Resource for the Computational Study of Peer Review, ACL 2023, [paper]
  • Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation, EMNLP 2023, [paper]
  • MOPRD: A multidisciplinary open peer review dataset, NCA 2023, [paper]
  • Scientific opinion summarization: Paper meta-review generation dataset, methods, and evaluation, IJCAI 2024, [paper]
  • ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews, ACL 2024, [paper]
  • Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions, arXiv.2406.05688, [paper]

Others

  • CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning, ICLR 2025, [paper]
  • MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers, arXiv.2503.16856,[paper]
  • PaperBench: Evaluating AI's Ability to Replicate AI Research, arXiv.2504.01848,[paper]

🚀 Tools

Tool NameResearch Paper RecommendationSystematic Literature ReviewHypothesis GenerationScientific Claim VerificationTheorem ProvingExperiment VerificationManuscript WritingPeer ReviewReading Assistance
Connected Paper
Inciteful
Litmaps
Pasa
Research Rabbit
Semantic Scholar
GenGO
Jenni AI
Elicit
Undermind
OpenScholar
ResearchBuddies
Hyperwrite
Concensus
Iris.ai
MirrorThink
SciSpace
AskYourPDF
Iflytek
FutureHouse
Enago Read
Aminer
OpenResearcher
ResearchFlow
You.com
GPT Researcher
PICO Portal
SurveyX
Scinence42:Dora
STORM
ChatDOC
Scite
Silatus
Agent Laboratory
Sider
Quillbot
Scholar AI
AI-Researcher
AI Scientist
Isabelle
LeanCopilot
Llmstep
Proverbot9001
chatgpt_academic
gpt_academic
HeadlineAnalyzer
Langsmith Editor
Textero.ai
Wordvice.AI
Writesonic
Writefull
Covidence
Penelope.ai
Byte-science
Cool Papers
Explainpaper
Uni-finder

📝 Citation

If you find our work helpful, you can cite this paper as:

@article{zhou2025hypothesis,
  title={From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems},
  author={Zhou, Zekun and Feng, Xiaocheng and Huang, Lei and Feng, Xiachong and Song, Ziyun and Chen, Ruihan and Zhao, Liang and Ma, Weitao and Gu, Yuxuan and Wang, Baoxin and others},
  journal={arXiv preprint arXiv:2503.01424},
  url={https://arxiv.org/abs/2503.01424},
  year={2025}
}