Awesome LLM-Based Human-Agent Collaboration and Interaction Systems

July 23, 2026 ยท View on GitHub

๐ŸŽ‰ Our survey has been accepted to ACL 2026!

ACL 2026 Awesome arXiv Maintenance Contribution Welcome Last Commit GitHub stars

Oryx Video-ChatGPT

image

Welcome to Awesome-Human-Agent-Collaboration-Interaction-Systems! ๐Ÿš€ This is the repo for our Survey on LLM-Based Human-Agent Collaboration and Interaction Systems, accepted to ACL 2026.

๐ŸŒŸ Introduction

Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-based agents still face significant challenges, including (1) limited reliability due to hallucinations, (2) difficulty in handling complex tasks, and (3) substantial safety and ethical risks, all of which limit their feasibility and trustworthiness in real-world applications.

LLM-based human-agent collaboration systems are interactive frameworks where humans actively provide (1) additional information, (2) feedback, or (3) control during interaction with LLM-powered agents to enhance system performance, reliability, and safety. These human-agent collaboration systems enable humans and LLM-based agents to collaborate effectively by leveraging their complementary strengths. For a detailed introduction, please refer to our survey paper: LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey.

Our goal with this project is to build an exhaustive collection of awesome resources relevant to LLM-Based Human-Agent/AI Collaboration and Interaction Systems, encompassing papers, repositories, and more to foster further research and innovation in this rapidly evolving interdisciplinary field of human-ai collaboration. ๐Ÿค— Contributions are welcome! ๐Ÿค— If you have recommended papers, resources or suggestions, please submit pull requests, open issues or contact us. We will keep updating our repo & survey paper.

๐Ÿ“„ Contents

๐Ÿ“„ Latest Research Papers

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

๐Ÿค— Contributions are welcome! If you have recommended papers and resources, please submit pull requests or open issues.

๐Ÿ“š Applications, Datasets & Benchmarks

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

๐Ÿ’ป Web Navigation & Computer Use

๐Ÿ‘จ๐Ÿปโ€๐Ÿ’ป Software Engineering, Coding

๐Ÿค– Embodied AI, Robotics

๐Ÿ’ฌ Conversation System

๐Ÿ“Š Data Science, Scientific Discovery

๐ŸŽฎ Gaming

๐Ÿ’ฐ Finance

๐Ÿฅ Healthcare, Medicine

๐Ÿงฉ General-Purpose Assistants, Cross-Domain

๐Ÿ›๏ธ Retail, Telecom

๐Ÿ›ฉ๏ธ Travel

โœ๏ธ Writing

๐Ÿ” Taxonomy

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

For a detailed introduction of the taxonomy, please refer to Section 3 in our survey paper: LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey.

image


๐Ÿค Human Feedback

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

Human Feedback can occur during different phases in various types and granularities. In the following table, we summarize different dimensions of Human Feedback in LLM-based human-agent systems, including feedback type, granularity, and phase. For each dimension, a summary, key characteristics, and example works are provided for comparison. More details are in Section 3.2 of our survey paper.

image

TitleDate & CodeFeedback TypeFeedback SubtypeFeedback GranularityFeedback Phase
Just A Rather Very Intelligent Spoken Agent (JarvisBench)2026/07Guidance, CorrectiveRefinement, CritiqueSegmentDuring Task
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability2026/07 GitHub starsGuidance, EvaluativeCritique, Binary AssessmentSegmentDuring Task
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation2026/07Guidance, Evaluative, Corrective, ImplicitCritique, Refinement, Binary Assessment, Human ControlHolistic, SegmentInitial Setup, During Task, Post Task
Uncertainty Decomposition for Clarification Seeking in LLM Agents2026/06GuidanceDemonstrationSegmentDuring Task
Learning User Simulators with Turing Rewards2026/06ImplicitUser ActionSegmentDuring Task
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE)2026/06 GitHub starsCorrective, GuidanceRefinement, CritiqueSegmentDuring Task, Post Task
Re-Centering Humans in LLM Personalization2026/06Evaluative, ImplicitPreference Ranking, Scalar Rating, User ActionHolistic, SegmentPost Task
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement2026/06Implicit, GuidanceUser Action, CritiqueSegmentDuring Task
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents2026/06Evaluative, Corrective, GuidanceBinary Assessment, Refinement, CritiqueHolistic, SegmentInitial Setup, During Task, Post Task
Uncertainty-Aware Clarification in LLM Agents with Information Gain2026/06GuidanceDemonstrationSegmentDuring Task
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making2026/05Evaluative, ImplicitBinary Assessment, User ActionSegment, HolisticDuring Task
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions2026/05 GitHub starsImplicit, GuidanceUser Action, DemonstrationSegmentInitial Setup, During Task
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL)2026/05Guidance, CorrectiveCritique, RefinementHolistic, SegmentPost Task
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement2026/05Corrective, GuidanceRefinement, DemonstrationSegmentDuring Task
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators2026/05Guidance, EvaluativeCritique, Binary AssessmentSegmentInitial Setup, During Task
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators2026/05Implicit, GuidanceUser Action, CritiqueSegmentDuring Task
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows2026/04Evaluative, Corrective, ImplicitBinary Assessment, Refinement, Human ControlSegmentInitial Setup, During Task
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks2026/04Guidance, Evaluative, ImplicitDemonstration, Scalar Rating, User ActionHolistic, SegmentDuring Task, Post Task
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench)2026/04 GitHub starsCorrective, GuidanceRefinement, CritiqueSegmentDuring Task
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents2026/03GuidanceDemonstration, CritiqueSegmentDuring Task
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS)2026/03 GitHub starsEvaluative, ImplicitScalar Rating, User ActionSegmentDuring Task, Post Task
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks2026/03 GitHub starsImplicit, GuidanceUser Action, DemonstrationSegmentInitial Setup, During Task
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science2026/03Guidance, Corrective, EvaluativeDemonstration, Refinement, Scalar RatingHolistic, SegmentInitial Setup, During Task, Post Task
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration2026/03 GitHub starsCorrective, GuidanceRefinement, DemonstrationSegmentDuring Task
InfoPO: Information-Driven Policy Optimization for User-Centric Agents2026/02 GitHub starsGuidance, ImplicitDemonstration, User ActionSegmentDuring Task
Modeling Distinct Human Interaction in Web Agents (CowCorpus)2026/02Implicit, CorrectiveUser Action, Human Control, RefinementSegmentDuring Task
Overseeing Agents Without Constant Oversight: Challenges and Opportunities2026/02Evaluative, CorrectiveBinary Assessment, CritiqueSegment, HolisticDuring Task, Post Task
Learning Personalized Agents from Human Feedback2026/02 GitHub starsGuidance, Corrective, EvaluativeCritique, Refinement, Preference RankingSegmentInitial Setup, During Task, Post Task
Value of Information: A Framework for Human-Agent Communication2026/01GuidanceDemonstrationSegmentDuring Task
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS)2026/01Evaluative, GuidancePreference Ranking, CritiqueSegmentDuring Task
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering2025/12Guidance, EvaluativeCritique, Preference RankingSegmentInitial Setup, During Task
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding2025/11Guidance, CorrectiveRefinements, CritiqueSegment, HolisticDuring Task
Training Proactive and Personalized LLM Agents2025/11 GitHub starsGuidance, EvaluativeScalar Rating, RefinementsSegment, HolisticDuring Task
Training LLM Agents to Empower Humans2025/10 GitHub starsImplicit, CorrectiveUser Action, RefinementsSegmentDuring Task
How can we assess human-agent interactions? Case studies in software agent design2025/10EvaluativeScalar RatingSegmentDuring Task
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback2025/10 GitHub starsGuidance, CorrectiveRefinements, Critique, DemonstrationSegmentDuring Task
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning2025/09 GitHub starsEvaluative, Guidance, ImplicitScalar Rating, Refinements, CritiqueSegmentDuring Task
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use2025/08 GitHub starsEvaluative, GuidanceBinary Assessment, RefinementHolisticInitial Setup, During Task
Magentic-UI: Towards Human-in-the-loop Agentic Systems2025/07 GitHub starsEvaluative, Corrective, Guidance, ImplicitBinary Assessment, Refinement, Critique, User ActionSegmentDuring Task, Post Task
UserBench: An Interactive Gym Environment for User-Centric Agents2025/07 GitHub starsImplicit, GuidanceUser Action, RefinementSegmentInitial Setup, During Task
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance2025/07 GitHub starsCorrective, GuidanceRefinement, CritiqueSegmentDuring Task
ฯ„2-Bench: Evaluating Conversational Agents in a Dual-Control Environment2025/06 GitHub starsEvaluative, ImplicitBinary Assessment, User ActionSegment, HolisticInitial Setup, During Task
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training2025/05CorrectiveRefinementSegmentDuring Task
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild2025/05 GitHub starsCorrective, Guidance, EvaluativeRefinement, Binary Assessment, DemonstrationSegment, HolisticDuring Task
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision2025/05 GitHub starsGuidance, EvaluativeCritique, Preference RankingSegmentInitial Setup, Post Task
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft2025/04 GitHub starsEvaluativeScalar RatingHolisticPost Task
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents2025/03ImplicitUser ActionSegmentDuring Task
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks2025/03 GitHub starsCorrective, ImplicitRefinement, User ActionSegmentDuring Task
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting2025/03GuidanceDemonstrationSegmentDuring Task
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration2025/02 GitHub starsGuidanceCritiqueHolisticDuring Task
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments2025/02 GitHub starsGuidanceDemonstration, CritiqueSegmentDuring Task
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation2025/01Corrective, ImplicitUser Action, RefinementSegmentDuring Task
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym)2024/12 GitHub starsCorrectiveRefinementSegmentDuring Task
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension2024/12 GitHub starsGuidanceDemonstrationHolisticInitial Setup, During Task, Post Task
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions2024/10 GitHub starsImplicit, GuidanceDemonstration, User ActionHolisticInitial Setup, During Task
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks2024/10 GitHub starsCorrective, GuidanceRefinement, CritiqueHolistic, SegmentDuring Task, Post Task
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment2024/09 GitHub starsImplicitHuman ControlHolisticInitial Setup
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task2024/09CorrectiveRefinementSegmentDuring Task
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations2024/08 GitHub starsGuidanceDemonstrationSegmentDuring Task
Human-LLM collaboration in generative design for customization2024/07Guidance, EvaluativeDemonstration, Binary Assessment, Preference RankingHolistic, SegmentInitial Setup, Post Task
WebCanvas: Benchmarking Web Agents in Online Environments2024/06EvaluativeScalar RatingHolisticPost Task
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models2024/06Corrective, GuidanceDemonstration, Refinement, CritiqueSegmentInitial Setup, During Task
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration2024/06CorrectiveRefinementHolisticDuring Task
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning2024/06 GitHub starsEvaluativeBinary AssessmentHolisticDuring Task, Post Task
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity2024/05ImplicitHuman ControlSegmentDuring Task
Autonomous Evaluation and Refinement of Digital Agents2024/04 GitHub starsEvaluativeBinary AssessmentHolisticPost Task
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples2024/04Corrective, GuidanceDemonstration, RefinementSegmentDuring Task
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly2024/04Guidance, CorrectiveDemonstration, RefinementSegmentDuring Task
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration2024/04 GitHub starsGuidance, CorrectiveDemonstration, RefinementSegmentInitial Setup
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs2024/04Corrective, GuidanceDemonstration, RefinementSegmentDuring Task
Embodied LLM Agents Learn to Cooperate in Organized Teams2024/03 GitHub starsGuidanceCritiqueHolisticDuring Task
Large Language Model-based Human-Agent Collaboration for Complex Task Solving2024/02 GitHub starsImplicitHuman ControlSegmentDuring Task
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue2024/02 GitHub starsGuidanceDemonstrationHolistic, SegmentDuring Task
Ask-before-Plan: Proactive Language Agents for Real-World Planning2024/01 GitHub starsGuidanceDemonstration, CritiqueSegmentDuring Task
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams2024/01Guidance, Implicit, Corrective, EvaluativeRefinement, Binary Assessment, Critique, Human ControlHolistic, SegmentDuring Task
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination2023/12 GitHub starsGuidance, CorrectiveDemonstration, RefinementSegmentDuring Task
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents2023/10Evaluative, ImplicitScalar Rating, User ActionHolistic, SegmentDuring Task, Post Task
MindAgent: Emergent Gaming Interaction2023/09 GitHub starsCorrectiveRefinementSegmentDuring Task
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles2023/09GuidanceDemonstrationSegmentDuring Task
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback2023/09EvaluativeBinary AssessmentHolisticDuring Task
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework2023/08 GitHub starsEvaluative, GuidanceBinary AssessmentHolisticInitial Setup, Post Task
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks2023/08Corrective, GuidanceDemonstration, RefinementSegmentDuring Task
Embodied Task Planning with Large Language Models2023/07 GitHub starsGuidance, EvaluativeDemonstration, Binary AssessmentHolistic, SegmentInitial Setup, Post Task
Building cooperative embodied agents modularly with large language models2023/07EvaluativeScalar RatingHolisticPost Task
Improved Trust in Human-Robot Collaboration With ChatGPT2023/06GuidanceDemonstration, CritiqueSegmentDuring Task
Investigating Agency of LLMs in Human-AI Collaboration Tasks2023/05GuidanceDemonstration, CritiqueSegmentDuring Task
Improving grounded language understanding in a collaborative environment by interacting with agents through help feedback2023/04Corrective, GuidanceDemonstration, RefinementHolisticDuring Task
PaLM-E: An Embodied Multimodal Language Model2023/03Guidance, ImplicitDemonstration, User ActionSegmentDuring Task
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts2021/10CorrectiveRefinementSegmentDuring Task

๐Ÿ”„ Interaction

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

TitleDate & CodeInteraction TypesInteraction Variant
Just A Rather Very Intelligent Spoken Agent (JarvisBench)2026/07CollaborationSupervision, Cooperation
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability2026/07 GitHub starsCollaborationCooperation, Coordination
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation2026/07CollaborationSupervision, Delegation, Coordination
Uncertainty Decomposition for Clarification Seeking in LLM Agents2026/06CollaborationCooperation
Learning User Simulators with Turing Rewards2026/06CollaborationCooperation
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE)2026/06 GitHub starsCollaborationSupervision, Cooperation
Re-Centering Humans in LLM Personalization2026/06CollaborationCooperation
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement2026/06CollaborationCooperation, Coordination
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents2026/06CollaborationSupervision, Delegation
Uncertainty-Aware Clarification in LLM Agents with Information Gain2026/06CollaborationCooperation
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making2026/05CollaborationSupervision
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions2026/05 GitHub starsCollaborationCooperation, Delegation
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL)2026/05CollaborationCooperation
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement2026/05CollaborationSupervision, Cooperation
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators2026/05CollaborationCoordination
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators2026/05CollaborationCooperation
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows2026/04CollaborationSupervision, Delegation
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks2026/04CollaborationCooperation, Delegation
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench)2026/04 GitHub starsCollaborationSupervision, Cooperation
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents2026/03CollaborationCooperation
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS)2026/03 GitHub starsCollaborationCooperation
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks2026/03 GitHub starsCollaborationDelegation
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science2026/03CollaborationCooperation, Delegation
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration2026/03 GitHub starsCollaborationCooperation
InfoPO: Information-Driven Policy Optimization for User-Centric Agents2026/02 GitHub starsCollaborationCooperation
Modeling Distinct Human Interaction in Web Agents (CowCorpus)2026/02CollaborationSupervision, Delegation, Coordination
Overseeing Agents Without Constant Oversight: Challenges and Opportunities2026/02CollaborationSupervision
Learning Personalized Agents from Human Feedback2026/02 GitHub starsCollaborationCooperation
Value of Information: A Framework for Human-Agent Communication2026/01CollaborationCooperation
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS)2026/01CollaborationCooperation
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering2025/12CollaborationSupervision, Cooperation
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding2025/11CollaborationDelegation, Cooperation
Training Proactive and Personalized LLM Agents2025/11 GitHub starsCollaborationCooperation
Training LLM Agents to Empower Humans2025/10 GitHub starsCollaborationCooperation
How can we assess human-agent interactions? Case studies in software agent design2025/10CollaborationDelegation, Supervision
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback2025/10 GitHub starsCollaborationSupervision, Cooperation
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning2025/09 GitHub starsCollaborationSupervision, Cooperation
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use2025/08 GitHub starsCollaborationDelegation
Magentic-UI: Towards Human-in-the-loop Agentic Systems2025/07 GitHub starsCollaborationCooperation, Coordination
UserBench: An Interactive Gym Environment for User-Centric Agents2025/07 GitHub starsCollaborationCooperation
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance2025/07 GitHub starsCollaborationSupervision
ฯ„2-Bench: Evaluating Conversational Agents in a Dual-Control Environment2025/06 GitHub starsCollaborationCooperation, Coordination
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training2025/05CollaborationCooperation
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild2025/05 GitHub starsCollaborationSupervision, Cooperation
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision2025/05 GitHub starsCollaborationDelegation, Supervision
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft2025/04 GitHub starsCollaborationDelegation
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents2025/03Coopetition-
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting2025/03CollaborationDelegation
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks2025/03 GitHub starsCollaborationDelegation
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments2025/02 GitHub starsCollaborationSupervision, Delegation
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration2025/02 GitHub starsCollaboration-
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation2025/01CollaborationSupervision, Delegation, Coordination
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym)2024/12 GitHub starsCollaborationSupervision, Delegation
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension2024/12 GitHub starsCollaborationCoordination
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions2024/10 GitHub starsCollaborationCoordination
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks2024/10 GitHub starsCollaborationCoordination, Cooperation
AssistantX: An Proactive Assistant in Collaborative Human-Populated Environment2024/09 GitHub starsCollaborationDelegation
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task2024/09CollaborationCoordination
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations2024/08 GitHub starsCollaborationDelegation
Human-LLM Collaboration in Generative Design for Customization2024/07CollaborationDelegation
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning2024/06 GitHub starsCollaborationDelegation
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models2024/06CollaborationDelegation
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration2024/06CollaborationDelegation
WebCanvas: Benchmarking Web Agents in Online Environments2024/06CollaborationDelegation
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity2024/05CollaborationDelegation
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples2024/04CollaborationDelegation, Supervision
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration2024/04 GitHub starsCollaborationCoordination
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly2024/04CollaborationDelegation
Autonomous Evaluation and Refinement of Digital Agents2024/04 GitHub starsCollaborationDelegation
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs2024/04CollaborationDelegation
Embodied LLM Agents Learn to Cooperate in Organized Teams2024/03 GitHub starsCollaborationDelegation
Large Language Model-based Human-Agent Collaboration for Complex Task Solving2024/02 GitHub starsCollaborationDelegation
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue2024/02 GitHub starsCollaborationDelegation
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams2024/01CollaborationCoordination
Ask-before-Plan: Proactive Language Agents for Real-World Planning2024/01 GitHub starsCollaborationCoordination
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination2023/12 GitHub starsCollaborationSupervision, Delegation
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents2023/10Collaboration, Competition, CoopetitionCoordination
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles2023/09CollaborationDelegation
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback2023/09CollaborationDelegation
MindAgent: Emergent Gaming Interaction2023/09 GitHub starsCollaborationCoordination
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks2023/08CollaborationSupervision, Delegation
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework2023/08 GitHub starsCollaborationCoordination
Building Cooperative Embodied Agents Modularly with Large Language Models2023/07CollaborationCooperation
Embodied Task Planning with Large Language Models2023/07 GitHub starsCollaborationDelegation
Improved Trust in Human-Robot Collaboration With ChatGPT2023/06CollaborationDelegation
Investigating Agency of LLMs in Human-AI Collaboration Tasks2023/05CollaborationCooperation, Delegation
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback2023/04CollaborationDelegation
PaLM-E: An Embodied Multimodal Language Model2023/03CollaborationDelegation
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts2021/10CollaborationDelegation

๐ŸŽ›๏ธ Orchestration

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

TitleDate & CodeOrchestration StrategyOrchestration Synchronization
Just A Rather Very Intelligent Spoken Agent (JarvisBench)2026/07SimultaneousAsynchronous
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability2026/07 GitHub starsOne-by-OneSynchronous
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation2026/07One-by-OneSynchronous
PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting2026/06 GitHub starsSimultaneousAsynchronous
Uncertainty Decomposition for Clarification Seeking in LLM Agents2026/06One-by-OneSynchronous
Learning User Simulators with Turing Rewards2026/06One-by-OneSynchronous
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE)2026/06 GitHub starsOne-by-OneSynchronous
Re-Centering Humans in LLM Personalization2026/06One-by-OneAsynchronous
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement2026/06SimultaneousSynchronous
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents2026/06One-by-OneAsynchronous
Uncertainty-Aware Clarification in LLM Agents with Information Gain2026/06One-by-OneSynchronous
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making2026/05One-by-OneSynchronous
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions2026/05 GitHub starsOne-by-OneSynchronous
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL)2026/05One-by-OneSynchronous
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement2026/05One-by-OneSynchronous
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators2026/05One-by-OneSynchronous
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators2026/05One-by-OneSynchronous
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows2026/04One-by-OneAsynchronous
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks2026/04One-by-OneSynchronous
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench)2026/04 GitHub starsOne-by-OneSynchronous
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents2026/03One-by-OneSynchronous
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS)2026/03 GitHub starsOne-by-OneSynchronous
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks2026/03 GitHub starsOne-by-OneAsynchronous
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science2026/03One-by-OneAsynchronous
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration2026/03 GitHub starsOne-by-OneSynchronous
InfoPO: Information-Driven Policy Optimization for User-Centric Agents2026/02 GitHub starsOne-by-OneSynchronous
Modeling Distinct Human Interaction in Web Agents (CowCorpus)2026/02SimultaneousSynchronous
Overseeing Agents Without Constant Oversight: Challenges and Opportunities2026/02One-by-OneAsynchronous
Learning Personalized Agents from Human Feedback2026/02 GitHub starsOne-by-OneSynchronous
Value of Information: A Framework for Human-Agent Communication2026/01One-by-OneSynchronous
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS)2026/01One-by-OneSynchronous
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering2025/12One-by-OneSynchronous
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding2025/11SimultaneousSynchronous
Training Proactive and Personalized LLM Agents2025/11 GitHub starsOne-by-OneSynchronous
Training LLM Agents to Empower Humans2025/10 GitHub starsOne-by-OneSynchronous
How can we assess human-agent interactions? Case studies in software agent design2025/10One-by-OneSynchronous
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback2025/10 GitHub starsOne-by-OneSynchronous
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning2025/09 GitHub starsOne-by-OneSynchronous
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use2025/08 GitHub starsOne-by-OneSynchronous
Magentic-UI: Towards Human-in-the-loop Agentic Systems2025/07 GitHub starsSimultaneousSynchronous
UserBench: An Interactive Gym Environment for User-Centric Agents2025/07 GitHub starsOne-by-OneSynchronous
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance2025/07 GitHub starsOne-by-OneSynchronous
ฯ„2-Bench: Evaluating Conversational Agents in a Dual-Control Environment2025/06 GitHub starsOne-by-OneSynchronous
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training2025/05One-by-OneSynchronous
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild2025/05 GitHub starsOne-by-OneSynchronous
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision2025/05 GitHub starsOne-by-OneSynchronous
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft2025/04 GitHub starsOne-by-OneSynchronous
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents2025/03One-by-OneSynchronous
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting2025/03One-by-OneSynchronous
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks2025/03 GitHub starsOne-by-OneSynchronous
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments2025/02 GitHub starsOne-by-OneAsynchronous
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration2025/02 GitHub starsSimultaneousAsynchronous
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation2025/01One-by-OneSynchronous
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym)2024/12 GitHub starsOne-by-OneAsynchronous
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension2024/12 GitHub starsOne-by-OneAsynchronous
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions2024/10 GitHub starsOne-by-OneSynchronous
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks2024/10 GitHub starsOne-by-OneAsynchronous
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment2024/09 GitHub starsOne-by-OneAsynchronous
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task2024/09SimultaneousSynchronous
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations2024/08 GitHub starsOne-by-OneSynchronous
Human-LLM Collaboration in Generative Design for Customization2024/07One-by-OneSynchronous
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning2024/06 GitHub starsOne-by-OneAsynchronous
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models2024/06One-by-OneSynchronous
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration2024/06One-by-OneSynchronous
WebCanvas: Benchmarking Web Agents in Online Environments2024/06One-by-OneSynchronous
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity2024/05One-by-OneAsynchronous
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples2024/04One-by-OneSynchronous
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration2024/04 GitHub starsOne-by-OneSynchronous
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly2024/04One-by-OneSynchronous
Autonomous Evaluation and Refinement of Digital Agents2024/04 GitHub starsOne-by-OneAsynchronous
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs2024/04One-by-OneSynchronous
Embodied LLM Agents Learn to Cooperate in Organized Teams2024/03 GitHub starsOne-by-OneSynchronous
Large Language Model-based Human-Agent Collaboration for Complex Task Solving2024/02 GitHub starsOne-by-OneSynchronous
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue2024/02 GitHub starsOne-by-OneSynchronous
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams2024/01One-by-OneAsynchronous
Ask-before-Plan: Proactive Language Agents for Real-World Planning2024/01 GitHub starsOne-by-OneSynchronous
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination2023/12 GitHub starsSimultaneousSynchronous
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents2023/10One-by-OneSynchronous
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles2023/09One-by-OneSynchronous
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback2023/09One-by-OneSynchronous
MindAgent: Emergent Gaming Interaction2023/09 GitHub starsOne-by-OneSynchronous
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks2023/08One-by-OneSynchronous
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework2023/08 GitHub starsOne-by-OneAsynchronous
Building Cooperative Embodied Agents Modularly with Large Language Models2023/07SimultaneousSynchronous
Embodied Task Planning with Large Language Models2023/07 GitHub starsOne-by-OneAsynchronous
Improved Trust in Human-Robot Collaboration With ChatGPT2023/06One-by-OneSynchronous
Investigating Agency of LLMs in Human-AI Collaboration Tasks2023/05One-by-OneSynchronous
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback2023/04One-by-OneSynchronous
PaLM-E: An Embodied Multimodal Language Model2023/03One-by-OneSynchronous
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts2021/10One-by-OneSynchronous

๐Ÿ’ฌ Communication

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

TitleDate & CodeCommunication StructureCommunication Mode
Just A Rather Very Intelligent Spoken Agent (JarvisBench)2026/07HierarchicalConversation, Observation
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability2026/07 GitHub starsDecentralizedConversation
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation2026/07Hierarchical, CentralizedConversation, Observation
Uncertainty Decomposition for Clarification Seeking in LLM Agents2026/06CentralizedConversation
Learning User Simulators with Turing Rewards2026/06DecentralizedConversation
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents (TRACE)2026/06 GitHub starsHierarchicalConversation
Re-Centering Humans in LLM Personalization2026/06CentralizedConversation
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement2026/06DecentralizedConversation, Observation
Human Oversight of Agentic Systems in Practice: Examining the Oversight Work, Challenges, and Heuristics of Developers Using Software Agents2026/06HierarchicalObservation, Conversation
Uncertainty-Aware Clarification in LLM Agents with Information Gain2026/06CentralizedConversation
Not All Uncertainty Is Equal: How Uncertainty Granularity Shapes Human Verification in LLM-Assisted Decision Making2026/05CentralizedConversation
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions2026/05 GitHub starsCentralizedConversation
Reinforcing Human Behavior Simulation via Verbal Feedback (DITTO & SOUL)2026/05DecentralizedConversation
ECHO: Explainable Co-editing with Human-in-the-loop Operations for Presentation Refinement2026/05HierarchicalConversation, Observation
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators2026/05DecentralizedMessage Pool
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators2026/05CentralizedConversation, Observation
A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows2026/04HierarchicalConversation
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks2026/04CentralizedConversation
When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation (InterruptBench)2026/04 GitHub starsCentralizedConversation
Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents2026/03HierarchicalConversation
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction (VARS)2026/03 GitHub starsCentralizedConversation
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks2026/03 GitHub starsCentralizedObservation, Conversation
AgentDS: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science2026/03CentralizedConversation
ViviDoc: Generating Interactive Documents through Human-Agent Collaboration2026/03 GitHub starsHierarchicalConversation
InfoPO: Information-Driven Policy Optimization for User-Centric Agents2026/02 GitHub starsCentralizedConversation
Modeling Distinct Human Interaction in Web Agents (CowCorpus)2026/02CentralizedObservation, Conversation
Overseeing Agents Without Constant Oversight: Challenges and Opportunities2026/02HierarchicalObservation
Learning Personalized Agents from Human Feedback2026/02 GitHub starsCentralizedConversation
Value of Information: A Framework for Human-Agent Communication2026/01CentralizedConversation
Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation (MIDAS)2026/01DecentralizedConversation
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering2025/12HierarchicalConversation
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding2025/11CentralizedConversation, Observation
Training Proactive and Personalized LLM Agents2025/11 GitHub starsDecentralizedConversation
Training LLM Agents to Empower Humans2025/10 GitHub starsHierarchicalConversation, Observation
How can we assess human-agent interactions? Case studies in software agent design2025/10HierarchicalConversation
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback2025/10 GitHub starsHierarchicalConversation
UserRL: Training Proactive User-Centric Agent via Reinforcement Learning2025/09 GitHub starsHierarchicalConversation
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use2025/08 GitHub starsHierarchicalConversation
Magentic-UI: Towards Human-in-the-loop Agentic Systems2025/07 GitHub starsHierarchical, CentralizedConversation, Observation
UserBench: An Interactive Gym Environment for User-Centric Agents2025/07 GitHub starsDecentralizedConversation
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance2025/07 GitHub starsHierarchicalConversation
ฯ„2-Bench: Evaluating Conversational Agents in a Dual-Control Environment2025/06 GitHub starsDecentralizedConversation, Observation
Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training2025/05DecentralizedConversation
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild2025/05 GitHub starsCentralizedConversation
XtraGPT: LLMs for Human-AI Collaboration on Controllable Academic Paper Revision2025/05 GitHub starsCentralizedConversation
MineWorld: A Real-Time and Open-Source Interactive World Model on Minecraft2025/04 GitHub starsDecentralizedConversation
Experimental Exploration: Investigating Cooperative Interaction Behavior Between Humans and Large Language Model Agents2025/03DecentralizedConversation
FinArena: A Human-Agent Collaboration Framework for Financial Market Analysis and Forecasting2025/03HierarchicalConversation
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks2025/03 GitHub starsDecentralizedConversation
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments2025/02 GitHub starsDecentralizedConversation
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration2025/02 GitHub starsDecentralizedObservation
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation2025/01DecentralizedConversation
Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension2024/12 GitHub starsDecentralizedConversation
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration (Co-Gym)2024/12 GitHub starsDecentralizedConversation
To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions2024/10 GitHub starsDecentralizedObservation
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-Agent Tasks2024/10 GitHub starsDecentralizedConversation
AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment2024/09 GitHub starsDecentralizedMessage Pool
Mutual Theory of Mind in Human-AI Collaboration: An Empirical Study with LLM-driven AI Agents in a Real-time Shared Workspace Task2024/09DecentralizedConversation
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations2024/08 GitHub starsDecentralizedConversation
Human-LLM Collaboration in Generative Design for Customization2024/07DecentralizedConversation
DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning2024/06 GitHub starsDecentralizedConversation
Enhancing Human-Robot Collaborative Assembly in Manufacturing Systems Using Large Language Models2024/06DecentralizedConversation
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration2024/06DecentralizedConversation
WebCanvas: Benchmarking Web Agents in Online Environments2024/06DecentralizedConversation
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity2024/05HierarchicalObservation
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples2024/04DecentralizedConversation
AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration2024/04 GitHub starsDecentralizedConversation
An LLM-based approach for Enabling Seamless Human-Robot Collaboration in Assembly2024/04CentralizedConversation
Autonomous Evaluation and Refinement of Digital Agents2024/04 GitHub starsDecentralizedConversation
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs2024/04DecentralizedConversation
Embodied LLM Agents Learn to Cooperate in Organized Teams2024/03 GitHub starsDecentralizedConversation
Large Language Model-based Human-Agent Collaboration for Complex Task Solving2024/02 GitHub starsDecentralizedConversation
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue2024/02 GitHub starsDecentralizedConversation
A2C: A Modular Multi-stage Collaborative Decision Framework for Human-AI Teams2024/01DecentralizedConversation
Ask-before-Plan: Proactive Language Agents for Real-World Planning2024/01 GitHub starsCentralizedConversation
LLM-Powered Hierarchical Language Agent for Real-time Human-AI Coordination2023/12 GitHub starsHierarchicalConversation
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents2023/10CentralizedConversation
Drive As You Speak: Enabling Human-Like Interaction With Large Language Models in Autonomous Vehicles2023/09CentralizedConversation
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback2023/09DecentralizedConversation
MindAgent: Emergent Gaming Interaction2023/09 GitHub starsCentralizedConversation
LLM-Based Human-Robot Collaboration Framework for Manipulation Tasks2023/08DecentralizedConversation
MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework2023/08 GitHub starsDecentralizedMessage Pool
Building Cooperative Embodied Agents Modularly with Large Language Models2023/07DecentralizedConversation
Embodied Task Planning with Large Language Models2023/07 GitHub starsDecentralizedConversation
Improved Trust in Human-Robot Collaboration With ChatGPT2023/06DecentralizedConversation
Investigating Agency of LLMs in Human-AI Collaboration Tasks2023/05DecentralizedConversation
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents through Help Feedback2023/04DecentralizedConversation
PaLM-E: An Embodied Multimodal Language Model2023/03DecentralizedConversation
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts2021/10HierarchicalConversation

๐Ÿ“Œ Contributing

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

Contributions are welcome! If you have relevant papers, code, or insights, feel free to submit a request ๐Ÿค—.

๐Ÿ“ Citation

(ยฉ๏ธclick here back to table of contents๐Ÿ‘†๐Ÿป)

If you find this repository useful, please consider citing our papers ๐Ÿ’•:

The survey has been accepted to ACL 2026; the BibTeX below will be updated to the proceedings entry once it is available.

@misc{zou2025llmbasedhumanagentcollaborationinteraction,
      title={LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey}, 
      author={Henry Peng Zou and Wei-Chieh Huang and Yaozu Wu and Yankai Chen and Chunyu Miao and Hoang Nguyen and Yue Zhou and Weizhi Zhang and Liancheng Fang and Langzhou He and Yangning Li and Dongyuan Li and Renhe Jiang and Xue Liu and Philip S. Yu},
      year={2025},
      eprint={2505.00753},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.00753}, 
}

@misc{zou2025collaborativeintelligencehumanagentsystems,
      title={A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy}, 
      author={Henry Peng Zou and Wei-Chieh Huang and Yaozu Wu and Chunyu Miao and Dongyuan Li and Aiwei Liu and Yue Zhou and Yankai Chen and Weizhi Zhang and Yangning Li and Liancheng Fang and Renhe Jiang and Philip S. Yu},
      year={2025},
      eprint={2506.09420},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2506.09420}, 
}

โญ Star History

Star History Chart