Large Language Models for Planning: A Comprehensive and Systematic Survey

July 7, 2025 · View on GitHub

Welcome to the Awesome-LLM-Planning repository! This repository contains a collection of the most influential papers, and benchmarks related to Large Language Models (LLMs) based Agent Planning. For a detailed introduction, please refer our survey paper:

Large Language Models for Planning: A Comprehensive and Systematic Survey arXiv

An overview of LLM-based agent planning, covering its definition, methods, evaluation approaches, and analysis and interpretation.

1 Planning Methods

1.1 External Module Augmented Methods

1.1.1 Planner Enhanced Methods

  • Hierarchical Planning For Complex Tasks With Knowledge Graph-Rag And Symbolic Verification arXiv
  • LLM+ Map: Bimanual Robot Task Planning Using Large Language Models And Planning Domain Definition Language arXiv Code
  • Robust Planning With Compound LLM Architectures: An LLM-Modulo Approach arXiv
  • Can Llms Plan Paths With Extra Hints From Solvers? arXiv
  • Planetarium: A Rigorous Benchmark For Translating Text To Structured Planning Languages NAACL
  • Dynamic Planning With A LLM NeurIPS Workshop
  • Trip-Pal: Travel Planning With Guarantees By Combining Large Language Models And Automated Planners arXiv
  • PDDLEGO: Iterative Planning In Textual Environments StarSEM Code
  • PROC2PDDL: Open-Domain Planning Representations From Texts NLRSE Code
  • Leveraging Environment Interaction For Automated PDDL Translation And Planning With Large Language Models NeurIPS Code
  • Position: Llms Can’t Plan, But Can Help Planning In LLM-Modulo Frameworks ICML
  • Memorybank: Enhancing Large Language Models With Long-Term Memory AAAI
  • Explicit Memory Learning With Expectation Maximization EMNLP
  • To The Globe (Ttg): Towards Language-Driven Guaranteed Travel Planning EMNLP
  • Autoplanbench: Automatically Generating Benchmarks For LLM Planners From PDDL arXiv
  • LLM+P: Empowering Large Language Models with Optimal Planning Proficiency arXiv Code
  • Translating Natural Language To Planning Goals With Large-Language Models arXiv
  • Generative Agents: Interactive Simulacra Of Human Behavior UIST
  • Think Before You Act: Decision Transformers With Working Memory ICML
  • Leveraging Pre-Trained Large Language Models To Construct And Utilize World Models For Model-Based Task Planning NeurIPS Code
  • Synapse: Trajectory-As-Exemplar Prompting With Memory For Computer Control ICLR
  • Coupling Large Language Models With Logic Programming For Robust And General Reasoning From Text ACL
  • Ldm2: A Large Decision Model Imitating Human Cognition With Dynamic Memory Enhancement EMNLP
  • Learning Adaptive Planning Representations With Natural Language Guidance arXiv

1.1.2 Memory Enhanced Methods

  • Crafting Personalized Agents Through Retrieval-Augmented Generation On Editable Memory Graphs EMNLP
  • Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following NeurIPS
  • Large Language Models Are Semi-Parametric Reinforcement Learning Agents NeurIPS
  • A Survey On Large Language Model Based Autonomous Agents Front. Comput. Sci Code
  • Jarvis-1: Open-World Multi-Task Agents With Memory-Augmented Multimodal Language Models TPAMI
  • Learning Memory Mechanisms For Decision Making Through Demonstrations arXiv Code
  • Mot: Memory-Of-Thought Enables Chatgpt To Self-Improve EMNLP
  • Memgpt: Towards Llms As Operating Systems arXiv Code
  • Clin: A Continually Learning Language Agent For Rapid Task Adaptation And Generalization arXiv Code
  • Open-Ended Instructable Embodied Agents With Memory-Augmented Large Language Models EMNLP Code
  • Think-In-Memory: Recalling And Post-Thinking Enable Llms With Long-Term Memory arXiv
  • The Rise And Potential Of Large Language Model Based Agents: A Survey arXiv Code
  • Ghost In The Minecraft: Generally Capable Agents For Open-World Environments Via Large Language Models With Text-Based Knowledge And Memory arXiv Code

1.2 Finetuning-based Methods

1.2.1 Imitation Learning-based Methods

  • Opencodereasoning: Advancing Data Distillation For Competitive Coding arXiv
  • Agentgen: Enhancing Planning Abilities For Large Language Model Based Agent Via Environment And Task Generation SIGKDD Code
  • Mixture-Of-Agents Enhances Large Language Model Capabilities ICLR Code
  • Dualformer: Controllable Fast And Slow Thinking By Learning With Randomized Reasoning Traces ICLR Code
  • Selp: Generating Safe And Efficient Task Plans For Robot Agents With Large Language Models ICRA Code
  • Knowagent: Knowledge-Augmented Planning For LLM-Based Agents NAACL Code
  • Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts In Multi-Objective Alignment arXiv Code
  • Enhancing Reasoning Capabilities Of Llms Via Principled Synthetic Logic Corpus NeurIPS
  • Horizon-Length Prediction: Advancing Fill-In-The-Middle Capabilities For Code Generation With Lookahead Planning arXiv
  • Cooperative Strategic Planning Enhances Reasoning Capabilities In Large Language Models arXiv
  • Camphor: Collaborative Agents For Multi-Input Planning And High-Order Reasoning On Device arXiv
  • Synatra: Turning Indirect Knowledge Into Direct Demonstrations For Digital Agents At Scale NeurIPS Code
  • Learning To Plan Long-Term For Language Modeling COLM
  • Retrieve-Plan-Generation: An Iterative Planning And Answering Framework For Knowledge-Intensive LLM Generation EMNLP Code
  • Ask-Before-Plan: Proactive Language Agents For Real-World Planning EMNLP
  • Learning To Plan For Retrieval-Augmented Large Language Models From Knowledge Graphs EMNLP
  • Distilling Instruction-Following Abilities Of Large Language Models With Task-Aware Curriculum Planning EMNLP
  • Monte Carlo Tree Search Boosts Reasoning Via Iterative Preference Learning arXiv Code
  • Multimodal Web Navigation With Instruction-Finetuned Foundation Models ICLR Code
  • Agent Planning With World Knowledge Model NeurIPS Code
  • Stream Of Search (Sos): Learning To Search In Language COLM Code
  • Agent-Flan: Designing Data And Methods Of Effective Agent Tuning For Large Language Models ACL Code
  • Beyond A*: Better Planning With Transformers Via Search Dynamics Bootstrapping arXiv Code
  • Autoact: Automatic Agent Learning From Scratch Via Self-Planning ACL Code
  • Adapting LLM Agents With Universal Feedback In Communication NAACL
  • You Only Look At Screens: Multimodal Chain-Of-Action Agents ACL Code
  • Fireact: Toward Language Agent Fine-Tuning arXiv Code
  • Training Verifiers To Solve Math Word Problems arXiv

1.2.2 Feedback-based Methods

  • Gui-R1: A Generalist R1-Style Vision-Language Action Model For Gui Agents arXiv Code
  • Flow Of Reasoning: Training Llms For Divergent Problem Solving With Minimal Examples ICML Code
  • Infigui-R1: Advancing Multimodal Gui Agents From Reactive Actors To Deliberative Reasoners arXiv Code
  • Embodied-R: Collaborative Framework For Activating Embodied Spatial Reasoning In Foundation Models Via Reinforcement Learning arXiv
  • Self-Play Preference Optimization For Language Model Alignment ICLR Code
  • Webrl: Training LLM Web Agents Via Self-Evolving Online Curriculum Reinforcement Learning ICLR Code
  • Closed-Loop Long-Horizon Robotic Planning Via Equilibrium Sequence Modeling ICML Code
  • Retool: Reinforcement Learning For Strategic Tool Use In Llms arXiv Code
  • Training Language Models To Self-Correct Via Reinforcement Learning ICLR
  • Toolrl: Reward Is All Tool Learning Needs arXiv Code
  • Can LLM Be A Good Path Planner Based On Prompt Engineering? Mitigating The Hallucination For Path Planning ICIC
  • Otc: Optimal Tool Calls Via Reinforcement Learning arXiv
  • Ui-R1: Enhancing Action Prediction Of Gui Agents By Reinforcement Learning arXivCode
  • Boost, Disentangle, And Customize: A Robust System2-To-System1 Pipeline For Code Generation arXiv
  • Alphamaze: Enhancing Large Language Models' Spatial Intelligence Via Grpo arXiv
  • Swe-Rl: Advancing LLM Reasoning Via Reinforcement Learning On Open Software Evolution arXiv Code
  • Think Smarter Not Harder: Adaptive Reasoning With Inference Aware Optimization arXiv
  • Error Classification Of Large Language Models On Math Word Problems: A Dynamically Adaptive Framework arXiv
  • Enhancing LLM Reasoning Via Critique Models With Test-Time And Training-Time Supervision arXiv Code
  • Thinking Llms: General Instruction Following With Thought Generation arXiv
  • Rest-Mcts*: LLM Self-Training Via Process Reward Guided Tree Search NeurIPS Code
  • Q*: Improving Multi-Step Reasoning For Llms With Deliberative Planning arXiv
  • Learn Beyond The Answer: Training Language Models With Reflection For Mathematical Reasoning EMNLP Code
  • Chain Of Preference Optimization: Improving Chain-Of-Thought Reasoning In Llms NeurIPS Code
  • Cpl: Critical Plan Step Learning Boosts LLM Generalization In Reasoning Tasks arXiv
  • Self-Playing Adversarial Language Game Enhances LLM Reasoning NeurIPS Code
  • Advancing LLM Reasoning Generalists With Preference Trees arXivCode
  • Autowebglm: A Large Language Model-Based Web Navigating Agent SIGKDD Code
  • Llms In The Imaginarium: Tool Learning Through Simulated Trial And Error ACL Code
  • Trial And Error: Exploration-Based Trajectory Optimization For LLM Agents ACL Code
  • React Meets Actre: When Language Agents Enjoy Training Data Autonomy COLM
  • Learning Planning-Based Reasoning By Trajectories Collection And Process Reward Synthesizing EMNLPCode
  • Self-Play Fine-Tuning Converts Weak Language Models To Strong Language Models ICML Code
  • Beyond Human Data: Scaling Self-Training For Problem-Solving With Language Models TMLR
  • A Real-World Webagent With Planning, Long Context Understanding, And Program Synthesis ICLR
  • Language Models Can Teach Themselves To Program Better ICLR

1.3 Searching-based Methods

1.3.1 Decomposition-based Methods

  • Verilogcoder: Autonomous Verilog Coding Agents With Graph-Based Planning And Abstract Syntax Tree (Ast)-Based Waveform Tracing Tool AAAI Code
  • Agent-Oriented Planning In Multi-Agent Systems arXiv Code
  • Optimizing Chain-Of-Thought Reasoning: Tackling Arranging Bottleneck Via Plan Augmentation arXiv
  • Can We Further Elicit Reasoning In Llms? Critic-Guided Planning With Retrieval-Augmentation For Solving Challenging Tasks arXiv
  • Oscar: Operating System Control Via State-Aware Reasoning And Re-Planning arXiv
  • Plan-Rag: Planning-Guided Retrieval Augmented Generation arXiv
  • Msi-Agent: Incorporating Multi-Scale Insight Into Embodied Agents For Superior Planning And Decision-Making EMNLP
  • Travelagent: An AI Assistant For Personalized Travel Planning arXiv
  • Meta-Task Planning For Language Agents arXiv
  • Adapt: As-Needed Decomposition And Planning With Language Models NAACL
  • A Human-Like Reasoning Framework For Multi-Phases Planning Task With Large Language Models arXiv
  • Urbanllm: Autonomous Urban Activity Planning And Management With Large Language Models arXiv
  • Paradise: Evaluating Implicit Planning Skills Of Language Models With Procedural Warnings And Tips Dataset ACL
  • Rada: Retrieval-Augmented Web Agent Planning With Llms ACL
  • Thoughts To Target: Enhance Planning For Target-Driven Conversation EMNLP
  • Personal Large Language Model Agents: A Case Study On Tailored Travel Planning EMNLP
  • Strength Lies In Differences! Improving Strategy Planning For Non-Collaborative Dialogues Via Diversified User Simulation EMNLP
  • Protrix: Building Models For Planning And Reasoning Over Tables With Sentence Context arXiv Code
  • Hugginggpt: Solving AI Tasks With Chatgpt And Its Friends In Hugging Face NeurIPS
  • Distilling Script Knowledge From Large Language Models For Constrained Language Planning ACL
  • Interpretable Math Word Problem Solution Generation Via Step-By-Step Planning ACL
  • Deductive Additivity For Planning Of Natural Language Proofs ACL
  • Least-To-Most Prompting Enables Complex Reasoning In Large Language Models ICLR

1.3.2 Exploration-based Methods

  • Worldcoder, A Model-Based LLM Agent: Building World Models By Writing Code And Interacting With The Environment NeurIPS
  • Scaling Autonomous Agents Via Automatic Reward Modeling And Planning ICLR Code
  • Codetree: Agent-Guided Tree Search For Code Generation With Large Language Models NAACL
  • Flow-Of-Options: Diversified And Improved LLM Reasoning By Thinking Through Options arXiv Code
  • Adaptive Graph Of Thoughts: Test-Time Adaptive Reasoning Unifying Chain, Tree, And Graph Structures arXiv
  • Step Back To Leap Forward: Self-Backtracking For Boosting Reasoning Of Language Models arXivCode
  • Doubly Robust Monte Carlo Tree Search arXiv
  • Webpilot: A Versatile And Autonomous Multi-Agent System For Web Task Execution With Strategic Exploration AAAI
  • Coat: Chain-Of-Associated-Thoughts Framework For Enhancing Large Language Models Reasoning arXiv
  • Qlass: Boosting Language Agent Inference Via Q-Guided Stepwise Search ICLR Workshop
  • Reflective Planning: Vision-Language Models For Multi-Stage Long-Horizon Robotic Manipulation arXiv Code
  • System-1. X: Learning To Balance Fast And Slow Planning With Language Models ICLR Code
  • Is Your LLM Secretly A World Model Of The Internet? Model-Based Planning For Web Agents arXiv Code
  • Reasonplanner: Enhancing Autonomous Planning In Dynamic Environments With Temporal Knowledge Graphs And Llms arXiv
  • Tree Search For Language Model Agents arXiv Code
  • LLM-A*: Large Language Model Enhanced Incremental Heuristic Search On Path Planning EMNLP
  • Graph Of Thoughts: Solving Elaborate Problems With Large Language Models AAAI
  • Language Agent Tree Search Unifies Reasoning Acting And Planning In Language Models ICML Code
  • Tree-Planner: Efficient Close-Loop Task Planning With Large Language Models ICLR
  • Toolllm: Facilitating Large Language Models To Master 16000+ Real-World Apis ICLR
  • Algorithm Of Thoughts: Enhancing Exploration Of Ideas In Large Language Models ICML Code
  • Toolchain*: Efficient Action Space Navigation In Large Language Models With A* Search ICLR
  • Thoughtsculpt: Reasoning With Intermediate Revision And Search NAACL
  • Alphamath Almost Zero: Process Supervision Without Process NeurIPS Code
  • Search-In-The-Chain: Interactively Enhancing Large Language Models With Search For Knowledge-Intensive Tasks WWW
  • Tree Of Thoughts: Deliberate Problem Solving With Large Language Models NeurIPS Code
  • Avis: Autonomous Visual Information Seeking With Large Language Model Agent NeurIPS
  • Large Language Models As Commonsense Knowledge For Large-Scale Task Planning NeurIPS Code
  • Reasoning With Language Model Is Planning With World Model EMNLP
  • Let'S Verify Step By Step ICLR
  • Solving Math Word Problems With Process-And Outcome-Based Feedback arXiv

1.3.3 Decoding-based Methods

  • Non-Myopic Generation Of Language Model For Reasoning And Planning ICLRCode
  • Flap: Flow-Adhering Planning With Constrained Decoding In Llms NAACL
  • Chain-Of-Thought Reasoning Without Prompting NeurIPS
  • Don't Throw Away Your Value Model! Generating More Preferable Text With Value-Guided Monte-Carlo Tree Search Decoding COLM
  • Contrastive Decoding Improves Reasoning In Large Language Models arXiv
  • Grounded Decoding: Guiding Text Generation With Grounded Models For Embodied Agents NeurIPS Code
  • Self-Evaluation Guided Beam Search For Reasoning NeurIPS Code
  • Chain-Of-Thought Prompting Elicits Reasoning In Large Language Models NeurIPS

2 Planning Evaluation

2.1 Datasets

2.1.1 Digital Scenarios

2.1.1.1 Web Navigation
  • WebLINX: Real-World Website Navigation with Multi-Turn Dialogue ICML Code
  • WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models ACL Code
  • VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks ACL Code
  • WebArena: A Realistic Web Environment for Building Autonomous Agents ICLR Code
  • Mind2Web: Towards a Generalist Agent for the Web NeurIPS Code
  • WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents NeurIPS Code
  • Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration ICLR Code
2.1.1.2 Mobile Navigation
  • A3: Android Agent Arena for Mobile GUI Agents arXiv
  • AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents ICLR Code
  • On the Effects of Data Scale on UI Control Agents NeurIPS Code
  • Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction arXiv Code
  • Android in the Wild: A Large-Scale Dataset for Android Device Control NeurIPS Code
  • META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI EMNLP Code
  • Mapping Natural Language Instructions to Mobile UI Action Sequences ACL Code
2.1.1.3 Desktop Navigation
  • AgentStudio: A Toolkit for Building General Virtual Agents ICLR Code
  • OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments NeurIPS Code
  • TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks arXiv Code
  • Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale arXiv Code
  • WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? ICML Code

2.1.2 Embodied Scenarios

2.1.2.1 Household Robot
  • PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks ICLR Code
  • LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments ACL Code
  • ActPlan-1K: Benchmarking the Procedural Planning Ability of Visual Language Models in Household Activities EMNLP Code
  • Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making NeurIPS Code
  • LoTa-Bench: Benchmarking Language-oriented Task Planners for Embodied Agents ICLR Code
  • GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation CVPR Code
  • ScienceWorld: Is your Agent Smarter than a 5th Grader? EMNLP Code
  • BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation CoRL Code
  • ALFWorld: Aligning Text and Embodied Environments for Interactive Learning ICLR Code
  • Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration ICLR Code
  • ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks CVPR Code
  • VirtualHome: Simulating Household Activities via Programs CVPR Code
2.1.2.2 Manipulation Robot
  • EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents ICML Code
  • VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks arXiv Code
  • RoCo: Dialectic Multi-Robot Collaboration with Large Language Models ICRA Code
  • VIMA: General Robot Manipulation with Multimodal Prompts ICML Code
  • VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation NeurIPS Code
2.1.2.3 Minecraft Robot
  • Plancraft: an evaluation dataset for planning with LLM agents arXiv Code
  • TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft arXiv Code
  • Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning ICML Code
  • MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds ICAPS Code
  • MindAgent: Emergent Gaming Interaction NAACL Code
  • MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge NeurIPS Code
  • On the Utility of Learning about Humans for Human-AI Coordination NeurIPS Code
2.1.2.4 Autonomous Driving
  • AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning arXiv Code
  • PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain ACL Code

2.1.3 Everyday Scenarios

2.1.3.1 Travel Planning
  • ChinaTravel: A Real-World Benchmark for Language Agents in Chinese Travel Planning arXiv Code
  • NATURAL PLAN: Benchmarking LLMs on Natural Language Planning arXiv Code
  • TravelPlanner: A Benchmark for Real-World Planning with Language Agents ICML Code
2.1.3.2 Workflow
  • Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment ICLR Code
  • Benchmarking Agentic Workflow Generation ICLR Code
  • FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents EMNLP Code
  • Open Grounded Planning: Challenges and Benchmark Construction ACL Code
  • TaskBench: Benchmarking Large Language Models for Task Automation NeurIPS Code
  • TaskLAMA: Probing the Complex Task Understanding of Language Models AAAI Code
  • Multimodal Procedural Planning via Dual Text-Image Prompting EMNLP Code
  • HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face NeurIPS Code
2.1.3.3 Tool Calling
  • ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark arXiv
  • ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities NAACL Code
  • ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs ICLR Code
  • AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents ACL Code
  • ToolTalk: Evaluating Tool-Usage in a Conversational Setting arXiv Code
  • API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs EMNLP Code
2.1.3.4 Code Generation
  • Training Software Engineering Agents and Verifiers with SWE-Gym ICML Code
  • SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR Code
  • Evaluating Large Language Models Trained on Code arXiv Code
2.1.3.5 Game Playing
  • VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs arXiv Code
  • PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change NeurIPS Code
  • BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning ICLR Code
  • TextWorld: A Learning Environment for Text-based Games IJCAI Code

2.1.4 Vertical Scenarios

2.1.4.1 Machine Learning
  • MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ICLR Code
2.1.4.2 AI Research
  • CycleResearcher: Improving Automated Research via Automated Review ICLR Code
  • ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents arXiv
2.1.4.3 Biological Research
  • BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology EMNLP Code
2.1.4.4 Financial Simulation
  • Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena arXiv Code
2.1.4.5 Interior Design
  • DStruct2Design: Data and Benchmarks for Data Structure Driven Generative Floor Plan Design arXiv Code
  • Tell2Design: A Dataset for Language-Guided Floor Plan Generation ACL Code
2.1.4.6 Comprehensive
  • VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents ICLR Code
  • AgentBench: Evaluating LLMs as Agents ICLR Code

2.2 Evaluation Metrics

The corresponding relationship between planning evaluation metrics and some typical planning datasets. The left three columns represent evaluation metrics of different granularities, while the rightmost column denotes the dataset.

2.3 Performance Comparisons

2.3.1 Web Navigation Performance

  • TongUI: Building Generalized GUI Agents by Learning from Multimodal Web Tutorials arXiv Code
  • SpiritSight Agent: Advanced GUI Agent with One Look CVPR Code
  • UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis arXiv
  • MP-GUI: Modality Perception with MLLMs for GUI Understanding CVPR Code
  • Magma: A Foundation Model for Multimodal AI Agents CVPR Code
  • Digi-Q: Learning Q-Value Functions for Training Device-Control Agents ICLR Code
  • MiniCPM-V: A GPT-4V Level MLLM on Your Phone CVPR Code
  • Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents ICLR Code
  • UI-TARS: Pioneering Automated GUI Interaction with Native Agents arXiv Code
  • Tree Search for Language Model Agents arXiv Code
  • GUICourse: From General Vision Language Models to Versatile GUI Agents arXiv Code
  • GPT-4V(ision) is a Generalist Web Agent, if Grounded ICML Code
  • SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents ACL Code
  • CogAgent: A Visual Language Model for GUI Agents CVPR Code
  • Agent Lumos: Unified and Modular Training for Open-Source Language Agents ACL Code
  • Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control ICLR Code
  • AgentTuning: Enabling Generalized Agent Abilities for LLMs ACL Code
  • Mind2Web: Towards a Generalist Agent for the Web NeurIPS Code

The performance comparison of different models and methods in web navigation.The value of Mind2Web is the average step success rate of the three subsets.The value of Webarena is task successrate.The value of AITW is step success rate of the subsets.The value of ScreenSpot is step success rate.

2.3.2 Embodied Scenarios Performance

  • ATLaS: Agent Tuning via Learning Critical Steps arXiv
  • DebFlow: Automating Agent Creation via Agent Debate CoLM
  • AgentRefine: Enhancing Agent Generalization through Refinement Tuning ICLR Code
  • KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents NAACL Code
  • Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training arXiv Code
  • AgentGym: Evolving Large Language Model-based Agents across Diverse Environments arXiv Code
  • Agent Planning with World Knowledge Model NeurIPS Code
  • Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents ACL Code
  • AgentTuning: Enabling Generalized Agent Abilities for LLMs ACL Code

The performance comparison of different models and methods in embodied.This refers to the average of seen and unseen in the original paper, or the value reported in the original paper.

3 Analysis And Interpretation

3.1 External Interpretation

  • Revealing the Barriers of Language Agents in Planning NAACL Code
  • To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning ICLR Code
  • The pitfalls of next-token prediction ICML Code
  • Chain of Thoughtlessness? An Analysis of CoT in Planning NeurIPS Code
  • Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models arXiv Code
  • Small Language Models Need Strong Verifiers to Self-Correct Reasoning ACL Code
  • A Theoretical Understanding of Self-Correction through In-context Alignment NeurIPS Code
  • Self-Refine: Iterative Refinement with Self-Feedback NeurIPS Code

3.2 Internal Interpretation

  • Do Large Language Models Latently Perform Multi-Hop Reasoning? ACL Code
  • Do language models plan ahead for future tokens? CoLM Code
  • Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language Models EMNLP
  • ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models NeurIPS
  • Iteration Head: A Mechanistic Study of Chain-of-Thought NeurIPS