We welcome issues and pull requests for missing work on harness agents, agent skills, memory, self-improvement, agent RL, evaluation, and safety.
[2026-06-25] ๐ Release: Our survey is now available on OpenReview .
If you find this survey or paper list helpful, please cite our work:
@article{jiang2026selfimprovingagents,
title={Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution},
author={Che Jiang and Jincheng Zhong and Yu Fu and Kai Tian and Junlin Yang and Kaikai Zhao and Yuchong Wang and Tianwei Luo and Weizhi Wang and Yuxin Zuo and Guoli Jia and Xingtai Lv and Dianqiao Lei and Sihang Zeng and Yuru Wang and Zhenzhao Yuan and Xinwei Long and Ermo Hua and Can Ren and Xin Jiang and Shulei Xie and Yuanchun Zheng and Youbang Sun and Biqing Qi and Ning Ding and Kaiyan Zhang and Bowen Zhou},
journal={OpenReview Archive},
year={2026},
url={https://openreview.net/pdf?id=IUltZSgLMm}
}
This repository collects papers, systems, benchmarks, and resources for studying how deployed agentic AI systems become more capable after deployment.
We organize the landscape around the harness agent : a deployed runtime system whose behavior is jointly shaped by a base model, a mutable harness, a user-facing interface, and an environment-facing interface. Under this view, self-improvement first appears as fast runtime adaptation over external surfaces such as skills, memory, context, tools, and execution environments. Repeated experience may later be consolidated into model parameters through reinforcement learning, fine-tuning, or continual learning.
We organize the survey into four parts:
Paradigm shift: from task-bounded tool-use loops to deployed harness-centered runtime systems.
External path: skills, memory, context, tools, and environments as editable runtime adaptation surfaces.
Parameter path and meta-evolution: agent RL, continual learning, and meta-agents for durable learning and update orchestration.
Conditions and limits: evaluation, safety, governance, and open problems for reliable post-deployment improvement.
This paper list follows the references cited by the LaTeX manuscript. It currently includes 331 unique cited entries from 379 unique manuscript citation keys and 379 cited BibTeX records.
Every row includes a date, display name, title, and at least one public source badge. Cited entries whose public source URL still needs verification are omitted from this table until complete metadata is available (44 currently omitted; 4 duplicate cited records collapsed).
Date Name Title Paper Github 2026-06 OpenSkillOpenSkill: Open-World Self-Evolution for LLM Agents - 2026-05 huang2026rawexperienceFrom Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills - 2026-05 SkillOptSkillOpt: Executive Strategy for Self-Evolving Agent Skills - 2026-04 cursor2026cursor3Meet the New Cursor - 2026-04 neuralcomputers2026_paperNeural Computers - 2026-04 SWE-chatSWE-chat: Coding Agent Interactions From Real Users in the Wild - 2026-01 AI Agent SystemsAI Agent Systems: Architectures, Applications, and Evaluation - 2025-08 fang2025selfevolvingagentsA Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems - 2025-07 A Survey of Self-Evolving AgentsA Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence - 2025-06 chen2025compoundaisystemsFrom Standalone LLMs to Integrated Intelligence: A Survey of Compound AI Systems - 2025-02 anthropic2025claudecodeClaude 3.7 Sonnet and Claude Code - 2025-01 silver2025eraexperienceWelcome to the Era of Experience - 2024-05 WildChatWildChat: 1M ChatGPT Interaction Logs in the Wild - 2024-04 tao2024selfevolutionsurveyA Survey on Self-Evolution of Large Language Models - 2024-01 langchain2024langgraphLangGraph - 2023-09 xi2023riseagentsurveyThe Rise and Potential of Large Language Model Based Agents: A Survey - 2023-08 wang2023llmagentsurveyA Survey on Large Language Model based Autonomous Agents - 2022-01 ahn2022canDo as i can, not as i say: Grounding language in robotic affordances -
Date Name Title Paper Github 2026-08 OuroborosOuroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution 2026-06 recursive2026automatedresearchFirst Steps Toward Automated AI Research - 2026-06 osmani2026loopengineeringLoop Engineering - 2026-06 Traj-EvolveTraj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection - 2026-05 Agent Harness EngineeringAgent Harness Engineering: A Survey - 2026-04 meng2026agentharnessAgent Harness for Large Language Model Agents: A Survey - 2026-04 boeckeler2026harnessengineeringHarness Engineering for Coding Agent Users - 2026-04 martin2026harnessfailuresMost AI Agent Failures Are Harness Failures - 2026-04 Scaling Managed AgentsScaling Managed Agents: Decoupling the Brain from the Hands - 2026-04 xu2026futureagentsopensourceThe Future of Agents is Open Source (part 1 of 2) - 2026-03 cursor2026composer2Composer 2 Technical Report - 2026-03 yue2026workflowsurveyFrom Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents - 2026-02 Unlocking the Codex harnessUnlocking the Codex harness: how we built the App Server - 2026-01 openharness2026_repoOpenHarness 2025-11 young2025effectiveharnessesEffective Harnesses for Long-Running Agents - 2025-07 mei2025contextengineeringA Survey of Context Engineering for Large Language Models - 2025-05 Darwin Godel MachineDarwin Godel Machine: Open-Ended Evolution of Self-Improving Agents - 2025-05 openai2025codexIntroducing Codex - 2020-01 Artificial IntelligenceArtificial Intelligence: A Modern Approach - 2003-01 Goedel MachinesGoedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements -
Date Name Title Paper Github 2026-07 Skill-SPSkill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills 2026-06 li2026agenticenvironmentengineeringAgentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application - 2026-05 RewardHarnessRewardHarness: Self-Evolving Agentic Post-Training 2026-05 zhou2026ctaCounterfactual Trace Auditing of LLM Agent Skills - 2026-05 Group of SkillsGroup of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries - 2026-05 MIND-SkillMIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction - 2026-05 MUSE-AutoskillMUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation - 2026-05 OpenClaw ResearchOpenClaw Research: A Systematic Survey of Large Language Model Agents in Open Deployment - 2026-05 SkillEvolverSkillEvolver: Skill Learning as a Meta-Skill - 2026-05 SkillRAESkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution - 2026-05 SkillRetSkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents - 2026-04 CoEvoSkillsCoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification - 2026-04 Corpus2SkillCorpus2Skill: Distilling Document Corpora into Hierarchical Skill Directories for Agent Navigation - 2026-04 Externalization in LLM AgentsExternalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering - 2026-04 hivemind2026_repoHivemind: Continual Learning Layer that Distills Coding-Agent Session Trajectories into Reusable Skills 2026-04 skillswild2026realisticHow Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings - 2026-04 SKILL0SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization - 2026-04 SkillXSkillX: Automatically Constructing Skill Knowledge Bases for Agents - 2026-03 bi2026repositoryminingAutomating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction - 2026-03 AutoSkillAutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution - 2026-03 d2skill2026dynamicDynamic Dual-Granularity Skill Bank for Agentic RL - 2026-03 EvoSkillEvoSkill: Automated Skill Discovery for Multi-Agent Systems - 2026-03 From Model to AgentFrom Model to Agent: Equipping the Responses API with a Computer Environment - 2026-03 MetaClawMetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild - 2026-03 SkillNetSkillNet: Create, Evaluate, and Connect AI Skills - 2026-03 SkillRouterSkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale - 2026-03 SWE-Skills-BenchSWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? - 2026-03 Trace2SkillTrace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills - 2026-03 XSkillXSkill: Continual Learning from Experience and Skills in Multimodal Agents - 2026-02 Skill-ProSkill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents - 2026-02 SkillRLSkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning - 2026-02 SkillsBenchSkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks - 2026-01 AutoRefineAutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement - 2026-01 hermesagent2026_repoHermes Agent 2026-01 li2026singleagentskillsWhen Single-Agent with Skills Replace Multi-Agent Systems and When They Fail - 2025-12 sage2025selfimprovingReinforcement Learning for Self-Improving Agent with Skill Library - 2025-10 composeincontext2025skillsCan Language Models Compose Skills In-Context? -
Date Name Title Paper Github 2026-05 Auto-DreamerAuto-Dreamer: Learning Offline Memory Consolidation for Language Agents - 2026-05 MemORAIMemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents - 2026-05 MEMOREPAIRMEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory - 2026-05 zou2026dememRemember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory - 2026-05 SAGESAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory - 2026-04 APEX-MEMAPEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI - 2026-04 CognisCognis: Context-Aware Memory for Conversational AI Agents - 2026-04 DeltaMemDeltaMem: Towards Agentic Memory Management via Reinforcement Learning - 2026-04 zhang2026lightmemLightweight LLM Agent Memory with Small Language Models - 2026-04 TricogExperience is the Best Teacher: Augmenting LLM Reasoning with Knowledge Learned from the Past - 2026-03 zhang2026amacAdaptive Memory Admission Control for LLM Agents - 2026-03 AriadneMemAriadneMem: Threading the Maze of Lifelong Memory for LLM Agents - 2026-03 ChronosChronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory - 2026-03 CLAGCLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents - 2026-03 GAAMAGAAMA: Graph Augmented Associative Memory for Agents - 2026-03 MemoriMemori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents - 2026-03 Memory for Autonomous LLM AgentsMemory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers - 2026-03 PlugMemPlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents - 2026-02 CASTCAST: Character-and-Scene Episodic Memory for Agents - 2026-02 From Lossy to VerifiedFrom Lossy to Verified: A Provenance-Aware Tiered Memory for Agents - 2026-02 Live-EvoLive-Evo: Online Evolution of Agentic Memory from Continuous Feedback - 2026-02 MemoryArenaMemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks - 2026-02 xinlewu2026umemTowards Autonomous Memory Agents - 2026-02 UI-MemUI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents - 2026-02 xMemoryxMemory: Beyond RAG for Agent Memory -- Retrieval by Decoupling and Aggregation - 2026-01 Active Context CompressionActive Context Compression: Autonomous Memory Management in LLM Agents - 2026-01 Agentic MemoryAgentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents - 2026-01 EMemBenchEMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents - 2026-01 H-MemH-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory via a Hybrid Structure - 2026-01 baldelli2026hangmanLLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents - 2026-01 Mem2ActBenchMem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents - 2026-01 MemRLMemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory - 2026-01 SimpleMemSimpleMem: Efficient Lifelong Memory for LLM Agents - 2025-11 WebCoachWebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance - 2025-08 NemoriNemori: Self-Organizing Agent Memory Inspired by Cognitive Science - 2025-03 Search-R1Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning - 2024-10 LongMemEvalLongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory - 2024-04 zhang2024memorymechanismA Survey on the Memory Mechanism of Large Language Model Based Agents - 2023-12 empoweringworkingmemory2023Empowering Working Memory for Large Language Model Agents - 2023-10 MemGPTMemGPT: Towards LLMs as Operating Systems - 2023-03 ReflexionReflexion: Language Agents with Verbal Reinforcement Learning - 2020-05 lewis2020ragRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks - 2016-12 kirkpatrick2017ewcOvercoming catastrophic forgetting in neural networks -
Date Name Title Paper Github 2026-05 zhong2026executablebenchmarkAn Executable Benchmarking Suite for Tool-Using Agents - 2026-05 wu2026chemcostCan Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning - 2026-05 CUA-GymCUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents - 2026-05 MANTRAMANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents - 2026-05 PhysicianBenchPhysicianBench: Evaluating LLM Agents in Real-World EHR Environments - 2026-05 SkillSmithSkillSmith: Compiling Agent Skills into Boundary-Guided Runtime Interfaces - 2026-05 When Simulation LiesWhen Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents - 2026-04 Agent-WorldAgent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence - 2026-04 Agentic World ModelingAgentic World Modeling: Foundations, Capabilities, Laws, and Beyond - 2026-04 Gym-AnythingGym-Anything: Turn any Software into an Agent Environment - 2026-04 zhou2026sandmleSynthetic Sandbox for Training Machine Learning Engineering Agents - 2026-04 ToolMisuseBenchToolMisuseBench: An Offline Deterministic Benchmark for Tool Misuse and Recovery in Agentic Systems - 2026-02 CLI-GymCLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion - 2026-02 shen2026wacWorld-Model-Augmented Web Agents with Action Correction - 2026-01 xiang2026selfevolvingcoevolutionA Systematic Survey of Self-Evolving Agents: From Model-Centric to Environment-Driven Co-Evolution - 2026-01 AG-UIAG-UI: The Agent-User Interaction Protocol 2026-01 a2a_spec_2026Agent2Agent (A2A) Protocol 2026-01 CLI-AnythingCLI-Anything: Making ALL Software Agent-Native 2026-01 HarborHarbor: A Framework for Running Agent Evaluations and Creating and Using RL Environments 2026-01 LiteCoder-TerminalLiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents - 2026-01 MEnvAgentMEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering - 2026-01 openclaw2026_repoOpenClaw 2026-01 SearchGymSearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation - 2026-01 Terminal-BenchTerminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces - 2026-01 WebGymWebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks - 2025-08 SEAgentSEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience - 2025-06 proceduraltooluse2025_paperProcedural Environment Generation for Tool-Use Agents - 2025-05 DeepResearchGymDeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research - 2025-05 MLE-DojoMLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering - 2025-03 chezelles2025browsergymThe BrowserGym Ecosystem for Web Agent Research - 2025-01 mcp_spec_2025Model Context Protocol Specification 2024-12 pan2024swegymTraining Software Engineering Agents and Verifiers with SWE-Gym - 2024-05 AndroidWorldAndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents - 2024-04 OSWorldOSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments - 2024-03 CRADLECRADLE: General Computer Agents with Tool Creation and Knowledge Discovery - 2024-03 DeepSeek-VLDeepSeek-VL: Towards Real-World Vision-Language Understanding - 2024-02 tang2024worldcoderWorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the Environment - 2024-01 AppWorldAppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents - 2024-01 WorkArenaWorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? - 2023-10 SWE-benchSWE-bench: Can Language Models Resolve Real-World GitHub Issues? - 2023-07 WebArenaWebArena: A Realistic Web Environment for Building Autonomous Agents - 2023-04 Generative AgentsGenerative Agents: Interactive Simulacra of Human Behavior - 2023-03 MM-ReActMM-ReAct: Prompting ChatGPT for Multimodal Reasoning and Action - 2022-10 ReActReAct: Synergizing Reasoning and Acting in Language Models - 2020-10 ALFWorldALFWorld: Aligning Text and Embodied Environments for Interactive Learning -
Date Name Title Paper Github 2026-04 Skill-SDSkill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents - 2026-03 cursor2026realtimerlImproving Composer through real-time RL - 2026-03 OpenClaw-RLOpenClaw-RL: Train Any Agent Simply by Talking - 2026-02 xue2026acurlAutonomous Continual Learning of Computer-Use Agents for Environment Adaptation - 2026-02 liu2026empo2Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization - 2026-02 ye2026opcdOn-Policy Context Distillation for Language Models - 2025-12 wei2025selfplayswerlToward Training Superintelligent Software Agents through Self-Play SWE-RL - 2025-11 MemSearcherMemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning - 2025-10 wang2025agenticrlguideA Practitioner's Guide to Multi-turn Agentic Reinforcement Learning - 2025-09 Kimi-DevKimi-Dev: Agentless Training as Skill Prior for SWE-Agents - 2025-09 jin2025rluserconversationsThe Era of Real-World Human Interaction: RL from User Conversations - 2025-09 zhang2026agenticrlsurveyThe Landscape of Agentic Reinforcement Learning for LLMs: A Survey - 2025-08 Agent LightningAgent Lightning: Train ANY AI Agents with Reinforcement Learning - 2025-08 ComputerRLComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents - 2025-05 Absolute ZeroAbsolute Zero: Reinforced Self-play Reasoning with Zero Data - 2025-04 ReToolReTool: Reinforcement Learning for Strategic Tool Use in LLMs - 2025-04 SWE-smithSWE-smith: Scaling Data for Software Engineering Agents - 2025-04 ToolRLToolRL: Reward is All Tool Learning Needs - 2025-02 SWE-RLSWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution - 2021-12 WebGPTWebGPT: Browser-Assisted Question-Answering with Human Feedback - 2021-01 kairouz2021advancesAdvances and Open Problems in Federated Learning - 2020-01 li2020federatedFederated Optimization in Heterogeneous Networks - 2019-01 Towards Federated Learning at ScaleTowards Federated Learning at Scale: System Design - 2017-06 lopezpaz2017gemGradient Episodic Memory for Continual Learning - 2017-01 mcmahan2017communicationCommunication-Efficient Learning of Deep Networks from Decentralized Data -
Date Name Title Paper Github 2026-06 AgonAgon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy - 2026-05 Ace-SkillAce-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution - 2026-05 Continual HarnessContinual Harness: Online Adaptation for Self-Improving Foundation Agents - 2026-05 Skill1Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning - 2026-05 SkillOSSkillOS: Learning Skill Curation for Self-Evolving Agents - 2026-04 Agentic Harness EngineeringAgentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses - 2026-04 AutogenesisAutogenesis: A Self-Evolving Agent Protocol - 2026-04 CORALCORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery 2026-04 Experience as a CompassExperience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts - 2026-04 Meta-TTLLearning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies 2026-04 pan2026mstarM^: Every Task Deserves Its Own Memory Harness - 2026-04 cheng2026mem2evolveMem^2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation - 2026-04 qiao2026miaMemory Intelligence Agent - 2026-04 PRIMEPRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent - 2026-04 RoboPhDRoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets - 2026-04 yang2026memoryextractionSelf-Evolving LLM Memory Extraction Across Heterogeneous Tasks - 2026-04 The World Leaks the FutureThe World Leaks the Future: Harness Evolution for Future Prediction Agents - 2026-03 AgentFactoryAgentFactory: A Self-Evolving Framework Through Executable Subagent Accumulation and Reuse - 2026-03 AI-SupervisorAI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model - 2026-03 ARISEARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning - 2026-03 AutoAgentAutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents - 2026-03 zhang2026hyperagentsHyperagents - 2026-03 Memento-SkillsMemento-Skills: Let Agents Design Agents - 2026-03 Meta-HarnessMeta-Harness: End-to-End Optimization of Model Harnesses - 2026-03 Mimosa FrameworkMimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research - 2026-03 Nurture-First Agent DevelopmentNurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization - 2026-03 RetroAgentRetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback - 2026-03 SAGESAGE: Multi-Agent Self-Evolution for LLM Reasoning - 2026-02 AOrchestraAOrchestra: Automating Sub-Agent Creation for Agentic Orchestration - 2026-02 Group-Evolving AgentsGroup-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing - 2026-02 KernelBlasterKernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning - 2026-02 xiong2026almaLearning to Continually Learn via Meta-learning Agentic Memory Designs - 2026-02 MemSkillMemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents - 2026-02 MetaMemMetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization - 2026-02 PositionPosition: Agentic Evolution is the Path to Evolving LLMs - 2026-02 ROMAROMA: Recursive Open Meta-Agent Framework for Long-Horizon Multi-Agent Systems - 2026-02 SkillOrchestraSkillOrchestra: Learning to Route Agents via Skill Transfer - 2026-02 Tool-R0Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data - 2026-01 GraphPlannerGraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs 2026-01 ye2026mceMeta Context Engineering via Agentic Skill Evolution - 2026-01 MetaGenMetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning - 2025-11 Agent0Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning - 2025-10 MLE-SmithMLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline - 2025-09 MetaEvoMetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution - 2025-08 MetaAgentMetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning - 2025-05 AlphaEvolveAlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms - 2025-05 dang2025evolvingorchestrationMulti-Agent Collaboration via Evolving Orchestration - 2025-04 FlowReasonerFlowReasoner: Reinforcing Query-Level Meta-Agents - 2025-04 TTRLTTRL: Test-Time Reinforcement Learning - 2025-02 A-MEMA-MEM: Agentic Memory for LLM Agents - 2025-02 zhang2025maasMulti-agent Architecture Search via Agentic Supernet - 2024-10 AFlowAFlow: Automating Agentic Workflow Generation - 2024-10 AgentSquareAgentSquare: Automatic LLM Agent Search in Modular Design Space - 2024-08 hu2024adasAutomated Design of Agentic Systems - 2023-05 VoyagerVoyager: An Open-Ended Embodied Agent with Large Language Models -
Date Name Title Paper Github 2026-05 ABRAABRA: Agent Benchmark for Radiology Applications - 2026-05 Agent-BRACEAgent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertainty - 2026-05 wang2026doraCan LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations - 2026-05 raj2026consistencyConsistency as a Testable Property - 2026-05 lee2026ctfusionCTFusion - 2026-05 DataClawDataClaw: A Process-Oriented Agent Benchmark for Exploratory Real-World Data Analysis - 2026-05 gao2026evidencesupportedboundsEvidence-Supported Score Bounds - 2026-05 Evolving-RLEvolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents - 2026-05 surana2026gfcrGenerate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning - 2026-05 wu2026longmemevalv2LongMemEval-V2 - 2026-05 MMTBMMTB: Evaluating Terminal Agents on Multimedia-File Tasks - 2026-05 zhao2026rethinkingexperienceRethinking Experience Utilization in Self-Evolving Language Model Agents - 2026-04 agentbeats_registry2026AgentBeats Dashboard / Agent Registry - 2026-04 agentbeats_docs_aaa2026Agentified Agent Assessment (AAA) & AgentBeats - 2026-04 gurram2026agentpropbenchAgentProp-Bench - 2026-04 ClawArenaClawArena: Benchmarking AI Agents in Evolving Information Environments - 2026-04 ClawBenchClawBench: Can AI Agents Complete Everyday Online Tasks? 2026-04 EvoAgentBenchEvoAgentBench: A Multi-Domain Benchmark for Self-Evolving Agents - 2026-04 chi2026frontierengFrontier-Eng - 2026-04 SEARLSEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents - 2026-04 SkillLearnBenchSkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks - 2026-03 agentmemorybench2026Benchmarking Continual Agent Memory for Online Learning, Transfer, and Forgetting - 2026-03 CUBECUBE: A Standard for Unifying Agent Benchmarks - 2026-03 DomusMindDomusMind: A Benchmark for Evaluating Lifelong Smart Home Agents Under Drift - 2026-02 Agent World ModelAgent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning - 2026-02 ResearchGymResearchGym: Evaluating Language Model Agents on Real-World AI Research - 2026-02 SE-BenchSE-Bench: Benchmarking Self-Evolution with Knowledge Internalization - 2026-02 When AI Benchmarks PlateauWhen AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation - 2026-01 DevOps-GymDevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle - 2026-01 SIP-BenchSIP-Bench: An Open Protocol for Longitudinal Self-Improvement Evaluation 2025-07 mohammadi2025agentbenchmarkingEvaluation and Benchmarking of LLM Agents: A Survey - 2025-07 SWE-MERASWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks - 2025-06 liu2025cerContextual Experience Replay for Self-Improvement of Language Agents - 2025-05 LifelongAgentBenchLifelongAgentBench: Evaluating LLM Agents as Lifelong Learners - 2025-05 zhang2025swebenchliveSWE-bench Goes Live! - 2025-05 SWE-rebenchSWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents - 2025-03 yehudai2025agentevaluationSurvey on Evaluation of LLM-based Agents - 2025-01 cai2025buildingBuilding self-evolving agents via experience-driven lifelong learning: A framework and benchmark - 2024-12 TheAgentCompanyTheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks - 2024-10 MLE-benchMLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering - 2024-09 wang2025awmAgent Workflow Memory - 2024-06 tau-benchtau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains - 2024-03 LiveCodeBenchLiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code - 2024-02 maharana2024locomoEvaluating Very Long-Term Conversational Memory of LLM Agents - 2023-08 AgentBenchAgentBench: Evaluating LLMs as Agents - 2023-07 ToolLLMToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs - 2019-04 vandeven2019threescenariosThree Scenarios for Continual Learning -
Date Name Title Paper Github 2026-05 wu2026bivBehavioral Integrity Verification for AI Agent Skills - 2026-05 liang2026mobiusCan a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection - 2026-05 LITMUSLITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments - 2026-05 yin2026fateOn-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment - 2026-05 ProteusProteus: A Self-Evolving Red Team for Agent Skill Ecosystems - 2026-05 SARCSARC: A Governance-by-Architecture Framework for Agentic AI Systems - 2026-05 ShadowMergeShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts - 2026-05 SkillScopeSkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills - 2026-05 SkillsVoteSkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution - 2026-05 STALESTALE: Can LLM Agents Know When Their Memories Are No Longer Valid? - 2026-05 Under the Hood of SKILL.mdUnder the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry - 2026-04 AgentWatcherAgentWatcher: A Rule-Based Prompt Injection Monitor - 2026-04 ATBenchATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis - 2026-04 Claw-EvalClaw-Eval: Toward Trustworthy Evaluation of Autonomous Agents - 2026-04 chang2026visualinjectionsIf You're Waiting for a Sign... That Might Not Be It! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems - 2026-04 MemEvoBenchMemEvoBench: Benchmarking Memory MisEvolution in LLM Agents - 2026-04 SafeAgentSafeAgent: A Runtime Protection Architecture for Agentic Systems - 2026-04 SkillClawSkillClaw: Let Skills Evolve Collectively with Agentic Evolver - 2026-04 SkillForgeSkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support - 2026-04 Sovereign Agentic LoopsSovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems - 2026-04 SporeSpore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing - 2026-03 From Storage to SteeringFrom Storage to Steering: Memory Control Flow Attacks on LLM Agents - 2026-03 lam2026ssgmGoverning Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework - 2026-03 SkillProbeSkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration - 2026-03 SkillTesterSkillTester: Benchmarking Utility and Security of Agent Skills - 2026-02 agentskills2026architectureAgent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward - 2026-02 AgentSysAgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management - 2026-02 ClawHavocClawHavoc: 341 Malicious Clawed Skills Found by the Bot They Were Targeting - 2026-02 SoKSoK: Agentic Skills -- Beyond Tool Use in LLM Agents - 2026-01 WildClawBenchWildClawBench: An In-the-Wild Benchmark for AI Agents in the OpenClaw Environment 2025-12 MemoryGraftMemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval - 2025-10 A-MemGuardA-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory - 2025-09 InjecMEMInjecMEM: Memory Injection Attack on LLM Agent Memory Systems - 2025-07 shanghai2025frontierriskFrontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report - 2025-07 SafeWork-R1SafeWork-R1: Coevolving Safety and Intelligence under the AI-45^ Law - 2025-06 su2025autonomyriskA Survey on Autonomy-Induced Security Risks in Large Model-Based Agents - 2025-06 Context Manipulation AttacksContext Manipulation Attacks: Web Agents Are Susceptible to Corrupted Memory - 2025-06 DRIFTDRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents - 2025-06 ferrag2025promptprotocolFrom Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows - 2025-06 RedDebateRedDebate: Safer Responses through Multi-Agent Red Teaming Debates - 2025-06 fang2025safemcpWe Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems - 2025-03 AutoRedTeamerAutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration - 2024-12 Agent-SafetyBenchAgent-SafetyBench: Evaluating the Safety of LLM Agents - 2024-12 yang2024ai45lawTowards AI-45^ Law: A Roadmap to Trustworthy AGI - 2024-10 zhang2025asbAgent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-Based Agents - 2024-03 IsolateGPTIsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems - 2024-02 Agent SmithAgent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast - 2024-02 dong2024conversationsafetyAttacks, Defenses and Evaluations for LLM Conversation Safety: A Survey - 2022-12 Constitutional AIConstitutional AI: Harmlessness from AI Feedback - 2022-03 ouyang2022instructgptTraining Language Models to Follow Instructions with Human Feedback -
This repository is maintained by the FrontisAI and Tsinghua University survey team. Its README structure follows the public awesome-list style of TsinghuaC3I/Awesome-RL-for-LRMs .