Awesome-Generalist-Agents [](https://github.com/sindresorhus/awesome) [](https://GitHub.com/Naereen/StrapDown.js/graphs/commit-activity) [](http://makeapullrequest.com)

February 20, 2025 ยท View on GitHub

A curated list of papers for generalist AI agents in both virtual and physical worlds.


Generalist Agents in Both Virtual and Physical Worlds

DatekeywordsPaperPublicationOthers
May 2022GatoA Generalist AgentTMLR'22Report
Feb 2024Interactive Agent Foundation ModelAn Interactive Agent Foundation ModelArXiv'24Report
Feb 2025MagmaMagma: A Foundation Model for Multimodal AI AgentsArXiv'25Project

Generalist Embodied Agents

Large Vision-Language (Action) Models

DatekeywordsPaperPublicationOthers
Dec 2022RT-1RT-1: Robotics Transformer for Real-World Control at ScaleRSS'23Project
Mar 2023PaLM-EPaLM-E: An Embodied Multimodal Language ModelArXiv'23Project
July 2023RT-2RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlArXiv'23Project
Nov 2023LEOAn embodied generalist agent in 3d worldICML'24Project
Nov 2023RoboFlamingoVision-Language Foundation Models as Effective Robot ImitatorsArXiv'23Project
Dec 2023GR-1Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationArXiv'23Project
Mar 20243D-VLA3D-VLA: A 3D Vision-Language-Action Generative World ModelICML'24Project
May 2024OctoOcto: An Open-Source Generalist Robot PolicyArXiv'24Project
Jun 2024OpenVLAOpenVLA: An Open-Source Vision-Language-Action ModelCORL'24Project
Jun 2024RoboUniViewRoboUniView: Visual-Language Model with Unified View Representation for Robotic ManipulationArXiv'24Project
Jul 2024Embodied-CoTRobotic Control via Embodied Chain-of-Thought ReasoningArXiv'24Project
Jun 2024LLARVALLARVA: Vision-Action Instruction Tuning Enhances Robot LearningArXiv'24Project
Sep 2024TinyVLATinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic ManipulationArXiv'24Project
Oct 2024GR-2GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationArXiv'24Project
Oct 2024LAPALatent Action Pretraining from VideosArXiv'24Project
Oct 2024ฯ€0ฯ€0: A Vision-Language-Action Flow Model for General Robot ControlArXiv'24Project
Oct 2024RDT-1BRDT-1B: a Diffusion Foundation Model for Bimanual ManipulationArXiv'24Project
Nov 2024CogACTCogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic ManipulationArXiv'24Project
Nov 2024DeeR-VLADeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionArXiv'24Project
Nov 2024RT-AffordanceRT-Affordance: Affordances are Versatile Intermediate Representations for Robot ManipulationArXiv'24Project
Dec 2024Diffusion-VLADiffusion-VLA: Scaling Robot Foundation Models via Unified Diffusion and AutoregressionArXiv'24Project
Dec 2024RoboVLMsTowards Generalist Robot Policies: What Matters in Building Vision-Language-Action ModelsArXiv'24Project
Dec 2024MotoMoto: Latent Motion Token as the Bridging Language for Robot ManipulationArXiv'24Project
Dec 2024TraceVLATraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic PoliciesArXiv'24Project
Dec 2024NaVILANaVILA: Legged Robot Vision-Language-Action Model for NavigationArXiv'24Project
Jan 2025FASTFAST: Efficient Action Tokenization for Vision-Language-Action ModelsArXiv'25Project
Feb 2025DexVLADexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot ControlArXiv'25Project

Generalist Robotics Policies

DatekeywordsPaperPublicationOthers
Apr 2021Mt-OptMt-Opt: Continuous Multi-Task Robotic Reinforcement Learning at ScaleArXiv'21Project
Jan 2023UniPiLearning Universal Policies via Text-Guided Video GenerationNeurIPS'23Project
Mar 2023MOOOpen-World Object Manipulation using Pre-trained Vision-Language ModelsCoRL'23Project
Jun 2023RoboCatRoboCat: A Self-Improving Generalist Agent for Robotic ManipulationArXiv'23Report
Sep 2023RoboAgentRoboAgent: Generalization and Efficiency in Robot Manipulation via Semantic Augmentations and Action ChunkingICRA'24Project
Feb 2024Extreme Cross-EmbodimentPushing the Limits of Cross-Embodiment Learning for Manipulation and NavigationRSS'24Project
Jun 2024RoboPointRoboPoint: A Vision-Language Model for Spatial Affordance Prediction for RoboticsCORL'24Project
Aug 2024CrossformerScaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and AviationCORL'24Project
Sep 2024HPTScaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersNeurIPS'24Project
Sep 2024RUMsRobot Utility Models: General Policies for Zero-Shot Deployment in New EnvironmentsArXiv'24Project
Sep 2024FLaReFLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-TuningArXiv'24Project
Sep 2024Neural MPNeural MP: A Generalist Neural Motion PlannerArXiv'24Project
Oct 2024Law in ILData Scaling Laws in Imitation Learning for Robotic ManipulationArXiv'24Project
Dec 2024RINGThe One RING: a Robotic Indoor Navigation GeneralistArXiv'24Project
Jan 2025FUSEBeyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language GroundingArXiv'25Project

Multimodal World Models

DatekeywordsPaperPublicationOthers
Mar 2018World ModelsWorld ModelsArXiv'18Project
Jan 2023DreamerV3Mastering Diverse Domains through World ModelsArXiv'23Project
Aug 2023Human World ModelStructured World Models from Human VideosRSS'23Project
Feb 2024World ModelsThe Essential Role of Causality in Foundation World Models for Embodied AIArXiv'24Project
Nov 2024WHALEWHALE: Towards Generalizable and Scalable World Models for Embodied Decision-makingArXiv'24Project

Generlist Web Agents

Generalist Agents for Simulated Worlds

DatekeywordsPaperPublicationOthers
Feb 2024Agent-ProAgent-Pro: Learning to Evolve via Policy-Level Reflection and OptimizationACL'24Project
Dec 2023LARPLARP: Language-Agent Role Play for Open-World GamesArXiv'23Project
Mar 2024SIMAScaling Instructable Agents Across Many Simulated WorldsArXiv'24Report
Aug 2024Optimus-1Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksArXiv'24Project

Generalist Agents for Realistic Tasks

DatekeywordsPaperPublicationOthers
Feb 2023ToolformerToolformer: Language Models Can Teach Themselves to Use ToolsNeurIPS'23Project
Mar 2023RCILanguage Models can Solve Computer TasksArXiv'23Project
Mar 2023HuggingGPTHuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceArXiv'23Project
May 2023Pix2ActFrom Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesNeurIPS'23Project
Jul 2023WebAgentA Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisICLR'24Project
Sep 2023LASERLLM Agent with State-Space Exploration for Web NavigationArXiv'23Project
Sep 2023Auto-GUIYou Only Look at Screens: Multimodal Chain-of-Action AgentsACL'24Project
Sep 2023AgentsAgents: An Open-source Framework for Autonomous Language AgentsArXiv'23Project
Oct 2023AgentTuningAgentTuning: Enabling Generalized Agent Abilities for LLMsArXiv'23Project
Dec 2023CogAgentCogAgent: A Visual Language Model for GUI AgentsCVPR'24Project
Dec 2023AppAgentAppAgent: Multimodal Agents as Smartphone UsersArXiv'23Project
Dec 2023CLOVACLOVA: A Closed-LOop Visual Assistant with Tool Usage and UpdateCVPR 2024Project
Jan 2024SeeActGPT-4V(ision) is a Generalist Web Agent, if GroundedICML'24Project
Jan 2024Mobile-AgentMobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual PerceptionArXiv'24Project
Jan 2024WebVoyagerWebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsACL'24Project
Jan 2024SeeClickSeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsArXiv'24Project
Jan 2024Mobile-AgentMobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual PerceptionArXiv'24Project
Feb 2024OS-CopilotOS-Copilot: Towards Generalist Computer Agents with Self-ImprovementArXiv'24Project
Feb 2024ScreenAgentScreenAgent: A Vision Language Model-driven Computer Control AgentArXiv'24Project
Feb 2024MiddlewareMiddleware for LLMs: Tools Are Instrumental for Language Agents in Complex EnvironmentsEMNLP'2024Project
Apr 2024WILBURWILBUR: Adaptive In-Context Learning for Robust and Accurate Web AgentsArXiv'24Project
Jul 2024OmniParserOmniParser for Pure Vision Based GUI AgentArXiv'24Project
Aug 2024Agent QAgent Q: Advanced Reasoning and Learning for Autonomous AI AgentsArXiv'24Project
Oct 2024OS-ATLASOS-ATLAS: A Foundation Action Model for Generalist GUI AgentsArXiv'24Project
Nov 2024ShowUIShowUI: One Vision-Language-Action Model for GUI Visual AgentArXiv'24Project
Jan 2025InfiGUIAgentInfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and ReflectionArXiv'25Project
Jan 2025UI-TARSUI-TARS: Pioneering Automated GUI Interaction with Native AgentsArXiv'25Project

Datasets & Benchmarks

For Embodied Agents

DatekeywordsPaperPublicationOthers
Jun 2023LIBEROLIBERO: Benchmarking Knowledge Transfer for Lifelong Robot LearningNeurIPS'23Project
Oct 2023Open X-EmbodimentOpen X-Embodiment: Robotic Learning Datasets and RT-X ModelsArXiv'24Project
Oct 2023GenSimGenSim: Generating Robotic Simulation Tasks via Large Language ModelsICLR'24Project
Aug 2024ARIOAll Robots in One: A New Standard and Unified Dataset for Versatile, General-Purpose Embodied AgentsArXiv'24Project
May 2024SimplerEvaluating Real-World Robot Manipulation Policies in SimulationArXiv'24Project
Jun 2024ManiSkill3ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AIArXiv'24Project
Jul 2024RoboCasaRoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist RobotsArXiv'24Project
Jul 2024GRUtopiaGRUtopia: Dream General Robots in a City at ScaleArXiv'24Project
Oct 2024GenesisGenesis: A Generative and Universal Physics Engine for Robotics and BeyondArXiv'24Project
Oct 2024GenSim2GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMsCORL'24Project
Dec 2024RoboMINDRoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot ManipulationArXiv'24Project
Dec 2024VLABenchVLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksArXiv'24Project
Jan 2024MuJoCo PlaygroundMuJoCo PlaygroundReport'24Project

For Web Agents

DatekeywordsPaperPublicationOthers
Jul 2022WebShopTowards Scalable Real-World Web Interaction with Grounded Language AgentsNeurIPS'22Project
May 2023Mobile-EnvMobile-Env: An Evaluation Platform and Benchmark for Interactive Agents in LLM EraArXiv'23Project
Jun 2023Mind2WebMind2Web: Towards a Generalist Agent for the WebNeurIPS'23Project
Jul 2023WebArenaWebArena: A Realistic Web Environment for Building Autonomous AgentsICLR'24Project
Jul 2023ToolBenchToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsICLR'24Project
Jul 2023AITWAndroid in the Wild: A Large-Scale Dataset for Android Device ControlArXiv'23Project
Aug 2023AgentBenchAgentBench: Evaluating LLMs as AgentsArXiv'23Project
Jan 2024VWAVisualwebarena: Evaluating multimodal agents on realistic visual web tasksACL'2024Project
Jan 2024A3A3: Android Agent Arena for Mobile GUI AgentsArXiv'24Project
Feb 2024TravelPlannerTravelplanner: A benchmark for real-world planning with language agentsICML'2024Project
Feb 2024OmniACTOmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and WebArXiv'24Dataset
Mar 2024WorkArenaWorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?ArXiv'24Project
Apr 2024OSWorldOSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsArXiv'24Project
Jul 2024MMAUMMAU: A Holistic Benchmark of Agent Capabilities Across Diverse DomainsArXiv'24Project
Sep 2024WindowsAgentArenaWindows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleArXiv'24Project

General Benchmarks

DatekeywordsPaperPublicationOthers
Aug 2024VisualAgentBenchVisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsArXiv'24Project

๐ŸŒท

We are currently under ongoing updates and always welcome contributions. If you find any interesting papers that are not included in this collection, feel free to open a pull request.

For any questions or suggestions, please contact Yongyuan Liang or Ruihan Yang.