Awesome-Minecraft-Agents
May 24, 2026 ยท View on GitHub
Our Minecraft Agent
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
We propose a Hybrid Multimodal Memory module that integrates structured knowledge and multimodal experiences into the memory mechanism of the agent. On top of it, we introduce a powerful Minecraft agent, Optimus-1, which achieves a 30% improvement over existing agents on 67 long-horizon tasks. โจ
Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy
We propose agent Optimus-2 which incorporates a Multimodal Large Language Model for high-level planning, alongside a Goal-Observation-Action Conditioned Policy (GOAP) for low-level control. Optimus-2 exhibits superior performance across atomic tasks, long-horizon tasks, and open-ended instruction tasks in Minecraft. โจ
Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task Experts
We propose generalist agent, Optimus-3, endowed with multidimensional capabilities including Captioning, Embodied QA, Planning, Action, Grounding, and Reflection. Our comprehensive evaluation demonstrates that it consistently surpasses existing agents in the Minecraft environment across all assessed dimensions. โจ
Awesome Policy
Visuomotor Policy
| Title | Venue | Year | Code | Demo |
|---|---|---|---|---|
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos | NeurIPS | 2022 | Github | - |
GROOT: Learning to Follow Instructions by Watching Gameplay Videos | ICLR | 2024 | Github | Demo |
Goal-conditioned Policy
Awesome Agent
End-to-end Architecture
| Title | Venue | Year | Code | Demo |
|---|---|---|---|---|
| JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse | Findings of ACL | 2025 | Github | Demo |
Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task Experts | Arxiv | 2025 | Github | Demo |