vLLM
May 10, 2026 · View on GitHub
vLLM is a high-performance, easy-to-use engine for LLM inference and serving. For an in-depth overview, see the official documentation.
This directory contains vLLM deployment guides for different MiniCPM versions. Use the links below to jump to the specific version.
Versioned Deployment Guides
| Version | English | 中文 |
|---|---|---|
| MiniCPM-V 4.6 (latest) | Guide | 指南 |
| MiniCPM-V 4.5 | Guide | 指南 |
| MiniCPM-V 4.0 | Guide | 指南 |
| MiniCPM-o 4.5 | Guide | 指南 |
| MiniCPM-o 2.6 | Guide | 指南 |
| MiniCPM-V 2.6 | Guide | 指南 |
| MiniCPM-V 2.5 | Guide | 指南 |