vLLM

May 10, 2026 · View on GitHub

vLLM is a high-performance, easy-to-use engine for LLM inference and serving. For an in-depth overview, see the official documentation.

This directory contains vLLM deployment guides for different MiniCPM versions. Use the links below to jump to the specific version.

Versioned Deployment Guides

VersionEnglish中文
MiniCPM-V 4.6 (latest)Guide指南
MiniCPM-V 4.5Guide指南
MiniCPM-V 4.0Guide指南
MiniCPM-o 4.5Guide指南
MiniCPM-o 2.6Guide指南
MiniCPM-V 2.6Guide指南
MiniCPM-V 2.5Guide指南