Large models on multiple GPUs

August 23, 2026 · View on GitHub

Recommended use
Start withStep 3.7 Flash on the qualified 2× RTX PRO 6000 Blackwell PP-2 path
TopologyPipeline parallelism for the published Step configuration
ValidatePlacement, cross-device transport, admission, and end-to-end output on the exact topology
Read nextStep model card, Serving: PP-2

Pipeline stages, tensor parallelism, expert parallelism, and replicas are different execution shapes. Use the name and gate for the topology actually running; do not call PP or replicas TP.