Cloud GPU performance benchmark recipes

July 22, 2026 ยท View on GitHub

License

This repository contains recipes that provide instructions to reproduce specific workload performance measurements, which are part of a confidential benchmarking program. These recipes focus on helping you reliably achieve performance metrics, such as throughput, that demonstrate the combined hardware and software stack on GPUs.

Note: The recipes in this repository are not designed as general-purpose code samples or tutorials for using Compute Engine-based products.

Intended audience

This content is for you if you are a customer or partner who needs to:

  • Validate hardware performance with your suppliers.
  • Inform purchasing decisions using the benchmarking data.
  • Reproduce optimal performance scenarios before you customize workflows for your own requirements.

How to use these recipes

To reproduce a benchmark, follow these steps:

  1. Identify your requirements: determine the model, GPU type, workload, framework, and orchestrator that you are interested in.
  2. Select a recipe: based on your requirements use the Benchmark support matrix to find a recipe that meets your needs.
  3. Follow the recipe: each recipe will provide you with procedures to complete the following tasks:
    • prepare your environment.
    • run the benchmark.
    • analyze the benchmarks results. This includes not just the results but detailed logs for further analysis. You can automate your infrastructure setup using Cluster Toolkit. For more information, see Automated GPU environment deployment with Cluster Toolkit.

Benchmarks support matrix

Training benchmarks A3 Mega

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
GPT3-175BA3 Mega (NVIDIA H100)NeMo (25.07)Pre-trainingGKELink
Llama-3-70BA3 Mega (NVIDIA H100)NeMo (25.07)Pre-trainingGKELink
Mixtral-8-7BA3 Mega (NVIDIA H100)NeMo (25.07)Pre-trainingGKELink

Training benchmarks A3 Ultra

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Llama-3.1-70BA3 Ultra (NVIDIA H200)MaxTextPre-trainingGKELink
Llama-3.1-70BA3 Ultra (NVIDIA H200)NeMo (24.07)Pre-trainingGKELink
Llama-3-70BA3 Ultra (NVIDIA H200)Megatron-Bridge (26.02)Pre-trainingGKELink
Llama-3-70BA3 Ultra (NVIDIA H200)Megatron-Bridge (25.11)Pre-trainingSlurmLink
Llama-3-8BA3 Ultra (NVIDIA H200)Megatron-Bridge (25.11)Pre-trainingSlurmLink
Llama-3.1-405BA3 Ultra (NVIDIA H200)MaxTextPre-trainingGKELink
Llama-3.1-405BA3 Ultra (NVIDIA H200)NeMo (24.12)Pre-trainingGKELink
Mixtral-8-7BA3 Ultra (NVIDIA H200)NeMo (24.07)Pre-trainingGKELink
DeepSeek-V3A3 Ultra (NVIDIA H200)Megatron-Bridge (26.02)Pre-trainingGKELink
GPT OSS 120BA3 Ultra (NVIDIA H200)NeMo (26.02)Pre-trainingGKELink
Qwen-3-30BA3 Ultra (NVIDIA H200)NeMo (26.02)Pre-trainingGKELink
Wan-2.1A3 Ultra (NVIDIA H200)Megatron-Bridge (26.02)Pre-trainingGKELink

Training benchmarks A4

ModelsGPU Machine TypeFramework / LibraryWorkload TypeOrchestratorLink to the recipe
Llama-3.1-70BA4 (NVIDIA B200)MaxTextPre-trainingGKELink
Llama-3.1-70BA4 (NVIDIA B200)NeMo (25.07)Pre-trainingGKELink
Llama-3.1-70BA4 (NVIDIA B200)NeMo (26.02)Pre-trainingGKELink
Llama-3.1-70BA4 (NVIDIA B200)Megatron-Bridge (25.09)Pre-trainingSlurmLink
Llama-3.1-405BA4 (NVIDIA B200)MaxTextPre-trainingGKELink
Llama-3.1-405BA4 (NVIDIA B200)NeMo (25.07)Pre-trainingGKELink
Llama-3.1-405BA4 (NVIDIA B200)NeMo (26.02)Pre-trainingGKELink
Llama-3.1-405BA4 (NVIDIA B200)Megatron-Bridge (25.09)Pre-trainingSlurmLink
Mixtral-8-7BA4 (NVIDIA B200)NeMo (25.07)Pre-trainingGKELink
PaliGemma2A4 (NVIDIA B200)Hugging Face AccelerateFinetuningGKELink
DeepSeek-V3A4 (NVIDIA B200)Megatron-Bridge (25.11)Pre-trainingGKELink
DeepSeek-V3A4 (NVIDIA B200)Megatron-Bridge (26.02)Pre-trainingGKELink
GPT OSS 120BA4 (NVIDIA B200)Megatron-Bridge (26.02)Pre-trainingGKELink
Llama-3-8BA4 (NVIDIA B200)Megatron-Bridge (26.02)Pre-trainingGKELink
Qwen-3-235BA4 (NVIDIA B200)Megatron-Bridge (25.11)Pre-trainingGKELink
Qwen-3-235BA4 (NVIDIA B200)Megatron-Bridge (26.02)Pre-trainingGKELink
Qwen-3-235BA4 (NVIDIA B200)Megatron-Bridge (25.11)Pre-trainingSlurmLink
Qwen-3-30BA4 (NVIDIA B200)NeMo (26.02)Pre-trainingGKELink
Wan-2.1-14BA4 (NVIDIA B200)NeMo (25.11)Pre-trainingGKELink

Training benchmarks A4X

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Llama-3.1-8BA4X (NVIDIA GB200)NeMo (25.07)Pre-trainingGKELink
Llama-3.1-8BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingGKELink
Llama-3.1-8BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingSlurmLink
Llama-3.1-70BA4X (NVIDIA GB200)NeMo (25.07)Pre-trainingGKELink
Llama-3.1-70BA4X (NVIDIA GB200)Megatron-Bridge (26.02)Pre-trainingGKELink
Llama-3.1-405BA4X (NVIDIA GB200)NeMo (25.07)Pre-trainingGKELink
Llama-3.1-405BA4X (NVIDIA GB200)NeMo (26.02)Pre-trainingGKELink
Llama-3.1-405BA4X (NVIDIA GB200)Megatron-Bridge (26.02)Pre-trainingGKELink
Llama-3.1-405BA4X (NVIDIA GB200)Megatron-Bridge (25.09)Pre-trainingSlurmLink
Nemotron-4-340BA4X (NVIDIA GB200)NeMo (25.09)Pre-trainingGKELink
Wan-2.1-14BA4X (NVIDIA GB200)NeMo (25.11)Pre-trainingGKELink
Wan-2.1-14BA4X (NVIDIA GB200)NeMo (26.02)Pre-trainingGKELink
Wan-2.1-14BA4X (NVIDIA GB200)NeMo (25.11)Pre-trainingSlurmLink
DeepSeek-V3A4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingGKELink
Qwen-3-235BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingGKELink
Qwen-3-235BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingSlurmLink
Qwen-3-30BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingGKELink
Qwen-3-30BA4X (NVIDIA GB200)Megatron-Bridge (25.11)Pre-trainingSlurmLink

Training benchmarks A4X MAX

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
DeepSeek-V3A4X MAX (NVIDIA GB300)Megatron-Bridge (25.11)Pre-trainingGKELink
DeepSeek-V3A4X MAX (NVIDIA GB300)Megatron-Bridge (26.04)Pre-trainingGKELink
DeepSeek-V3A4X MAX (NVIDIA GB300)Megatron-Bridge (26.06)Pre-trainingGKELink
GPT OSS 120BA4X MAX (NVIDIA GB300)Megatron-Bridge (25.11)Pre-trainingGKELink
GPT OSS 120BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.04)Pre-trainingGKELink
Kimi-k2A4X MAX (NVIDIA GB300)Megatron-Bridge (26.04)Pre-trainingGKELink
Llama-3.1-405BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.02)Pre-trainingGKELink
Llama-3.1-405BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.04)Pre-trainingGKELink
Qwen-3-235BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.02)Pre-trainingGKELink
Qwen-3-235BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.04)Pre-trainingGKELink
Wan-2.1-14BA4X MAX (NVIDIA GB300)Megatron-Bridge (26.02)Pre-trainingGKELink

Inference benchmarks A3 Mega

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Llama-4A3 Mega (NVIDIA H100)SGLangInferenceGKELink
DeepSeek R1 671BA3 Mega (NVIDIA H100)SGLangInferenceGKELink
DeepSeek R1 671BA3 Mega (NVIDIA H100)vLLMInferenceGKELink

Inference benchmarks A3 Ultra

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
GPT OSS 120BA3 Ultra (NVIDIA H200)vLLMInferenceGKELink
Llama-4A3 Ultra (NVIDIA H200)vLLMInferenceGKELink
Llama-3.1-405BA3 Ultra (NVIDIA H200)TensorRT-LLMInferenceGKELink
DeepSeek R1 671BA3 Ultra (NVIDIA H200)SGLangInferenceGKELink
DeepSeek R1 671BA3 Ultra (NVIDIA H200)vLLMInferenceGKELink

Inference benchmarks A4

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
DeepSeek R1 671BA4 (NVIDIA B200)vLLMInferenceGKELink
DeepSeek R1 671BA4 (NVIDIA B200)SGLangInferenceGKELink
DeepSeek R1 671BA4 (NVIDIA B200)TensorRT-LLMInferenceGKELink
Llama 3.1 405BA4 (NVIDIA B200)TensorRT-LLMInferenceGKELink
Qwen 2.5 VL 7BA4 (NVIDIA B200)TensorRT-LLMInferenceGKELink
Qwen 3 235B A22BA4 (NVIDIA B200)TensorRT-LLMInferenceGKELink
Qwen 3 32BA4 (NVIDIA B200)TensorRT-LLMInferenceGKELink

Inference benchmarks A4X

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
DeepSeek R1 671BA4X (NVIDIA GB200)vLLM (v0.14.0rc1)InferenceGKELink
Wan2.2 T2V A14B DiffusersA4X (NVIDIA GB200)SGLang (latest)InferenceGKELink
Wan2.2 I2V A14B DiffusersA4X (NVIDIA GB200)SGLang (latest)InferenceGKELink
DeepSeek R1 671BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink

Link for Using Google Cloud Storage (GCS) as Storage Option

Link for Using Lustre as Storage Option
Llama 3.1 405BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Llama 3.1 70BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Llama 3.1 8BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Qwen 2.5 VL 7BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Qwen 3 235B A22BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Qwen 3 32BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink
Qwen 3 4BA4X (NVIDIA GB200)TensorRT-LLM (1.3.0rc5)InferenceGKELink

Inference benchmarks G4

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Qwen3 8BG4 (NVIDIA RTX PRO 6000 Blackwell)vLLMInferenceGCELink
Qwen3 30B A3BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Qwen3 4BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Qwen3 8BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Qwen3 32BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Qwen3 32BG4 (NVIDIA RTX PRO 6000 Blackwell)vLLMInferenceGCELink
Llama3.1 70BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
DeepSeek R1G4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Qwen3 235BG4 (NVIDIA RTX PRO 6000 Blackwell)TensorRT-LLMInferenceGCELink
Wan2.2 14BG4 (NVIDIA RTX PRO 6000 Blackwell)SGLangInferenceGCELink

Checkpointing benchmarks

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Llama-3.1-70BA3 Mega (NVIDIA H100)NeMoPre-training using Google Cloud Storage buckets for checkpointsGKELink

Goodput benchmarks

ModelsGPU Machine TypeFrameworkWorkload TypeOrchestratorLink to the recipe
Llama-3.1-70BA3 Mega (NVIDIA H100)NeMoPre-training using the Google Cloud Resiliency libraryGKELink
Llama-3.1-405BA3 Ultra (NVIDIA H200)NeMoPre-training using the Google Cloud Resiliency libraryGKELink
Mixtral-8x7BA3 Ultra (NVIDIA H200)NeMoPre-training using the Google Cloud Resiliency libraryGKELink

Repository organization

  • ./training: this directory contains recipes with instructions to reproduce training benchmarks with GPUs.
  • ./inference: this directory contains recipes with instructions to reproduce inference benchmarks with GPUs.
  • ./src: this directory contains the shared dependencies required to run benchmarks, such as Docker images and Helm charts.
  • ./docs: this directory contains supporting documentation for explanations of benchmark methodologies or configurations.

Repository scope

This repository provides the steps that you can use to reproduce a specific benchmark. The actual performance measurements and the complete, confidential benchmark report are not included.

Methodology

Performance benchmarks measure the performance of various workloads on the platform. These benchmarks are primarily used to validate performance with hardware suppliers and to provide you with data for purchasing decisions.

Maintenance policy

Benchmark data is considered a point-in-time measurement and completed benchmarks are not repeated. We maintain and update the recipes in this repository on a best-effort basis.

Resources

For general guidance on how to get started using Compute products, refer to the official documentation and tutorials:

Report issues

If you have questions or encounter problems with this repository, report them through GitHub Issues or reach out to your Google Cloud account team for assistance.

Contributor notes

Note: This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.