Inference Benchmarks

July 20, 2026 · View on GitHub

These tables collect inference benchmarks for Cosmos3. Generator sections measure visual-generation and world-model latency across PyTorch, vLLM-Omni, Diffusers, and NIM, including image and video generation, forward and inverse dynamics, and policy generation. Reasoner sections measure VLM serving and token-generation performance for text, image, and video inputs through vLLM and Hugging Face Transformers.

Generator results are published incrementally from internal benchmark runs. Empty cells mean that combination has not been measured yet — not that it is unsupported. See the notes under each table for workload details and data-source definitions.

Table of Contents

Cosmos3-Edge Generator

These tables report Cosmos3-Edge Generator latency in seconds for image-to-video, forward dynamics, inverse dynamics, and DROID policy generation. Measurements use one GPU or one integrated computing platform. Lower latency is better, and empty cells indicate that a run has not been completed.

Unless otherwise noted, visual-generation benchmarks use 480p resolution, and image-to-video benchmarks generate 189 frames. vLLM-Omni values report end-to-end latency, while PyTorch values report average generation latency.

vLLM-Omni

GPU or PlatformImage-to-VideoForward DynamicsInverse DynamicsPolicy DROID
B200 SXM 192 GB2.443.980.99
H100 SXM 80 GB27.643.915.601.41
H100 NVL 96 GB35.604.736.391.37
H20 SXM 96 GB108.1612.7715.493.41
RTX PRO 6000 Blackwell Server Edition36.295.657.461.87
DGX Station12.174.336.348.11
DGX Spark165.9626.4330.867.66
Jetson AGX Thor T5000, 128 GB, MAXN137.506.057.196.32
Jetson T3000, 32 GB, 1100 MHz194.768.6710.258.63
Jetson T2000, 16 GB, 702 MHz, THOR_NANO101.20

PyTorch

GPU or PlatformImage-to-VideoForward DynamicsInverse DynamicsPolicy DROID
H100 SXM 80 GB23.923.693.561.25
H100 NVL 96 GB32.244.644.521.28
H20 SXM 96 GB97.5112.7812.642.92
RTX PRO 6000 Blackwell Server Edition38.985.265.661.32
DGX Station10.572.162.261.30
DGX Spark179.8024.5926.765.44
Jetson AGX Thor T5000, 128 GB, MAXN153.00
Jetson T3000, 32 GB, 1100 MHz227.80

Notes:

  1. All measurements use one GPU or one integrated computing platform.
  2. Values are average end-to-end or generation latency in seconds; lower is better.
  3. Unless otherwise specified, visual-generation measurements use 480p resolution.
  4. Image-to-video measurements generate 189 output frames.
  5. Jetson AGX Thor T5000 and Jetson T3000 visual-generation measurements use 832 × 480 resolution.
  6. Jetson T2000 visual-generation measurements use 448 × 256 resolution and therefore should not be compared directly with the 480p results. Its image-to-video values are warm-run measurements generating 189 frames.
  7. PyTorch values report average generation latency rather than diffusion-only latency.
  8. Datacenter and enterprise forward- and inverse-dynamics results use the autonomous-driving (AV) configuration.
  9. Jetson AGX Thor T5000 and Jetson T3000 forward-dynamics, inverse-dynamics, and policy measurements use the DROID configuration with action chunk [16, 8].

Cosmos3-Nano Generator

These tables report Cosmos3-Nano generator latency in seconds for image-to-video (i2v), text-to-image (t2i), and text-to-video (t2v) - the primary vision-generation modes of the omni-model. Benchmarks use BF16 precision, batch size 1, and identical prompts, seeds, and sampler settings across engines where noted below. Video workloads follow the standard Cosmos3 generation profile (189 frames at 24 FPS unless a resolution tier limits frame count).

Four integration paths are compared. PyTorch reports average generation (sampling) time from OSS reference inference with CUDA Graphs enabled where supported. vLLM-Omni reports total pipeline time at 720p on supported GPUs. Diffusers reports end-to-end generation time through the Hugging Face Cosmos3OmniPipeline without custom CUDA graphs at 256p/1, 480p/1, and 720p/1 (320×192, 832×480, and 1280×720). NIM reports end-to-end request latency using NIM latency profiles with FP8 precision, including request processing, video generation, output encoding, and returning the response. Empty cells indicate a run has not been completed for that GPU, engine, resolution, or tensor-parallel width - tables are filled in as benchmark campaigns finish.

Text-to-Video (t2v)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch13.954.90180.81786.37225.45127.57
vLLM-Omni10.655.063.78105.6135.9323.76369.67114.3068.66
Diffusers11.20112.00392.00
NIM7.934.694.0082.6933.5724.49318.69107.8368.50
H20PyTorch30.57257.51931.39268.88157.71
vLLM-Omni28.5810.207.70256.9777.4247.53929.81260.75148.46
Diffusers30.20258.00926.00
NIM17.817.846.14192.0762.8041.67771.37223.24132.68
H100 NVLPyTorch10.034.273.9584.1229.1821.46297.2794.1561.63
vLLM-Omni9.253.683.1580.7527.4818.77311.1388.25(*)54.01(*)
Diffusers11.0090.00324.20
NIM7.093.773.6568.0125.6720.51267.7386.1857.38
H200 NVLPyTorch8.1769.79244.3977.3545.70
vLLM-Omni7.443.272.3364.5821.3112.92240.0569.6339.17
Diffusers9.0074.00276.20
NIM5.873.383.0057.1021.7415.04229.6371.3243.34
H100 80GB HBM3PyTorch7.613.503.1759.8321.2314.37207.7866.9441.81
vLLM-Omni6.973.453.4958.1719.9513.46202.2962.8237.80
Diffusers9.0068.00240.00
NIM5.723.363.1251.7320.2614.82199.4665.3241.66
H200 141GB HBM3PyTorch7.533.343.1960.1820.8413.97214.2867.4841.26
vLLM-Omni6.793.253.4258.1419.7712.97208.3663.2737.49
Diffusers9.0067.00239.60
NIM5.723.263.1051.9020.0914.61200.6364.5740.70
B200PyTorch4.562.782.7933.2013.209.69114.8539.7526.27
vLLM-Omni4.032.433.4932.0412.6310.09107.8435.2922.87
Diffusers7.0036.80117.00
NIM3.722.742.8926.6812.6610.1093.3335.0424.64
B300PyTorch
vLLM-Omni4.464.115.4432.1813.8311.57102.1035.6824.33
Diffusers39.4063.40139.40
NIM4.513.544.1127.4914.5111.8690.3937.2826.19

Image-to-Video (i2v)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch182.14788.80226.25127.79
vLLM-Omni11.045.484.24107.7738.0525.95375.01119.2773.57
Diffusers12.00112.00397.00
NIM8.235.284.6484.5835.7426.63326.96112.2773.31
H20PyTorch31.36257.10933.07268.99158.10
vLLM-Omni29.5011.268.64261.5681.9352.06940.16271.37158.76
Diffusers31.00258.00925.00
NIM18.679.047.38195.1067.6446.30774.88233.92143.15
H100 NVLPyTorch10.194.313.9984.5028.6921.52298.5795.7660.58
vLLM-Omni9.624.113.6382.6129.3520.73286.3392.23(*)58.02(*)
Diffusers11.0091.00325.20
NIM7.394.364.2869.3927.6722.43272.2990.2661.55
H200 NVLPyTorch8.2769.99246.6277.6945.99
vLLM-Omni7.833.692.7866.3922.9314.58243.5273.2642.86
Diffusers9.0074.00275.20
NIM6.253.973.6058.6423.3316.75232.4775.1347.22
H100 80GB HBM3PyTorch7.643.473.2159.9521.4014.43207.8767.5241.66
vLLM-Omni7.373.813.9759.7721.6815.12205.9766.5241.51
Diffusers9.0068.00239.80
NIM6.083.893.7153.0222.2016.61202.5969.0245.16
H200 141GB HBM3PyTorch7.653.373.1760.5121.0114.07214.8067.1441.00
vLLM-Omni7.283.633.8359.6421.3514.67209.6566.6540.77
Diffusers9.0067.20240.00
NIM6.043.803.6553.2521.8916.23203.6668.3044.29
B200PyTorch4.602.772.8113.079.66113.9040.0126.58
vLLM-Omni4.332.773.8433.0913.7911.39110.1937.7625.68
Diffusers116.00
NIM4.053.433.6027.7214.0211.4395.5737.7627.29
B300PyTorch
vLLM-Omni5.614.675.9033.4515.0613.13104.7538.2726.87
Diffusers28.6065.60139.60
NIM4.504.965.0328.9016.0613.2592.5939.4929.00

Text-to-Image (t2i)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch2.994.517.123.182.70
vLLM-Omni1.591.541.592.871.551.814.992.321.96
Diffusers2.004.005.00
H20PyTorch3.066.5112.314.283.06
vLLM-Omni1.732.463.224.922.577.2410.734.243.59
Diffusers3.006.0010.00
H100 NVLPyTorch2.772.452.572.832.562.514.212.572.64
vLLM-Omni1.551.751.911.921.8110.823.441.831.90
Diffusers3.003.004.00
H200 NVLPyTorch2.853.582.622.64
vLLM-Omni1.532.011.961.581.9117.712.811.941.94
Diffusers3.003.004.00
H100 80GB HBM3PyTorch3.012.662.563.012.592.753.452.732.77
vLLM-Omni1.612.453.181.532.357.022.612.453.03
Diffusers3.003.004.00
H200 141GB HBM3PyTorch2.962.592.703.042.782.773.282.842.77
vLLM-Omni1.572.383.161.522.377.052.602.333.20
Diffusers3.003.004.00
B200PyTorch2.392.592.752.432.562.872.582.62
vLLM-Omni1.492.213.271.202.057.581.772.203.41
Diffusers3.00
B300PyTorch
vLLM-Omni1.974.525.821.814.1671.192.344.095.62
Diffusers36.2041.00

Notes:

  1. All times measured on identical workloads (same seed, sampler settings, prompt).
  2. 4×/8× GPU configurations use tensor parallelism.
  3. vLLM-Omni numbers are for the upcoming public release in the vLLM-Omni repo; subject to change before GA. Values marked with (*) are pre-release vLLM-Omni measurements on H100 NVL and may change before GA.
  4. Diffusers numbers use the HuggingFace diffusers integration without custom CUDA graphs; reported at 256p/1, 480p/1, and 720p/1 (single-GPU only).
  5. PyTorch numbers report average generation (sampling) time from OSS inference benchmarking.
  6. At 256p, multi-GPU configurations on B300 may underperform single-GPU due to small-workload TP overhead; single-GPU is recommended at this resolution.
  7. NIM numbers use latency profiles with FP8 precision and report end-to-end Request Latency s, including request processing, video generation, output encoding, and returning the response.

Cosmos3-Super Generator

These tables report Cosmos3-Super generator latency in seconds for image-to-video (i2v), text-to-image (t2i), and text-to-video (t2v). The 32B checkpoint targets higher-quality world generation; expect longer runtimes than Nano at the same resolution. Benchmarks use BF16 precision, batch size 1, and matched prompts, seeds, and sampler settings. Video workloads follow the standard Cosmos3 profile (189 frames at 24 FPS where applicable).

As with Nano, four engines are tracked: PyTorch (OSS generation/sampling time), vLLM-Omni (total pipeline time at 720p on supported GPUs), Diffusers (Hugging Face Cosmos3OmniPipeline end-to-end time at 256p/1, 480p/1, and 720p/1 — 320×192, 832×480, and 1280×720), and NIM (end-to-end request latency using latency profiles with FP8 precision, including request processing, video generation, output encoding, and returning the response). Super coverage is narrower than Nano in early releases - for example, vLLM-Omni and Diffusers runs exist primarily on B200 and select H200 configurations. Empty cells are pending measurements, not unsupported configurations.

Text-to-Video (t2v)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch789.03427.16
vLLM-Omni
Diffusers
NIM12.6513.99104.2599.05350.74286.02
H20PyTorch492.41
vLLM-Omni
Diffusers
NIM20.0712.95192.45110.71734.37395.56
H100 NVLPyTorch16.83101.2764.14330.04186.19
vLLM-Omni
Diffusers
NIM8.7712.7373.3766.07267.64197.32
H200 NVLPyTorch258.34139.37
vLLM-Omni27.545.06252.3336.66911.49245.51123.85
Diffusers33.00286.801036.00
NIM17.136.794.43200.0058.5532.87811.41223.00117.98
H100 80GB HBM3PyTorch
vLLM-Omni
Diffusers
NIM6.985.8955.1035.52198.13114.92
H200 141GB HBM3PyTorch14.8211.8270.2741.78224.43123.49
vLLM-Omni25.615.87219.1135.26769.63212.30111.94
Diffusers31.00251.60886.20
NIM15.956.144.28174.7152.9430.94695.89194.34106.16
B200PyTorch5.594.09114.3835.7321.39407.50118.3865.93
vLLM-Omni13.844.76114.0822.09390.28113.3162.11
Diffusers127.20414.40
NIM9.094.263.3882.3927.8317.74314.6892.2553.43
B300PyTorch
vLLM-Omni14.576.68109.0322.67366.66108.5860.73
Diffusers54.20155.40424.80
NIM9.675.235.1979.7328.9718.39292.3592.3154.07

Image-to-Video (i2v)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch795.14427.96
vLLM-Omni
Diffusers
NIM13.1714.48106.50100.23356.60289.93
H20PyTorch931.74
vLLM-Omni
Diffusers
NIM21.1614.06196.86114.49745.12405.68
H100 NVLPyTorch20.8516.9699.56331.40186.47
vLLM-Omni
Diffusers
NIM9.3213.3075.3167.34271.46201.39
H200 NVLPyTorch265.33138.31
vLLM-Omni27.905.52254.2938.51915.05248.89127.32
Diffusers33.00287.201034.60
NIM17.517.435.04201.4560.1534.48817.35226.38121.35
H100 80GB HBM3PyTorch
vLLM-Omni
Diffusers
NIM7.506.4956.3336.81200.97118.77
H200 141GB HBM3PyTorch14.8711.8070.4542.10224.36123.57
vLLM-Omni25.476.32220.7036.90766.33215.03117.52
Diffusers31.00249.20879.20
NIM16.396.744.90175.9554.4632.51699.13197.96109.55
B200PyTorch14.715.634.1235.7021.25397.31117.9865.91
vLLM-Omni14.135.31115.1723.26393.02115.6964.82
Diffusers19.20414.80
NIM9.364.834.1283.1929.0919.14316.7694.6555.92
B300PyTorch
vLLM-Omni14.147.19111.4223.91368.73111.4163.25
Diffusers54.20151.80425.00
NIM9.735.585.9480.5130.1720.62294.1193.7756.76

Text-to-Image (t2i)

GPUEngine256p/1256p/4256p/8480p/1480p/4480p/8720p/1720p/4720p/8
RTX PRO 6000 BlackwellPyTorch92.6193.11
vLLM-Omni
Diffusers
H20PyTorch20.92
vLLM-Omni
Diffusers
H100 NVLPyTorch19.7319.8620.6819.87
vLLM-Omni
Diffusers
H200 NVLPyTorch32.6433.16
vLLM-Omni2.733.206.2817.7111.024.223.26
Diffusers5.008.0012.00
H100 80GB HBM3PyTorch
vLLM-Omni
Diffusers
H200 141GB HBM3PyTorch13.6213.4813.3313.5313.7813.50
vLLM-Omni2.834.505.707.1610.244.234.43
Diffusers5.008.0011.00
B200PyTorch4.104.274.784.134.487.254.284.65
vLLM-Omni2.324.583.299.106.023.094.43
Diffusers8.00
B300PyTorch
vLLM-Omni5.057.623.7972.657.085.797.24
Diffusers38.8039.4040.40

Notes:

  1. All times measured on identical workloads (same seed, sampler settings, prompt).
  2. 4×/8× GPU configurations use tensor parallelism.
  3. vLLM-Omni numbers are for the upcoming public release in the vLLM-Omni repo; subject to change before GA. Current vLLM-Omni coverage is B200 at 720p.
  4. Diffusers numbers use the HuggingFace diffusers integration without custom CUDA graphs; reported at 256p/1, 480p/1, and 720p/1 (single-GPU only).
  5. At 256p, multi-GPU configurations on B300 may underperform single-GPU due to small-workload TP overhead; single-GPU is recommended at this resolution.
  6. PyTorch numbers report average generation (sampling) time from OSS inference benchmarking.
  7. NIM numbers use latency profiles with FP8 precision and report end-to-end Request Latency s, including request processing, video generation, output encoding, and returning the response.

Cosmos3-Edge Reasoner

These tables report Cosmos3-Edge reasoner serving performance through vLLM. Unlike the Generator benchmarks, Reasoner workloads produce autoregressively generated text and measure time to first token (TTFT), end-to-end request latency, request throughput, and output-token throughput. Lower is better for latency metrics; higher is better for throughput.

All vLLM runs use the nvidia/Cosmos3-Edge checkpoint with one GPU. Metrics were collected at client-side concurrency levels of 1, 64, 128, and 256. Each GPU section contains four workload tables that vary input sequence length, output sequence length, and video frame rate.

RTX PRO 4500 Blackwell Server Edition

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)165.798817.3314702.2029482.39
Request Latency (ms)165.798817.3314702.2029482.39
Request Count (requests)50320256512
Request Throughput (Req/s)6.006.556.556.52
Output Token Throughput (Tok/s)6.006.556.556.52

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)371.6720375.9833812.4568201.55
Request Latency (ms)371.6720375.9833812.4568201.55
Request Count (requests)50313249492
Request Throughput (Req/s)2.682.772.762.71
Output Token Throughput (Tok/s)2.682.772.762.71

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)166.866900.9019625.8345729.55
Request Latency (ms)764.1516667.0129196.8455749.62
Request Count (requests)50320256512
Request Throughput (Req/s)1.313.733.743.70
Output Token Throughput (Tok/s)130.63372.40373.98369.87

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)374.9323526.6547550.99101553.31
Request Latency (ms)1041.2933712.5457641.53111895.20
Request Count (requests)50320256512
Request Throughput (Req/s)0.961.791.791.78
Output Token Throughput (Tok/s)95.74178.73178.89178.15

RTX PRO 6000 Blackwell Server Edition

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)141.993213.915384.5110792.72
Request Latency (ms)141.993213.915384.5110792.72
Request Count (requests)50320254512
Request Throughput (Req/s)6.9618.0017.9517.89
Output Token Throughput (Tok/s)6.9618.0017.9517.89

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)239.867483.2212552.6925259.11
Request Latency (ms)239.867483.2212552.6925259.11
Request Count (requests)49303249491
Request Throughput (Req/s)4.067.287.497.34
Output Token Throughput (Tok/s)4.067.287.497.34

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)138.74943.462680.1711599.63
Request Latency (ms)503.446188.9013022.0726388.89
Request Count (requests)50320256512
Request Throughput (Req/s)1.9810.279.578.95
Output Token Throughput (Tok/s)197.751026.14956.47893.91

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)239.241798.9611644.8433293.32
Request Latency (ms)638.7113599.8926299.9049165.91
Request Count (requests)50320256512
Request Throughput (Req/s)1.564.664.504.45
Output Token Throughput (Tok/s)155.93465.28449.57444.17

Embedded-Platform Eager Transformers

These preliminary measurements use raw Hugging Face Transformers in eager mode rather than vLLM. They are presented separately because their runtime, workloads, and metric definitions differ from the vLLM serving benchmarks above.

BoardSpecificationInputPrompt TokensPrefill ThroughputPrefill LatencyDecode ThroughputEnd-to-End Latency
Jetson AGX Thor T5000128 GB / MAXNText17058717 Tok/s0.20 s37.3 Tok/s3.60 s
Jetson AGX Thor T5000128 GB / MAXNImage9114845 Tok/s0.19 s42.6 Tok/s3.17 s
Jetson AGX Thor T5000128 GB / MAXNVideo12636032 Tok/s0.21 s41.8 Tok/s3.25 s
Jetson AGX Thor T400064 GB / MAXN, 1530 MHzText17056519 Tok/s0.26 s34.1 Tok/s3.99 s
Jetson AGX Thor T400064 GB / MAXN, 1530 MHzImage9113471 Tok/s0.26 s40.3 Tok/s3.41 s
Jetson AGX Thor T400064 GB / MAXN, 1530 MHzVideo12634164 Tok/s0.30 s38.1 Tok/s3.64 s
Jetson Thor T300032 GB / 1100 MHzText17055230 Tok/s0.33 s29.7 Tok/s4.61 s
Jetson Thor T300032 GB / 1100 MHzImage9112710 Tok/s0.34 s36.3 Tok/s3.83 s
Jetson Thor T300032 GB / 1100 MHzVideo12633388 Tok/s0.37 s33.7 Tok/s4.14 s
Jetson Thor T200016 GB / 702 MHz, THOR_NANOText17052355 Tok/s0.72 s15.7 Tok/s8.80 s
Jetson Thor T200016 GB / 702 MHz, THOR_NANOImage9111233 Tok/s0.74 s19.6 Tok/s7.21 s
Jetson Thor T200016 GB / 702 MHz, THOR_NANOVideo12631543 Tok/s0.82 s18.0 Tok/s7.87 s
Jetson AGX Orin64 GBText17053260 Tok/s0.52 s12.3 Tok/s10.83 s
Jetson AGX Orin64 GBImage9111840 Tok/s0.50 s12.3 Tok/s10.81 s
Jetson AGX Orin64 GBVideo12632103 Tok/s0.60 s12.2 Tok/s10.97 s

Notes:

  1. Source: vLLM inference benchmarking for nvidia/Cosmos3-Edge; metrics were collected with one GPU at client-side concurrency levels of 1, 64, 128, and 256.
  2. Time To First Token (TTFT) measures latency until the first output token is emitted. Request Latency is end-to-end time per request. For single-token outputs (Output 1), TTFT and request latency are identical.
  3. Request Throughput is completed requests per second. Output Token Throughput is generated tokens per second. For Output 1 workloads, the two throughput values match.
  4. Concurrency is the number of simultaneous client requests, not tensor-parallel GPU count.
  5. Embedded-platform measurements use Hugging Face Transformers in eager mode and should not be compared directly with the vLLM serving results.

Cosmos3-Nano Reasoner

These tables report Cosmos3-Nano reasoner serving performance through vLLM. Unlike the generator sections, Reasoner benchmarks measure text understanding and generation latency - time to first token (TTFT) in milliseconds, end-to-end request latency in milliseconds, and token throughput under concurrent load - not diffusion sampling time. Workloads vary input sequence length, output sequence length, and video frame rate to reflect common captioning, VQA, and video-understanding request profiles.

All runs use the nvidia/Cosmos3-Nano checkpoint. Metrics are collected with the AIPerf client at client-side concurrency levels of 1, 64, 128, and 256. Each GPU section below contains four workload tables (Input 50 / Output 1 or 100 / Video 1 or 2 FPS). Lower is better for latency metrics; higher is better for throughput.

RTX PRO 6000 Blackwell

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)187.595826.849742.4319541.84
Request Latency (ms)187.595826.849742.4319541.84
Request Count (requests)50320256512
Request Throughput (Req/s)5.299.959.979.89
Output Token Throughput (Tok/s)5.299.959.979.89

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)316.9012223.0020364.0440929.42
Request Latency (ms)316.9012223.0020364.0440929.42
Request Count (requests)50320256512
Request Throughput (Req/s)3.144.734.754.71
Output Token Throughput (Tok/s)3.144.734.754.71

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)186.462280.434627.0814419.32
Request Latency (ms)1402.129309.9318541.9039202.74
Request Count (requests)50320256512
Request Throughput (Req/s)0.716.856.826.22
Output Token Throughput (Tok/s)71.22684.76682.18622.49

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)315.773248.7213795.4544476.55
Request Latency (ms)1553.5318532.3437994.0571534.87
Request Count (requests)50320256512
Request Throughput (Req/s)0.643.443.223.15
Output Token Throughput (Tok/s)64.28343.79322.21314.62

H20

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)358.6214953.9025086.3049549.94
Request Latency (ms)358.6214953.9025086.3049549.94
Request Count (requests)50320256512
Request Throughput (Req/s)2.773.883.873.90
Output Token Throughput (Tok/s)2.773.883.873.90

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)648.4830604.9151364.19101597.85
Request Latency (ms)648.4830604.9151364.19101597.85
Request Count (requests)50320256512
Request Throughput (Req/s)1.541.891.891.90
Output Token Throughput (Tok/s)1.541.891.891.90

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)360.406607.8010973.0529404.60
Request Latency (ms)1026.9718990.4837514.2174287.25
Request Count (requests)50320256512
Request Throughput (Req/s)0.973.373.403.33
Output Token Throughput (Tok/s)97.14336.55339.57332.93

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)646.278329.8229036.8186145.80
Request Latency (ms)1331.0037577.6874291.62136416.07
Request Count (requests)50320256512
Request Throughput (Req/s)0.751.701.671.67
Output Token Throughput (Tok/s)75.00170.03167.36167.12

H100 NVL

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)170.696527.1310726.8021881.52
Request Latency (ms)170.696527.1310726.8021881.52
Request Count (requests)50320256512
Request Throughput (Req/s)5.818.889.058.83
Output Token Throughput (Tok/s)5.818.889.058.83

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)303.6513480.2922431.9344352.53
Request Latency (ms)303.6513480.2922431.9344352.53
Request Count (requests)50320256512
Request Throughput (Req/s)3.274.294.314.35
Output Token Throughput (Tok/s)3.274.294.314.35

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)172.662890.815022.5813929.58
Request Latency (ms)867.359192.1918061.4335151.09
Request Count (requests)50320256512
Request Throughput (Req/s)1.156.947.026.95
Output Token Throughput (Tok/s)115.03694.37702.48695.12

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)296.993572.3113808.0341101.98
Request Latency (ms)1009.4118030.4135239.2564485.08
Request Count (requests)50320256512
Request Throughput (Req/s)0.993.543.483.50
Output Token Throughput (Tok/s)98.87353.81348.37350.07

H200 NVL

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)142.793614.376050.5812094.34
Request Latency (ms)142.793614.376050.5812094.34
Request Count (requests)50320256512
Request Throughput (Req/s)6.9216.0416.0815.96
Output Token Throughput (Tok/s)6.9216.0416.0815.96

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)228.847569.4612515.9125646.99
Request Latency (ms)228.847569.4612515.9125646.99
Request Count (requests)50320256512
Request Throughput (Req/s)4.347.627.717.48
Output Token Throughput (Tok/s)4.347.627.717.48

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)142.231948.063180.205271.37
Request Latency (ms)770.155284.5810054.5519831.69
Request Count (requests)50320256512
Request Throughput (Req/s)1.3012.0712.6012.71
Output Token Throughput (Tok/s)129.531206.861259.601270.44

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)227.402718.135522.0517729.47
Request Latency (ms)862.9210249.1419775.3339089.75
Request Count (requests)50320256512
Request Throughput (Req/s)1.166.226.386.18
Output Token Throughput (Tok/s)115.63621.97638.33618.21

H100 80GB HBM3 (SXM)

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)145.523332.725608.4111133.76
Request Latency (ms)145.523332.725608.4111133.76
Request Count (requests)50320256512
Request Throughput (Req/s)6.7817.4117.3617.38
Output Token Throughput (Tok/s)6.7817.4117.3617.38

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)228.806876.2511556.3322836.32
Request Latency (ms)228.806876.2511556.3322836.32
Request Count (requests)50320256512
Request Throughput (Req/s)4.348.428.388.46
Output Token Throughput (Tok/s)4.348.428.388.46

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)143.361720.732906.009000.63
Request Latency (ms)865.565251.569818.7318353.75
Request Count (requests)50320256512
Request Throughput (Req/s)1.1512.1412.8712.83
Output Token Throughput (Tok/s)115.241213.611286.891282.48

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)231.812295.448767.7823214.88
Request Latency (ms)967.499738.1818061.7933190.08
Request Count (requests)50320256512
Request Throughput (Req/s)1.036.546.536.60
Output Token Throughput (Tok/s)103.12653.91653.39659.92

H200 141GB HBM3

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)142.453363.015656.5811271.28
Request Latency (ms)142.453363.015656.5811271.28
Request Count (requests)50320256512
Request Throughput (Req/s)6.9317.2517.2117.17
Output Token Throughput (Tok/s)6.9317.2517.2117.17

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)229.706932.5511640.6823173.10
Request Latency (ms)229.706932.5511640.6823173.10
Request Count (requests)50320256512
Request Throughput (Req/s)4.318.358.328.33
Output Token Throughput (Tok/s)4.318.358.328.33

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)143.252060.052839.974713.83
Request Latency (ms)711.774965.099364.2018325.25
Request Count (requests)50320256512
Request Throughput (Req/s)1.4012.8413.5313.75
Output Token Throughput (Tok/s)140.021284.381352.411374.57

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)228.242715.855015.1915991.33
Request Latency (ms)807.079341.2118285.6535043.34
Request Count (requests)50320256512
Request Throughput (Req/s)1.246.826.906.90
Output Token Throughput (Tok/s)123.55682.33690.26689.50

B200

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)115.551661.572819.225550.74
Request Latency (ms)115.551661.572819.225550.74
Request Count (requests)50320256512
Request Throughput (Req/s)8.5334.9634.7234.94
Output Token Throughput (Tok/s)8.5334.9634.7234.94

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)168.953410.275699.8811422.16
Request Latency (ms)168.953410.275699.8811422.16
Request Count (requests)50320256512
Request Throughput (Req/s)5.8616.9817.0116.93
Output Token Throughput (Tok/s)5.8616.9817.0116.93

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)115.271106.352111.972549.79
Request Latency (ms)553.012736.535001.209279.25
Request Count (requests)50320256512
Request Throughput (Req/s)1.8023.2825.2327.01
Output Token Throughput (Tok/s)180.162328.012523.072701.08

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)166.521914.242596.367277.89
Request Latency (ms)622.364881.389220.3017548.01
Request Count (requests)50320256512
Request Throughput (Req/s)1.6013.0413.6313.87
Output Token Throughput (Tok/s)160.111303.921362.491386.99

B300

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)80.701617.122742.685421.22
Request Latency (ms)80.701617.122742.685421.22
Request Count (requests)50320256511
Request Throughput (Req/s)12.2435.9235.6935.79
Output Token Throughput (Tok/s)12.2435.9235.6935.79

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)126.253304.765551.1211054.47
Request Latency (ms)126.253304.765551.1211054.47
Request Count (requests)50320256512
Request Throughput (Req/s)7.8617.5317.4917.49
Output Token Throughput (Tok/s)7.8617.5317.4917.49

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)83.191070.931444.682739.50
Request Latency (ms)490.112657.064750.028975.21
Request Count (requests)50320256512
Request Throughput (Req/s)2.0323.9626.5727.92
Output Token Throughput (Tok/s)203.292396.352657.142791.79

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)129.221602.002404.486982.58
Request Latency (ms)550.024684.788813.8316813.61
Request Count (requests)50320256512
Request Throughput (Req/s)1.8113.5914.2514.47
Output Token Throughput (Tok/s)181.251358.621425.021447.23

Notes:

  1. Source: vLLM inference benchmarking for nvidia/Cosmos3-Nano; AIPerf client was used as the benchmarking tool.
  2. Hardware: results are grouped by GPU product (RTX PRO 6000 Blackwell, H20, H100 NVL, H200 NVL, H100 80GB HBM3 SXM, H200 141GB HBM3, B200, B300). All metrics are averages for a number of requests.
  3. Time To First Token (TTFT) measures latency until the first output token is emitted. Request Latency is end-to-end time per request. For single-token outputs (Output 1), TTFT and request latency are identical.
  4. Request Throughput is completed requests per second. Output Token Throughput is generated tokens per second (for Output 1 workloads, the two throughputs match).
  5. Concurrency is the number of simultaneous client requests issued by AIPerf, not tensor-parallel GPU count.

Cosmos3-Super Reasoner

These tables report Cosmos3-Super reasoner serving performance through vLLM. Unlike the generator sections, Reasoner benchmarks measure text understanding and generation latency - time to first token (TTFT) in milliseconds, end-to-end request latency in milliseconds, and token throughput under concurrent load - not diffusion sampling time. Workloads vary input sequence length, output sequence length, and video frame rate to reflect common captioning, VQA, and video-understanding request profiles.

All runs use the nvidia/Cosmos3-Super checkpoint. Metrics are collected with the AIPerf client at client-side concurrency levels of 1, 64, 128, and 256. Each GPU section below contains four workload tables (Input 50 / Output 1 or 100 / Video 1 or 2 FPS). Lower is better for latency metrics; higher is better for throughput. Empty cells indicate a run has not been completed for that GPU, workload, or concurrency level.

RTX PRO 6000 Blackwell

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)534.7324781.4741467.4582626.36
Request Latency (ms)534.7324781.4741467.4582626.36
Request Count (requests)50320256509
Request Throughput (Req/s)1.862.342.342.32
Output Token Throughput (Tok/s)1.862.342.342.32

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)978.7851145.6185476.50
Request Latency (ms)978.7851145.6185476.50
Request Count (requests)50320256
Request Throughput (Req/s)1.021.131.13
Output Token Throughput (Tok/s)1.021.131.13

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)530.4725094.9054400.80117849.75
Request Latency (ms)5225.7940064.2269193.77133019.50
Request Count (requests)50320256512
Request Throughput (Req/s)0.191.511.501.51
Output Token Throughput (Tok/s)19.12151.27149.69151.15

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)981.55114704.46
Request Latency (ms)5716.49130177.35
Request Count (requests)50256
Request Throughput (Req/s)0.170.77
Output Token Throughput (Tok/s)17.4977.00

H20

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)1241.58108912.73
Request Latency (ms)1241.58108912.73
Request Count (requests)50256
Request Throughput (Req/s)0.800.89
Output Token Throughput (Tok/s)0.800.89

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)2399.91
Request Latency (ms)2399.91
Request Count (requests)50
Request Throughput (Req/s)0.42
Output Token Throughput (Tok/s)0.42

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)1230.95108988.21
Request Latency (ms)3523.74135784.46
Request Count (requests)50256
Request Throughput (Req/s)0.280.78
Output Token Throughput (Tok/s)28.3677.81

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)2380.85
Request Latency (ms)4707.46
Request Count (requests)50
Request Throughput (Req/s)0.21
Output Token Throughput (Tok/s)21.22

H100 NVL

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)521.8727004.4245688.9590353.80
Request Latency (ms)521.8727004.4245688.9590353.80
Request Count (requests)50320256512
Request Throughput (Req/s)1.912.152.132.14
Output Token Throughput (Tok/s)1.912.152.132.14

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)993.4855409.8392484.18
Request Latency (ms)993.4855409.8392484.18
Request Count (requests)50320256
Request Throughput (Req/s)1.001.051.05
Output Token Throughput (Tok/s)1.001.051.05

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)508.9726567.6554861.82116733.31
Request Latency (ms)3119.7639090.1667203.81129435.49
Request Count (requests)50320256512
Request Throughput (Req/s)0.321.541.551.54
Output Token Throughput (Tok/s)32.03153.87154.58154.04

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)999.1049069.05116178.34
Request Latency (ms)3638.2359084.57128875.27
Request Count (requests)50320256
Request Throughput (Req/s)0.271.000.77
Output Token Throughput (Tok/s)27.46100.1477.38

H200 NVL

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)357.6916243.1126759.3554470.31
Request Latency (ms)357.6916243.1126759.3554470.31
Request Count (requests)50319254510
Request Throughput (Req/s)2.783.563.583.53
Output Token Throughput (Tok/s)2.783.563.583.53

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)641.0533640.6856090.59111965.21
Request Latency (ms)641.0533640.6856090.59111965.21
Request Count (requests)50320255510
Request Throughput (Req/s)1.561.721.721.71
Output Token Throughput (Tok/s)1.561.721.721.71

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)348.885805.6316053.6848411.16
Request Latency (ms)2385.9521240.6240187.1375354.93
Request Count (requests)50320256512
Request Throughput (Req/s)0.423.013.062.99
Output Token Throughput (Tok/s)41.87300.56305.57298.96

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)640.1413800.4047514.83110513.13
Request Latency (ms)2692.4641460.8074683.53138991.66
Request Count (requests)50320256512
Request Throughput (Req/s)0.371.521.511.52
Output Token Throughput (Tok/s)37.12151.97151.42152.07

H200 141GB HBM3

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)327.2114045.0123809.0046893.25
Request Latency (ms)327.2114045.0123809.0046893.25
Request Count (requests)50320256507
Request Throughput (Req/s)3.044.144.094.08
Output Token Throughput (Tok/s)3.044.144.094.08

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)592.1928769.6848595.0495884.55
Request Latency (ms)592.1928769.6848595.0495884.55
Request Count (requests)50320256512
Request Throughput (Req/s)1.682.011.992.01
Output Token Throughput (Tok/s)1.682.011.992.01

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)327.735374.5514328.5342108.40
Request Latency (ms)2254.0718553.5635613.8465558.20
Request Count (requests)50320256512
Request Throughput (Req/s)0.443.443.443.43
Output Token Throughput (Tok/s)44.31344.02344.42343.29

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)592.9711969.3441372.5795995.33
Request Latency (ms)2533.9336021.0065053.92120751.01
Request Count (requests)50320256512
Request Throughput (Req/s)0.391.751.741.75
Output Token Throughput (Tok/s)39.43174.83173.72174.85

B200

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)212.106902.5111412.3522707.11
Request Latency (ms)212.106902.5111412.3522707.11
Request Count (requests)50320256512
Request Throughput (Req/s)4.688.418.528.52
Output Token Throughput (Tok/s)4.688.418.528.52

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)350.3013909.7023275.7146780.29
Request Latency (ms)350.3013909.7023275.7146780.29
Request Count (requests)50320256510
Request Throughput (Req/s)2.844.164.164.10
Output Token Throughput (Tok/s)2.844.164.164.10

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)212.402723.575574.2716228.42
Request Latency (ms)1552.879594.9417572.4134293.88
Request Count (requests)50320256512
Request Throughput (Req/s)0.646.657.216.97
Output Token Throughput (Tok/s)64.30664.83721.29696.84

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)325.213821.2215872.0142120.75
Request Latency (ms)1686.8217970.1534042.2761444.78
Request Count (requests)50320256512
Request Throughput (Req/s)0.593.553.523.60
Output Token Throughput (Tok/s)59.21354.77352.10360.15

B300

Input 50 / Output 1 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)176.246665.8611233.6822238.90
Request Latency (ms)176.246665.8611233.6822238.90
Request Count (requests)50320256510
Request Throughput (Req/s)5.648.718.678.65
Output Token Throughput (Tok/s)5.648.718.678.65

Input 50 / Output 1 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)301.4913570.8322688.8745276.68
Request Latency (ms)301.4913570.8322688.8745276.68
Request Count (requests)50320255512
Request Throughput (Req/s)3.314.264.264.26
Output Token Throughput (Tok/s)3.314.264.264.26

Input 50 / Output 100 / Video 1 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)175.832491.564999.718916.09
Request Latency (ms)1492.469254.0017203.6033189.36
Request Count (requests)50320256512
Request Throughput (Req/s)0.676.897.377.61
Output Token Throughput (Tok/s)66.93689.29736.64761.12

Input 50 / Output 100 / Video 2 FPS

MetricConcurrency 1Concurrency 64Concurrency 128Concurrency 256
Time To First Token (ms)303.523655.158924.4930494.72
Request Latency (ms)1637.0217223.7833088.0862798.50
Request Count (requests)50320256512
Request Throughput (Req/s)0.613.703.823.83
Output Token Throughput (Tok/s)61.02370.09382.18383.31

Notes:

  1. Source: vLLM inference benchmarking for nvidia/Cosmos3-Super; AIPerf client was used as the benchmarking tool.
  2. Hardware: results are grouped by GPU product (RTX PRO 6000 Blackwell, H20, H100 NVL, H200 NVL, H200 141GB HBM3, B200, B300). All metrics are averages for a number of requests.
  3. Time To First Token (TTFT) measures latency until the first output token is emitted. Request Latency is end-to-end time per request. For single-token outputs (Output 1), TTFT and request latency are identical.
  4. Request Throughput is completed requests per second. Output Token Throughput is generated tokens per second (for Output 1 workloads, the two throughputs match).
  5. Concurrency is the number of simultaneous client requests issued by AIPerf, not tensor-parallel GPU count.