README_nvidia.md
February 27, 2024 ยท View on GitHub
[ Back to index ]
Build Nvidia Docker Container (from 3.1 Inference round)
cm docker script --tags=build,nvidia,inference,server
Run this benchmark via CM
Do a test run to detect and record the system performance
cmr "run-mlperf inference _find-performance" --scenario=Offline \
--model=bert-99 --implementation=nvidia-original --device=cuda --backend=tensorrt \
--category=edge --division=open --quiet
- Use
--model=bert-99.9to run the high-accuracy model (only for datacenter) - Use
--rerunto force a rerun even when result files (from a previous run) exist
Do full accuracy and performance runs for all the scenarios
cmr "run-mlperf inference _submission _all-scenarios" --model=bert-99 \
--device=cuda --implementation=nvidia-original --backend=tensorrt \
--execution-mode=valid --category=edge --division=open --quiet
- Use
--category=datacenterto run datacenter scenarios (only for bert-99.9) - Use
--power=yesfor measuring power. It is ignored for accuracy and compliance runs - Use
--division=closedto run all scenarios for the closed division including the compliance tests --offline_target_qps,--server_target_qps, and--singlestream_target_latencycan be used to pass in the performance numbers
Generate and upload MLPerf submission
Follow this guide to generate the submission tree and upload your results.
Questions? Suggestions?
Check the MLCommons Task Force on Automation and Reproducibility and get in touch via public Discord server.
Acknowledgments
- CM automation for Nvidia's MLPerf inference implementation was developed by Arjun Suresh and Grigori Fursin.
- Nvidia's MLPerf inference implementation was developed by Zhihan Jiang, Ethan Cheng, Yiheng Zhang and Jinho Suh.