Overview of OpenVINO™ Toolkit Public Pre-Trained Models

August 6, 2024 · View on GitHub

:maxdepth: 1
:hidden:
:caption: Device Support

   Public Pre-Trained Models Device Support <omz_models_public_device_support>
   aclnet <omz_models_model_aclnet>
   aclnet-int8 <omz_models_model_aclnet_int8>
   anti-spoof-mn3 <omz_models_model_anti_spoof_mn3>
   background-matting-mobilenetv2 <omz_models_model_background_matting_mobilenetv2>
   bert-base-ner <omz_models_model_bert_base_ner>
   brain-tumor-segmentation-0002 <omz_models_model_brain_tumor_segmentation_0002>
   cocosnet <omz_models_model_cocosnet>
   colorization-siggraph <omz_models_model_colorization_siggraph>
   colorization-v2 <omz_models_model_colorization_v2>
   common-sign-language-0001 <omz_models_model_common_sign_language_0001>
   convnext-tiny <omz_models_model_convnext_tiny>
   ctdet_coco_dlav0_512 <omz_models_model_ctdet_coco_dlav0_512>
   ctpn <omz_models_model_ctpn>
   deeplabv3 <omz_models_model_deeplabv3>
   densenet-121-tf <omz_models_model_densenet_121_tf>
   detr-resnet50 <omz_models_model_detr_resnet50>
   dla-34 <omz_models_model_dla_34>
   drn-d-38 <omz_models_model_drn_d_38>
   efficientdet-d0-tf <omz_models_model_efficientdet_d0_tf>
   efficientdet-d1-tf <omz_models_model_efficientdet_d1_tf>
   efficientnet-b0 <omz_models_model_efficientnet_b0>
   efficientnet-b0-pytorch <omz_models_model_efficientnet_b0_pytorch>
   efficientnet-v2-b0 <omz_models_model_efficientnet_v2_b0>
   efficientnet-v2-s <omz_models_model_efficientnet_v2_s>
   erfnet <omz_models_model_erfnet>
   f3net <omz_models_model_f3net>
   face-recognition-resnet100-arcface-onnx <omz_models_model_face_recognition_resnet100_arcface_onnx>
   faceboxes-pytorch <omz_models_model_faceboxes_pytorch>
   facenet-20180408-102900 <omz_models_model_facenet_20180408_102900>
   fast-neural-style-mosaic-onnx <omz_models_model_fast_neural_style_mosaic_onnx>
   faster_rcnn_inception_resnet_v2_atrous_coco <omz_models_model_faster_rcnn_inception_resnet_v2_atrous_coco>
   faster_rcnn_resnet50_coco <omz_models_model_faster_rcnn_resnet50_coco>
   fastseg-large <omz_models_model_fastseg_large>
   fastseg-small <omz_models_model_fastseg_small>
   fbcnn <omz_models_model_fbcnn>
   fcrn-dp-nyu-depth-v2-tf <omz_models_model_fcrn_dp_nyu_depth_v2_tf>
   forward-tacotron (composite) <omz_models_model_forward_tacotron>
   gmcnn-places2-tf <omz_models_model_gmcnn_places2_tf>
   googlenet-v1-tf <omz_models_model_googlenet_v1_tf>
   googlenet-v2-tf <omz_models_model_googlenet_v2_tf>
   googlenet-v3 <omz_models_model_googlenet_v3>
   googlenet-v3-pytorch <omz_models_model_googlenet_v3_pytorch>
   googlenet-v4-tf <omz_models_model_googlenet_v4_tf>
   gpt-2 <omz_models_model_gpt_2>
   hbonet-0.25 <omz_models_model_hbonet_0_25>
   hbonet-1.0 <omz_models_model_hbonet_1_0>
   higher-hrnet-w32-human-pose-estimation <omz_models_model_higher_hrnet_w32_human_pose_estimation>
   hrnet-v2-c1-segmentation <omz_models_model_hrnet_v2_c1_segmentation>
   human-pose-estimation-3d-0001 <omz_models_model_human_pose_estimation_3d_0001>
   hybrid-cs-model-mri <omz_models_model_hybrid_cs_model_mri>
   i3d-rgb-tf <omz_models_model_i3d_rgb_tf>
   inception-resnet-v2-tf <omz_models_model_inception_resnet_v2_tf>
   levit-128s <omz_models_model_levit_128s>
   license-plate-recognition-barrier-0007 <omz_models_model_license_plate_recognition_barrier_0007>
   mask_rcnn_inception_resnet_v2_atrous_coco <omz_models_model_mask_rcnn_inception_resnet_v2_atrous_coco>
   mask_rcnn_resnet50_atrous_coco <omz_models_model_mask_rcnn_resnet50_atrous_coco>
   midasnet <omz_models_model_midasnet>
   mixnet-l <omz_models_model_mixnet_l>
   mobilenet-v1-0.25-128 <omz_models_model_mobilenet_v1_0_25_128>
   mobilenet-v1-1.0-224-tf <omz_models_model_mobilenet_v1_1_0_224_tf>
   mobilenet-v2-1.0-224 <omz_models_model_mobilenet_v2_1_0_224>
   mobilenet-v2-1.4-224 <omz_models_model_mobilenet_v2_1_4_224>
   mobilenet-v2-pytorch <omz_models_model_mobilenet_v2_pytorch>
   mobilenet-v3-large-1.0-224-tf <omz_models_model_mobilenet_v3_large_1_0_224_tf>
   mobilenet-v3-small-1.0-224-tf <omz_models_model_mobilenet_v3_small_1_0_224_tf>
   mobilenet-yolo-v4-syg <omz_models_model_mobilenet_yolo_v4_syg>
   modnet-photographic-portrait-matting <omz_models_model_modnet_photographic_portrait_matting>
   modnet-webcam-portrait-matting <omz_models_model_modnet_webcam_portrait_matting>
   mozilla-deepspeech-0.6.1 <omz_models_model_mozilla_deepspeech_0_6_1>
   mozilla-deepspeech-0.8.2 <omz_models_model_mozilla_deepspeech_0_8_2>
   nanodet-m-1.5x-416 <omz_models_model_nanodet_m_1_5x_416>
   nanodet-plus-m-1.5x-416 <omz_models_model_nanodet_plus_m_1_5x_416>
   netvlad-tf <omz_models_model_netvlad_tf>
   nfnet-f0 <omz_models_model_nfnet_f0>
   open-closed-eye-0001 <omz_models_model_open_closed_eye_0001>
   pspnet-pytorch <omz_models_model_pspnet_pytorch>
   quartznet-15x5-en <omz_models_model_quartznet_15x5_en>
   regnetx-3.2gf <omz_models_model_regnetx_3_2gf>
   repvgg-a0 <omz_models_model_repvgg_a0>
   repvgg-b1 <omz_models_model_repvgg_b1>
   repvgg-b3 <omz_models_model_repvgg_b3>
   resnest-50-pytorch <omz_models_model_resnest_50_pytorch>
   resnet-18-pytorch <omz_models_model_resnet_18_pytorch>
   resnet-34-pytorch <omz_models_model_resnet_34_pytorch>
   resnet-50-pytorch <omz_models_model_resnet_50_pytorch>
   resnet-50-tf <omz_models_model_resnet_50_tf>
   retinaface-resnet50-pytorch <omz_models_model_retinaface_resnet50_pytorch>
   retinanet-tf <omz_models_model_retinanet_tf>
   rexnet-v1-x1.0 <omz_models_model_rexnet_v1_x1_0>
   rfcn-resnet101-coco-tf <omz_models_model_rfcn_resnet101_coco_tf>
   robust-video-matting-mobilenetv3 <omz_models_model_robust_video_matting_mobilenetv3>
   shufflenet-v2-x1.0 <omz_models_model_shufflenet_v2_x1_0>
   single-human-pose-estimation-0001 <omz_models_model_single_human_pose_estimation_0001>
   ssd_mobilenet_v1_coco <omz_models_model_ssd_mobilenet_v1_coco>
   ssd_mobilenet_v1_fpn_coco <omz_models_model_ssd_mobilenet_v1_fpn_coco>
   ssd-resnet34-1200-onnx <omz_models_model_ssd_resnet34_1200_onnx>
   ssdlite_mobilenet_v2 <omz_models_model_ssdlite_mobilenet_v2>
   swin-tiny-patch4-window7-224 <omz_models_model_swin_tiny_patch4_window7_224>
   t2t-vit-14 <omz_models_model_t2t_vit_14>
   text-recognition-resnet-fc <omz_models_model_text_recognition_resnet_fc>
   ultra-lightweight-face-detection-rfb-320 <omz_models_model_ultra_lightweight_face_detection_rfb_320>
   ultra-lightweight-face-detection-slim-320 <omz_models_model_ultra_lightweight_face_detection_slim_320>
   vehicle-license-plate-detection-barrier-0123 <omz_models_model_vehicle_license_plate_detection_barrier_0123>
   vehicle-reid-0001 <omz_models_model_vehicle_reid_0001>
   vitstr-small-patch16-224 <omz_models_model_vitstr_small_patch16_224>
   wav2vec2-base <omz_models_model_wav2vec2_base>
   wavernn (composite) <omz_models_model_wavernn>
   yolact-resnet50-fpn-pytorch <omz_models_model_yolact_resnet50_fpn_pytorch>
   yolo-v1-tiny-tf <omz_models_model_yolo_v1_tiny_tf>
   yolo-v2-tf <omz_models_model_yolo_v2_tf>
   yolo-v2-tiny-tf <omz_models_model_yolo_v2_tiny_tf>
   yolo-v3-onnx <omz_models_model_yolo_v3_onnx>
   yolo-v3-tf <omz_models_model_yolo_v3_tf>
   yolo-v3-tiny-onnx <omz_models_model_yolo_v3_tiny_onnx>
   yolo-v3-tiny-tf <omz_models_model_yolo_v3_tiny_tf>
   yolo-v4-tf <omz_models_model_yolo_v4_tf>
   yolo-v4-tiny-tf <omz_models_model_yolo_v4_tiny_tf>
   yolof <omz_models_model_yolof>
   yolox-tiny <omz_models_model_yolox_tiny>

OpenVINO™ toolkit provides a set of public pre-trained models that you can use for learning and demo purposes or for developing deep learning software. Most recent version is available in the repo on Github. The table Public Pre-Trained Models Device Support summarizes devices supported by each model.

You can download models and convert them into OpenVINO™ IR format (*.xml + *.bin) using the OpenVINO™ Model Downloader and other automation tools.

Classification Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
AntiSpoofNetPyTorch*anti-spoof-mn33.81%0.153.02
ConvNeXt TinyPyTorch*convnext-tiny82.05%/95.86%8.941928.5892
DenseNet 121densenet-121-tf74.46%/92.13%5.723~5.72877.971
DLA 34PyTorch*dla-3474.64%/92.06%6.136815.7344
EfficientNet B0TensorFlow*
PyTorch*
efficientnet-b0
efficientnet-b0-pytorch
75.70%/92.76%
77.70%/93.52%
0.8195.268
EfficientNet V2 B0PyTorch*efficientnet-v2-b078.36%/94.02%1.46417.1094
EfficientNet V2 SmallPyTorch*efficientnet-v2-s84.29%/97.26%16.940621.3816
HBONet 1.0PyTorch*hbonet-1.073.1%/91.0%0.62084.5443
HBONet 0.25PyTorch*hbonet-0.2557.3%/79.8%0.07581.9299
Inception (GoogleNet) V1TensorFlow*googlenet-v1-tf69.814%/89.6%3.016~3.2666.619~6.999
Inception (GoogleNet) V2TensorFlow*googlenet-v2-tf74.084%/91.798%4.05811.185
Inception (GoogleNet) V3TensorFlow*
PyTorch*
googlenet-v3
googlenet-v3-pytorch
77.904%/93.808%
77.69%/93.7%
11.46923.817
Inception (GoogleNet) V4TensorFlow*googlenet-v4-tf80.204%/95.21%24.58442.648
Inception-ResNet V2TensorFlow*inception-resnet-v2-tf77.82%/94.03%22.22730.223
LeViT 128SPyTorch*levit-128s76.54%/92.85%0.61778.2199
MixNet LTensorFlow*mixnet-l78.30%/93.91%0.5657.3
MobileNet V1 0.25 128Caffe*mobilenet-v1-0.25-12840.54%/65%0.0280.468
MobileNet V1 1.0 224Caffe*
TensorFlow*
mobilenet-v1-1.0-224-tf71.03%/89.94%1.1484.221
MobileNet V2 1.0 224TensorFlow*
PyTorch*
mobilenet-v2-1.0-224
mobilenet-v2-pytorch
71.85%/90.69%
71.81%/90.396%
0.615~0.8763.489
MobileNet V2 1.4 224TensorFlow*mobilenet-v2-1.4-22474.09%/91.97%1.1836.087
MobileNet V3 Small 1.0TensorFlow*mobilenet-v3-small-1.0-224-tf67.36%/87.44%0.11682.537
MobileNet V3 Large 1.0TensorFlow*mobilenet-v3-large-1.0-224-tf75.30%/92.62%0.44505.4721
NFNet F0PyTorch*nfnet-f083.34%/96.56%24.805371.4444
RegNetX-3.2GFPyTorch*regnetx-3.2gf78.17%/94.08%6.389315.2653
open-closed-eye-0001PyTorch*open-closed-eye-000195.84%0.00140.0113
RepVGG A0PyTorch*repvgg-a072.40%/90.49%2.72868.3094
RepVGG B1PyTorch*repvgg-b178.37%/94.09%23.647251.8295
RepVGG B3PyTorch*repvgg-b380.50%/95.25%52.4407110.9609
ResNeSt 50PyTorch*resnest-50-pytorch81.11%/95.36%10.814827.4493
ResNet 18PyTorch*resnet-18-pytorch69.754%/89.088%3.63711.68
ResNet 34PyTorch*resnet-34-pytorch73.30%/91.42%7.340921.7892
ResNet 50PyTorch*
TensorFlow*
resnet-50-pytorchresnet-50-tf75.168%/92.212%
76.38%/93.188%
76.17%/92.98%
6.996~8.21625.53
ReXNet V1 x1.0PyTorch*rexnet-v1-x1.077.86%/93.87%0.83254.7779
Shufflenet V2 x1.0PyTorch*shufflenet-v2-x1.069.36%/88.32%0.29572.2705
Swin Transformer Tiny, window size=7PyTorch*swin-tiny-patch4-window7-22481.38%/95.51%9.028028.8173
T2T-ViT, transformer layers number=14PyTorch*t2t-vit-1481.44%/95.66%9.545121.5498

Segmentation Models

Semantic segmentation is an extension of object detection problem. Instead of returning bounding boxes, semantic segmentation models return a "painted" version of the input image, where the "color" of each pixel represents a certain class. These networks are much bigger than respective object detection networks, but they provide a better (pixel-level) localization of objects and they can detect areas with complex shape.

Semantic Segmentation Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
DeepLab V3TensorFlow*deeplabv368.41%11.46923.819
DRN-D-38PyTorch*drn-d-3871.31%1768.327625.9939
ErfnetPyTorch*erfnet76.47%11.137.87
HRNet V2 C1 SegmentationPyTorch*hrnet-v2-c1-segmentation77.69%81.99366.4768
Fastseg MobileV3Large LR-ASPP, F=128PyTorch*fastseg-large72.67%140.96113.2
Fastseg MobileV3Small LR-ASPP, F=128PyTorch*fastseg-small67.15%69.22041.1
PSPNet R-50-D8PyTorch*pspnet-pytorch70.6%357.171946.5827

Instance Segmentation Models

Instance segmentation is an extension of object detection and semantic segmentation problems. Instead of predicting a bounding box around each object instance instance segmentation model outputs pixel-wise masks for all instances.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
Mask R-CNN Inception ResNet V2TensorFlow*mask_rcnn_inception_resnet_v2_atrous_coco39.86%/35.36%675.31492.368
Mask R-CNN ResNet 50TensorFlow*mask_rcnn_resnet50_atrous_coco29.75%/27.46%294.73850.222
YOLACT ResNet 50 FPNPyTorch*yolact-resnet50-fpn-pytorch28.0%/30.69%118.57536.829

3D Semantic Segmentation Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
Brain Tumor Segmentation 2PyTorch*brain-tumor-segmentation-000291.4826%300.8014.51

Object Detection Models

Several detection models can be used to detect a set of the most popular objects - for example, faces, people, vehicles. Most of the networks are SSD-based and provide reasonable accuracy/performance trade-offs.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
CTPNTensorFlow*ctpn73.67%55.81317.237
CenterNet (CTDET with DLAV0) 512x512ONNX*ctdet_coco_dlav0_51244.2756%62.21117.911
DETR-ResNet50PyTorch*detr-resnet5039.27% / 42.36%174.470841.3293
EfficientDet-D0TensorFlow*efficientdet-d0-tf31.95%2.543.9
EfficientDet-D1TensorFlow*efficientdet-d1-tf37.54%6.16.6
FaceBoxesPyTorch*faceboxes-pytorch83.565%1.89751.0059
Faster R-CNN with Inception-ResNet v2TensorFlow*faster_rcnn_inception_resnet_v2_atrous_coco40.69%30.68713.307
Faster R-CNN with ResNet 50TensorFlow*faster_rcnn_resnet50_coco31.09%57.20329.162
Mobilenet-yolo-v4-sygKeras*mobilenet-yolo-v4-syg86.35%65.98161.922
NanoDet with ShuffleNetV2 1.5x, size=416PyTorch*nanodet-m-1.5x-41627.38%/26.63%2.38952.0534
NanoDet Plus with ShuffleNetV2 1.5x, size=416PyTorch*nanodet-plus-m-1.5x-41634.53%/33.77%3.01472.4614
RetinaFace with ResNet 50PyTorch*retinaface-resnet50-pytorch91.78%88.862727.2646
RetinaNet with Resnet 50TensorFlow*retinanet-tf33.15%238.946964.9706
R-FCN with Resnet-101TensorFlow*rfcn-resnet101-coco-tf28.40%/45.02%53.462171.85
SSD with MobileNetTensorFlow*ssd_mobilenet_v1_coco23.32%2.316~2.4945.783~6.807
SSD with MobileNet FPNTensorFlow*ssd_mobilenet_v1_fpn_coco35.5453%123.30936.188
SSD lite with MobileNet V2TensorFlow*ssdlite_mobilenet_v224.2946%1.5254.475
SSD with ResNet 34 1200x1200PyTorch*ssd-resnet34-1200-onnx20.7198%/39.2752%433.41120.058
Ultra Lightweight Face Detection RFB 320PyTorch*ultra-lightweight-face-detection-rfb-32084.78%0.21060.3004
Ultra Lightweight Face Detection slim 320PyTorch*ultra-lightweight-face-detection-slim-32083.32%0.17240.2844
Vehicle License Plate Detection BarrierTensorFlow*vehicle-license-plate-detection-barrier-012399.52%0.2710.547
YOLO v1 TinyTensorFlow.js*yolo-v1-tiny-tf54.79%6.988315.8587
YOLO v2 TinyKeras*yolo-v2-tiny-tf27.3443%/29.1184%5.423611.2295
YOLO v2Keras*yolo-v2-tf53.1453%/56.483%63.030150.9526
YOLO v3Keras*
ONNX*
yolo-v3-tf
yolo-v3-onnx
62.2759%/67.7221%
48.30%/47.07%
65.9843~65.99861.9221~61.930
YOLO v3 TinyKeras*
ONNX*
yolo-v3-tiny-tf
yolo-v3-tiny-onnx
35.9%/39.7%
17.07%/13.64%
5.5828.848~8.8509
YOLO v4Keras*yolo-v4-tf71.23%/77.40%/50.26%129.556764.33
YOLO v4 TinyKeras*yolo-v4-tiny-tf6.92896.0535
YOLOFPyTorch*yolof60.69%/66.23%/43.63%175.3794248.228
YOLOX TinyPyTorch*yolox-tiny47.85%/52.56%/31.82%6.48135.0472

Face Recognition Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
FaceNetTensorFlow*facenet-20180408-10290099.14%2.84623.469
LResNet100E-IR,ArcFace@ms1m-refine-v2MXNet*face-recognition-resnet100-arcface-onnx99.68%24.211565.1320

Human Pose Estimation Models

Human pose estimation task is to predict a pose: body skeleton, which consists of keypoints and connections between them, for every person in an input image or video. Keypoints are body joints, i.e. ears, eyes, nose, shoulders, knees, etc. There are two major groups of such methods: top-down and bottom-up. The first detects persons in a given frame, crops or rescales detections, then runs pose estimation network for every detection. These methods are very accurate. The second finds all keypoints in a given frame, then groups them by person instances, thus faster than previous, because network runs once.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
human-pose-estimation-3d-0001PyTorch*human-pose-estimation-3d-0001100.44437mm18.9985.074
single-human-pose-estimation-0001PyTorch*single-human-pose-estimation-000169.0491%60.12533.165
higher-hrnet-w32-human-pose-estimationPyTorch*higher-hrnet-w32-human-pose-estimation64.64%92.836428.6180

Monocular Depth Estimation Models

The task of monocular depth estimation is to predict a depth (or inverse depth) map based on a single input image. Since this task contains - in the general setting - some ambiguity, the resulting depth maps are often only defined up to an unknown scaling factor.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
midasnetPyTorch*midasnet0.07071207.25144104.081
FCRN ResNet50-UpprojTensorFlow*fcrn-dp-nyu-depth-v2-tf0.57363.542134.5255

Image Inpainting Models

Image inpainting task is to estimate suitable pixel information to fill holes in images.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
GMCNN InpaintingTensorFlow*gmcnn-places2-tf33.47Db691.158912.7773
Hybrid-CS-Model-MRITensorFlow*hybrid-cs-model-mri34.27Db146.603711.3313

Style Transfer Models

Style transfer task is to transfer the style of one image to another.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
fast-neural-style-mosaic-onnxONNX*fast-neural-style-mosaic-onnx12.04dB15.5181.679

Action Recognition Models

The task of action recognition is to predict action that is being performed on a short video clip (tensor formed by stacking sampled frames from input video).

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
RGB-I3D, pretrained on ImageNet*TensorFlow*i3d-rgb-tf64.83%/84.58%278.981512.6900
common-sign-language-0001PyTorch*common-sign-language-000193.58%4.22694.1128

Colorization Models

Colorization task is to predict colors of scene from grayscale image.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
colorization-v2PyTorch*colorization-v226.99dB83.604532.2360
colorization-siggraphPyTorch*colorization-siggraph27.73dB150.544134.0511

Sound Classification Models

The task of sound classification is to predict what sounds are in an audio fragment.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
ACLNetPyTorch*aclnet86%/92%1.42.7
ACLNet-int8PyTorch*aclnet-int887%/93%1.412.71

Speech Recognition Models

The task of speech recognition is to recognize and translate spoken language into text.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
DeepSpeech V0.6.1TensorFlow*mozilla-deepspeech-0.6.17.55%0.047247.2
DeepSpeech V0.8.2TensorFlow*mozilla-deepspeech-0.8.26.13%0.047247.2
QuartzNetPyTorch*quartznet-15x5-en3.86%2.419518.8857
Wav2Vec 2.0 BasePyTorch*wav2vec2-base3.39%26.84394.3965

Image Translation Models

The task of image translation is to generate the output based on exemplar.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
CoCosNetPyTorch*cocosnet12.93dB1080.7032167.9141

Optical Character Recognition Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
license-plate-recognition-barrier-0007TensorFlow*license-plate-recognition-barrier-000798%0.3471.435

Place Recognition Models

The task of place recognition is to quickly and accurately recognize the location of a given query photograph.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
NetVLADTensorFlow*netvlad-tf82.0321%36.6374149.0021

JPEG Artifacts Removal Models

The task of restoration images from jpeg format.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
FBCNNPyTorch*fbcnn34.34Db1420.7823571.922

Salient Object Detection Models

Salient object detection is a task-based on a visual attention mechanism, in which algorithms aim to explore objects or regions more attentive than the surrounding areas on the scene or images.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
F3NetPyTorch*f3net84.21%31.288325.2791

Text Prediction Models

Text prediction is a task to predict the next word, given all of the previous words within some text.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
GPT-2PyTorch*gpt-229.00%293.0489175.6203

Text Recognition Models

Scene text recognition is a task to recognize text on a given image. Researchers compete on creating algorithms which are able to recognize text of different shapes, fonts and background. See details about datasets in here The reported metric is collected over the alphanumeric subset of ICDAR13 (1015 images) in case-insensitive mode.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
Resnet-FCPyTorch*text-recognition-resnet-fc90.94%40.3704177.9668
ViTSTR Small patch=16, size=224PyTorch*vitstr-small-patch16-22490.34%9.154421.5061

Text to Speech Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
ForwardTacotronPyTorch*forward-tacotron:
forward-tacotron-duration-prediction
forward-tacotron-regression

6.66
4.91

13.81
3.05
WaveRNNPyTorch*wavernn:
wavernn-upsampler
wavernn-rnn

0.37
0.06

0.4
3.83

Named Entity Recognition Models

Named entity recognition (NER) is the task of tagging entities in text with their corresponding type.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
bert-base-NERPyTorch*bert-base-ner94.45%22.3874107.4319

Vehicle Reidentification Models

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
vehicle-reid-0001PyTorch*vehicle-reid-000196.31%/85.15 %2.6432.183

Background Matting Models

Background matting is a method of separating a foreground from a background in an image or video, wherein some pixels may belong to foreground as well as background, such pixels are called partial or mixed pixels. This distinguishes background matting from segmentation approaches where the result is a binary mask.

Model NameImplementationOMZ Model NameAccuracyGFlopsmParams
background-matting-mobilenetv2PyTorch*background-matting-mobilenetv24.32/1.0/2.48/2.76.74195.052
modnet-photographic-portrait-mattingPyTorch*modnet-photographic-portrait-matting5.21/727.9531.15646.4597
modnet-webcam-portrait-mattingPyTorch*modnet-webcam-portrait-matting5.66/762.5231.15646.4597
robust-video-matting-mobilenetv3PyTorch*robust-video-matting-mobilenetv320.8/15.1/4.42/4.059.38923.7363

See Also

Caffe, Caffe2, Keras, MXNet, PyTorch, and TensorFlow are trademarks or brand names of their respective owners. All company, product and service names used in this website are for identification purposes only. Use of these names,trademarks and brands does not imply endorsement.