Compact Generalized Non-local Network

May 10, 2019 ยท View on GitHub

By Kaiyu Yue, Ming Sun, Yuchen Yuan, Feng Zhou, Errui Ding and Fuxin Xu

Introduction

This is a PyTorch re-implementation for the paper Compact Generalized Non-local Network. It brings the CGNL models trained on the CUB-200, ImageNet and COCO based on maskrcnn-benchmark from FAIR.

introfig

Update

Citation

If you think this code is useful in your research or wish to refer to the baseline results published in our paper, please use the following BibTeX entry.

@article{CGNLNetwork2018,
    author={Kaiyu Yue and Ming Sun and Yuchen Yuan and Feng Zhou and Errui Ding and Fuxin Xu},
    title={Compact Generalized Non-local Network},
    journal={NIPS},
    year={2018}
}

Requirements

  • PyTorch >= 0.4.1 or 1.0 from a nightly release
  • Python >= 3.5
  • torchvision >= 0.2.1
  • termcolor >= 1.1.0

Environment

The code is developed and tested under 8 Tesla P40 / V100-SXM2-16GB GPUS cards on CentOS with installed CUDA-9.2/8.0 and cuDNN-7.1.

Baselines and Main Results on CUB-200 Dataset

File IDModelBest Top-1 (%)Top-5 (%)Google DriveBaidu Pan
1832260500R-50 Base86.4597.00linklink
1832260501R-50 w/ 1 NL Block86.6996.95linklink
1832260502R-50 w/ 1 CGNL Block87.0696.91linklink
1832261010R-101 Base86.7696.91linklink
1832261011R-101 w/ 1 NL Block87.0497.01linklink
1832261012R-101 w/ 1 CGNL Block87.2897.20linklink

Notes:

  • The input size is 448.
  • The CGNL block with dot production kernel is configured within 8 groups.
File IDModelBest Top-1 (%)Top-5 (%)Google DriveBaidu Pan
1832260503xR-50 w/ 1 CGNLx Block86.5696.63linklink
1832261013xR-101 w/ 1 CGNLx Block87.1897.03linklink

Notes:

  • The input size is 448.
  • The CGNLx block with Gaussian RBF [0][1] kernel is configured within 8 groups.
  • The Taylor Expansion order for the kernel function is 3.

Experiments on ImageNet Dataset

File IDModelBest Top-1 (%)Top-5 (%)Google DriveBaidu Pan
torchvisionR-50 Base76.1592.87--
1832261502R-50 w/ 1 CGNL Block77.6993.63linklink
1832261503R-50 w/ 1 CGNLx Block77.3293.40linklink
torchvisionR-152 Base78.3194.06--
1832261522R-152 w/ 1 CGNL Block79.5394.52linklink
1832261523R-152 w/ 1 CGNLx Block79.3794.47linklink

Notes:

  • The input size is 224.
  • The CGNL and CGNLx blocks are configured as same as above experiments on CUB-200.

Experiments on COCO based on Mask R-CNN in PyTorch 1.0

backbonetypelr schedim / gputrain mem(GB)train time (s/iter)total train time(hr)inference time(s/im)box APmask APmodel idGoogle DriveBaidu Pan
R-50-C4Mask1x15.6410.543427.30.18329 + 0.01135.631.56358801--
R-50-C4 w/ 1 CGNL BlockMask1x15.8680.578528.50.20326 + 0.00836.332.1-linklink
R-50-C4 w/ 1 CGNLx BlockMask
s1x_C.SOLVER.WARMUP_ITERS = 20000
STEPS: (140000, 180000)
MAX_ITER: 200000
15.9770.585532.30.18571 + 0.01036.231.9-linklink

Notes:

  • The CGNL model is simply trained using the same experimental strategy as in maskrcnn-benchmark. It is configured as same as above experiments on CUB-200.
  • If you want to add the CGNL / CGNLx / NL blocks to the backbone of Mask-RCNN models, you can use the maskrcnn-benchmark/modeling/backbone/resnet.py and maskrcnn-benchmark/utils/c2_model_loading.py to replace the original py-files. Please refer to the code for specific configurations.
  • Prolonging the WARMUP_ITERS appropriately would produce the better results for CGNL models. The long training schedule is also recommended, like 2x or 1.44x in Detectron.
  • Due to some reasons of the Linux virtual environment or the data I/O speed, the numbers of train time, total train time and inference time in above table are both larger than the benchmarks. But this does not affect the demonstration of the efficiency of CGNL block.

Getting Start

Prepare Dataset

  • Download pytorch imagenet pretrained models from pytorch model zoo. The optional download links can be found in torchvision. Put them in the pretrained folder.

  • Download the training and validation lists for CUB-200 dataset from Google Drive or Baidu Pan. Download the ImageNet dataset and move validation images to labeled subfolders following the tutorial. The training and validation lists can be found in Google Drive or Baidu Pan. Put them in the data folder and make them look like:

    ${THIS REPO ROOT}
     `-- pretrained
         |-- resnet50-19c8e357.pth
         |-- resnet101-5d3b4d8f.pth
         |-- resnet152-b121ed2d.pth
     `-- data
         `-- cub
             `-- images
             |   |-- 001.Black_footed_Albatross
             |   |-- 002.Laysan_Albatross
             |   |-- ...
             |   |-- 200.Common_Yellowthroat
             |-- cub_train.list
             |-- cub_val.list
             |-- images.txt
             |-- image_class_labels.txt
             |-- README
         `-- imagenet
             `-- img_train
             |   |-- n01440764
             |   |-- n01734418
             |   |-- ...
             |   |-- n15075141
             `-- img_val
             |   |-- n01440764
             |   |-- n01734418
             |   |-- ...
             |   |-- n15075141
             |-- imagenet_train.list
             |-- imagenet_val.list
    

Perform Validating

$ python train_val.py --arch '50' --dataset 'cub' --nl-type 'cgnl' --nl-num 1 --checkpoints ${FOLDER_DIR} --valid

Perform Training Baselines

$ python train_val.py --arch '50' --dataset 'cub' --nl-num 0

Perform Training NL and CGNL Networks

$ python train_val.py --arch '50' --dataset 'cub' --nl-type 'cgnl' --nl-num 1 --warmup

Reference

License

This code is released under the MIT License. See LICENSE for additional details.