Model Card

February 7, 2023 · View on GitHub

This page lists the MaskAlign model weights. CLIP-L/14* denotes input 196 × 196 resolution image to CLIP-L/14. This will keep the same feature map size as the student model. PT epochs and FT Acc denotes pre-training epochs and fine-tuning accuracy on ImageNet-1K, respectively.

ModelTeacher ModelPT epochsLinkFT Acc.
ViT-B/16CLIP-B/16200gdrive85.4
ViT-L/16CLIP-B/16200gdrive86.5
ViT-L/16CLIP-L/14*200gdrive87.4