Iwin Transformer    

April 17, 2026 · View on GitHub

Iwin Transformer

Introduction

Iwin Transformer (the name Iwin stands for Interleaved window) is initially described in arxiv. It is a position-embedding-free hierarchical vision transformer, which can be fine-tuned directly from low to high resolution, through the collaboration of innovative interleaved window attention and depthwise convolution. teaser teaser teaser teaser

Results on ImageNet with Pretrained Models

ImageNet-1K and ImageNet-22K Pretrained Iwin Models

namepretrainresolutionacc@1#paramsFLOPs22K model1K model
Iwin-TImageNet-1K224x22482.030.2M4.7G-github/config
Iwin-SImageNet-1K224x22483.451.6M9.0G-github/config
Iwin-SImageNet-1K384x38484.351.6M27.7G-github/config
Iwin-SImageNet-1K512x51284.451.6M52.0G-github/config
Iwin-SImageNet-1K1024x102483.851.6M207.9G-github/config
Iwin-BImageNet-1K224x22483.591.2M15.9G-github/config
Iwin-BImageNet-1K384x38484.991.2M48.3G-github/config
Iwin-BImageNet-1K512x51285.191.3M89.5G-github/config
Iwin-BImageNet-1K1024x102485.091.3M358.2G-github/config
Iwin-BImageNet-22K224x22485.591.2M15.9Ggithub/configgithub/config
Iwin-BImageNet-22K384x38486.691.2M48.3G-github/config
Iwin-BImageNet-22K512x51286.191.2M89.5G-github/config
Iwin-BImageNet-22K1024x102485.691.2M358.2G-github/config
Iwin-LImageNet-22K224x22486.4204.3M35.4Ggithub/configgithub/config
Iwin-LImageNet-22K384x38487.4204.3M106.6G-github/config

Results on Downstream Tasks

COCO Object Detection (2017 val)

BackboneMethodpretrainLr Schdbox mAPmask mAP#paramsFLOPsmodel
Iwin-TMask R-CNNImageNet-1K1x42.238.948M268Ggithub
Iwin-SMask R-CNNImageNet-1K1x43.740.069M358Ggithub
Iwin-TMask R-CNNImageNet-1K3x44.740.948M268Ggithub
Iwin-SMask R-CNNImageNet-1K3x45.541.069M358Ggithub
Iwin-TCascade Mask R-CNNImageNet-1K1x47.240.986M747Ggithub
Iwin-TCascade Mask R-CNNImageNet-1K3x49.442.986M747Ggithub
Iwin-SCascade Mask R-CNNImageNet-1K3x49.443.0107M837Ggithub

ADE20K Semantic Segmentation (val)

BackboneMethodpretrainCrop SizeLr SchdmIoU#paramsFLOPsmodel
Iwin-TUPerNetImageNet-1K512x512160K44.7061.9M946Ggithub
Iwin-SUperNetImageNet-1K512x512160K47.5083.2M1038Ggithub
Iwin-BUperNetImageNet-1K512x512160K48.90124.8M1189Ggithub

Kinetics 400 Recognition

Kinetics-400 Video Recognition

BackbonePretrainLr Schdspatial cropacc@1acc@5#paramsFLOPsconfigmodel
Iwin-TImageNet-1K30ep22479.193.829.8M74Gconfiggithub
Iwin-SImageNet-1K30ep22480.094.151.1M140Gconfiggithub

Image generation

We built an image generation model FlashDiT without position embedding using iwin 2D attention. It has less memory usage and faster training speed than Lightningdit.

TokenizerGeneration ModelFIDFID cfg
VA-VAEFlashDiT-win8-XL-56ep5.372.30

Language Model

We built a language model "miniwin" based on the iwin's philosophy "no token is an island." See miniwin for more details.

Acknowledgements

This repo is mainly built on Swin. Thanks for the great work.

License

MIT.

Citation

If you find our work helpful, feel free to give us a cite and a star🌟.

@misc{huo2025iwin,
      title={Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows}, 
      author={Simin Huo and Ning Li},
      year={2025},
      eprint={2507.18405},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.18405}, 
}