Iwin Transformer    
April 17, 2026 · View on GitHub
Iwin Transformer
Introduction
Iwin Transformer (the name Iwin stands for Interleaved window) is initially described in arxiv. It is a position-embedding-free hierarchical vision transformer, which can be fine-tuned directly from low to high resolution, through the collaboration of innovative interleaved window attention and depthwise convolution.

Results on ImageNet with Pretrained Models
ImageNet-1K and ImageNet-22K Pretrained Iwin Models
| name | pretrain | resolution | acc@1 | #params | FLOPs | 22K model | 1K model |
|---|---|---|---|---|---|---|---|
| Iwin-T | ImageNet-1K | 224x224 | 82.0 | 30.2M | 4.7G | - | github/config |
| Iwin-S | ImageNet-1K | 224x224 | 83.4 | 51.6M | 9.0G | - | github/config |
| Iwin-S | ImageNet-1K | 384x384 | 84.3 | 51.6M | 27.7G | - | github/config |
| Iwin-S | ImageNet-1K | 512x512 | 84.4 | 51.6M | 52.0G | - | github/config |
| Iwin-S | ImageNet-1K | 1024x1024 | 83.8 | 51.6M | 207.9G | - | github/config |
| Iwin-B | ImageNet-1K | 224x224 | 83.5 | 91.2M | 15.9G | - | github/config |
| Iwin-B | ImageNet-1K | 384x384 | 84.9 | 91.2M | 48.3G | - | github/config |
| Iwin-B | ImageNet-1K | 512x512 | 85.1 | 91.3M | 89.5G | - | github/config |
| Iwin-B | ImageNet-1K | 1024x1024 | 85.0 | 91.3M | 358.2G | - | github/config |
| Iwin-B | ImageNet-22K | 224x224 | 85.5 | 91.2M | 15.9G | github/config | github/config |
| Iwin-B | ImageNet-22K | 384x384 | 86.6 | 91.2M | 48.3G | - | github/config |
| Iwin-B | ImageNet-22K | 512x512 | 86.1 | 91.2M | 89.5G | - | github/config |
| Iwin-B | ImageNet-22K | 1024x1024 | 85.6 | 91.2M | 358.2G | - | github/config |
| Iwin-L | ImageNet-22K | 224x224 | 86.4 | 204.3M | 35.4G | github/config | github/config |
| Iwin-L | ImageNet-22K | 384x384 | 87.4 | 204.3M | 106.6G | - | github/config |
Results on Downstream Tasks
COCO Object Detection (2017 val)
| Backbone | Method | pretrain | Lr Schd | box mAP | mask mAP | #params | FLOPs | model |
|---|---|---|---|---|---|---|---|---|
| Iwin-T | Mask R-CNN | ImageNet-1K | 1x | 42.2 | 38.9 | 48M | 268G | github |
| Iwin-S | Mask R-CNN | ImageNet-1K | 1x | 43.7 | 40.0 | 69M | 358G | github |
| Iwin-T | Mask R-CNN | ImageNet-1K | 3x | 44.7 | 40.9 | 48M | 268G | github |
| Iwin-S | Mask R-CNN | ImageNet-1K | 3x | 45.5 | 41.0 | 69M | 358G | github |
| Iwin-T | Cascade Mask R-CNN | ImageNet-1K | 1x | 47.2 | 40.9 | 86M | 747G | github |
| Iwin-T | Cascade Mask R-CNN | ImageNet-1K | 3x | 49.4 | 42.9 | 86M | 747G | github |
| Iwin-S | Cascade Mask R-CNN | ImageNet-1K | 3x | 49.4 | 43.0 | 107M | 837G | github |
ADE20K Semantic Segmentation (val)
| Backbone | Method | pretrain | Crop Size | Lr Schd | mIoU | #params | FLOPs | model |
|---|---|---|---|---|---|---|---|---|
| Iwin-T | UPerNet | ImageNet-1K | 512x512 | 160K | 44.70 | 61.9M | 946G | github |
| Iwin-S | UperNet | ImageNet-1K | 512x512 | 160K | 47.50 | 83.2M | 1038G | github |
| Iwin-B | UperNet | ImageNet-1K | 512x512 | 160K | 48.90 | 124.8M | 1189G | github |
Kinetics 400 Recognition
Kinetics-400 Video Recognition
| Backbone | Pretrain | Lr Schd | spatial crop | acc@1 | acc@5 | #params | FLOPs | config | model |
|---|---|---|---|---|---|---|---|---|---|
| Iwin-T | ImageNet-1K | 30ep | 224 | 79.1 | 93.8 | 29.8M | 74G | config | github |
| Iwin-S | ImageNet-1K | 30ep | 224 | 80.0 | 94.1 | 51.1M | 140G | config | github |
Image generation
We built an image generation model FlashDiT without position embedding using iwin 2D attention. It has less memory usage and faster training speed than Lightningdit.
| Tokenizer | Generation Model | FID | FID cfg |
|---|---|---|---|
| VA-VAE | FlashDiT-win8-XL-56ep | 5.37 | 2.30 |
Language Model
We built a language model "miniwin" based on the iwin's philosophy "no token is an island." See miniwin for more details.
Acknowledgements
This repo is mainly built on Swin. Thanks for the great work.
License
MIT.
Citation
If you find our work helpful, feel free to give us a cite and a star🌟.
@misc{huo2025iwin,
title={Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows},
author={Simin Huo and Ning Li},
year={2025},
eprint={2507.18405},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.18405},
}