Attention Gym
September 7, 2026 Β· View on GitHub
Attention Gym is a collection of kernels, guides, and examples for FlexAttention and other novel attention variants.
π Docs | π― Features | π Getting Started | π» Usage | π οΈ Dev | π€ Contributing | βοΈ License
π Overview
Attention Gym began as a library of examples showing the many ways to express attention variants with the FlexAttention API. It is growing into a broader playground for attention, with the addition of sparse attention kernels, linear attention APIs for training and inference as well as showcasing how to use FlexAttention and friends in real workloads.

π― Features
- FlexAttention masks and score modifications
- Sparse attention patterns
- APIs and kernels for efficient GDN and KDA
- Utility functions for creating and combining attention masks
- Examples of how to use FlexAttention in real-world scenarios
π Getting Started
Prerequisites
- PyTorch (version 2.5 or higher)
Installation
Install the official wheel from PyPI:
pip install attn-gym
The base package intentionally keeps its runtime dependency surface small: it depends only on PyTorch (unpinned). Optional features live behind extras, so install only what you need:
pip install "attn-gym[linear]" # Linear-attention APIs and kernels
pip install "attn-gym[viz]" # Visualization and example dependencies
Warning
Attention Gym is under active development. We reserve the right to make
backward-incompatible changes between releases. If you depend on a particular API or kernel
behavior, hard-pin the version you test, for example:
pip install "attn-gym[linear]==X.Y.Z".
π» Usage
Attention Gym supports three complementary workflows:
- Compose FlexAttention building blocks. Import
mask_modandscore_modfunctions and pass them directly to PyTorch's FlexAttention APIs. - Use sparse and linear-attention APIs and kernels. Build with
selected_attention, GDN and KDA chunk, recurrent, and decode paths, and short-convolution primitives. See the compressed sparse attention and delta-rule (KDA/GDN) training for working examples. - Run real workloads and benchmarks. The
examples/directory covers paged, ring, and variable sparse attention, CUDA Graphs, determinism, compilation, and profiling. Most of this should serve as inspiration for fun things you might build from our building blocks :)
π οΈ Dev
Install dev requirements
pip install -e ".[dev]"
Install and run the repository hooks:
prek install
prek run --all-files
π€ Contributing
We welcome contributions to Attention Gym, especially new Masks or score mods! Here's how you can contribute:
Contributing Mods
- Create a new file in the attn_gym/masks/ for mask_mods or attn_gym/mods/ for score_mods.
- Implement your function, and add a simple main function that showcases your new function.
- Update the
attn_gym/*/__init__.pyfile to include your new function. - Optionally, add an end-to-end example using your new function in the examples/ directory.
See CONTRIBUTING.md for more details.
βοΈ License
Attention Gym-authored code is released under the BSD 3-Clause License. Vendored third-party components retain the licenses and notices included alongside their source.