Contributing
June 9, 2026 · View on GitHub
Reporting issues
Use the Bug report / Feature request templates under Issues → New issue.
Development setup
git clone https://github.com/hcompai/late-interaction-kernels
cd late-interaction-kernels
uv sync --extra dev --extra pylate --extra torch-cuda # or --extra torch-cpu on CPU-only boxes
uv run ruff check . && uv run ruff format --check .
uv run pytest -q # CUDA tests auto-skip without a GPU
Pick exactly one of torch-cuda (CUDA index, cu124) or torch-cpu (CPU-only
wheel, what CI uses) — the two are declared conflicting in pyproject.toml. On
macOS, torch-cpu falls back to PyPI's default MPS-capable wheel.
GPU tests
GPU tests run on AWS CodeBuild (A10G). They do not fire on pushes to main
(CodeBuild spend); they run automatically on v* tag pushes and on PRs
carrying the run-gpu-tests label (applying the label requires triage+, so
ping a maintainer if your PR needs it). Maintainers can also trigger an
on-demand run via the GPU CI workflow
workflow_dispatch.
Style
- Python 3.10+; type hints on public APIs.
- Comments explain why, not what.
- Docstrings: short, concrete, no marketing.
Autotune for a new GPU
Run the relevant benchmarks/kernels/bench_*.py, add a shortlist in
late_interaction_kernels/_autotune.py keyed on the device-name prefix,
re-run and include before / after in the PR.
New kernel variant
Keep it in its own module under late_interaction_kernels/ and follow
the existing split: internal _forward returning (scores, argmax),
torch.autograd.Function wrapper, pure-PyTorch reference in
reference.py, parity tests in tests/.
Publishing a release
mainis green andCHANGELOG.mdUnreleasedis filled in.- Releases → Draft a new release on GitHub, tag
vX.Y.Zoffmain. - Paste the matching
CHANGELOG.mdsection as the body, then Publish.
publish.yml builds and uploads to PyPI
via OIDC — no token needed.
Note
Pushing the tag also triggers CI, CPU CI, and GPU CI on the tagged
commit in parallel with publish.yml. Confirm they're green on the
Actions page
so a broken release is caught immediately.
License
By contributing you agree your work is licensed under Apache 2.0
(see LICENSE).