kernels.md
February 4, 2026 ยท View on GitHub
Setting up kernels
-
Custom CUDA layernorm kernels modified from FastFold and Oneflow accelerate about 30%-50% during different training stages.
fast_layernormis used by default, and no explicit setting is required.- If you wish to disable it and use the native PyTorch layernorm, you can set the environment variable to
torch:export LAYERNORM_TYPE=torch - The kernels will be JIT-compiled when
fast_layernormis called for the first time.
-
Triangle_attention Kernel Options The model supports four implementations for triangle attention, configurable in configs_base.py:
triangle_attention = "cuequivariance" # or "triattention"/"deepspeed"/"torch"-
TriAttention kernel(Default)
Custom kernel implementation from protenix/model/tri_attention/
-
cuEquivariance Kernel Optimized implementation using NVIDIA's cuEquivariance library.
-
DeepSpeed DS4Sci_EvoformerAttention kernel is a memory-efficient attention kernel developed as part of a collaboration between OpenFold and the DeepSpeed4Science initiative.
DS4Sci_EvoformerAttention is implemented based on CUTLASS. If you use this feature, you need to clone the CUTLASS repository and specify the path to it in the environment variable
CUTLASS_PATH. The Dockerfile already includes this setting:RUN git clone -b v3.5.1 https://github.com/NVIDIA/cutlass.git /opt/cutlass ENV CUTLASS_PATH=/opt/cutlassIf you set up
Protenixbypip, you can set environment variableCUTLASS_PATHas follows:git clone -b v3.5.1 https://github.com/NVIDIA/cutlass.git /path/to/cutlass export CUTLASS_PATH=/path/to/cutlassThe kernels will be compiled when DS4Sci_EvoformerAttention is called for the first time.
-
-
Triangle_multiplicative Kernel Options
The Triangle Multiplicative operation supports two implementations, configurable in configs_base.py:
triangle_multiplicative = "cuequivariance" # or "torch"-
cuEquivariance Kernel (Default) Optimized implementation using NVIDIA's cuEquivariance library.
-
Torch Native Standard PyTorch implementation.
-