Compiler Options
August 8, 2026 ยท View on GitHub
Code Generation Options
These flags control how TT-Lang compiles operations. Pass them on the command line,
or print the list with --ttl-help:
python my_kernel.py --ttl-help
python my_kernel.py --no-ttl-maximize-dst
| Flag | Default | Description |
|---|---|---|
--ttl-maximize-dst / --no-ttl-maximize-dst | enabled | Partition compute iteration spaces into subblocks that maximize DST register utilization, and reorder tile operations within sync regions to group by kind. Disabling falls back to per-tile synchronization. |
--ttl-fpu-binary-ops / --no-ttl-fpu-binary-ops | enabled | Allow FPU strategy selection for binary add, subtract, and multiply when their operands permit it. Disabling selects SFPU. |
--ttl-block-matmul / --no-ttl-block-matmul | enabled | Emit matmul_block (processes the full tile block atomically) instead of per-tile matmul loops. Disabling this option is not yet supported. |
--ttl-subblock-sync / --no-ttl-subblock-sync | disabled | Refine DFB reserve/push to per-subblock granularity, enabling pack_tile_block for contiguous subblocks. When disabled, user-placed reserve/push is preserved as written. |
--ttl-combine-pack-tiles / --no-ttl-combine-pack-tiles | enabled | Combine consecutive pack_tile ops on the same DFB with contiguous DST and DFB indices into a single pack_tile_block call. |
--ttl-reduce-full-fp32 / --no-ttl-reduce-full-fp32 | enabled | Prefer full-fp32 accumulation for reduce operations when supported by the target and the complete kernel configuration. |
--ttl-matmul-full-fp32 / --no-ttl-matmul-full-fp32 | enabled | Prefer full-fp32 accumulation for matmul operations when supported by the target and the complete kernel configuration. |
--ttl-strict-f32-acc / --no-ttl-strict-f32-acc | disabled | Error at compile time if a += accumulation loop's output block exceeds f32 DST capacity (4 tiles with double-buffering). When enabled, guarantees each accumulation step fits in a single DST section without subblocking. |
--ttl-compiler-dfbs / --no-ttl-compiler-dfbs | enabled | Insert compiler-allocated intermediate DFBs when an operation requires DFB-attached inputs, fusion would read a source after its DFB is released, or a computed value is stored by operations in multiple MLIR basic blocks. When disabled, the compiler emits an error if materialization is required. |
--ttl-pipe-computed-addresses / --no-ttl-pipe-computed-addresses | enabled | Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
--ttl-pipe-capacity-sync / --no-ttl-pipe-capacity-sync | enabled | Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
--ttl-pipe-global-semaphores-only / --no-ttl-pipe-global-semaphores-only | disabled | Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
--ttl-pipe-batch-tiles N | 0 (auto) | Limit the logical transfers in one PipeTransport group. 0 selects automatically and 1 disables grouping. |
--ttl-l1-budget N | target-dependent | Override the L1 allocation budget used by DFB validation and PipeTransport selection. |
--ttl-reuse-user-dfbs / --no-ttl-reuse-user-dfbs | enabled | Reuse physical DFB indices when concurrent-kernel liveness proves that compatible logical DFB lifetimes do not overlap. Disabling compacts provisional user indices without introducing new user-DFB sharing. |
--ttl-dfb-exact-coloring-search-limit N | 1000000 | Examine at most N states during deterministic exact DFB allocation when order-dependent first-fit does not satisfy a DFB index or L1 limit. This bounds compile time; reaching the limit reports an inconclusive result, not a capacity proof. |
--ttl-specialize-cores / --no-ttl-specialize-cores | disabled | Clone each TTKernel function whose control flow branches on a core coordinate once per launch coordinate (ttkernel-specialize-cores), replacing my_logical_x_ / my_logical_y_ with constants and tagging clones with ttl.core_coord for per-core dispatch. Opt-in. |
Other Ways to Set These
Besides the command line, the same flags can be set through three other mechanisms. When the same flag is set in multiple places, higher-priority sources win and unmentioned flags fall through from lower levels:
| Priority | Mechanism | Example |
|---|---|---|
| 1 (lowest) | CompilerOptions class defaults | โ |
| 2 | @ttl.operation decorator options= parameter | @ttl.operation(grid=(2,2), options="--no-ttl-maximize-dst") |
| 3 | TTLANG_COMPILER_OPTIONS environment variable | export TTLANG_COMPILER_OPTIONS="--no-ttl-fpu-binary-ops" |
| 4 (highest) | Command-line arguments (sys.argv) | python my_kernel.py --no-ttl-maximize-dst |
The options keyword can also be passed at call time to override the decorator
for a single invocation:
my_kernel(tensor_a, tensor_b, options="--no-ttl-fpu-binary-ops")
Compute Configuration
These parameters are set on the @ttl.operation decorator (not via command-line
flags) and control the TTNN compute kernel hardware configuration:
| Parameter | Type | Default | Description |
|---|---|---|---|
fp32_dest_acc_en | bool or None | None | Constrain the Wormhole B0/Blackhole DST register-file element width: true selects 32-bit elements and false selects 16-bit elements. When None, resolve the width from target capabilities and tile-operation requirements. |
dst_full_sync_en | bool or None | None | Enable full DST synchronization (single-buffering mode). Doubles DST capacity (32-bit elements: 8, 16-bit elements: 16) at the cost of a full sync between math and pack threads. |
math_fidelity | str or None | None | Set the compute math fidelity to LoFi, HiFi2, HiFi3, or HiFi4. When None, retain the TTNN default. |
@ttl.operation(
grid=(2, 2),
fp32_dest_acc_en=True,
dst_full_sync_en=False,
math_fidelity="HiFi4",
)
def my_kernel(a, b): ...
Environment Variables
These environment variables control compilation behavior and diagnostic output. They are independent of the code generation flags above.
| Variable | Type | Default | Description |
|---|---|---|---|
TTLANG_COMPILE_ONLY | 0/1 | 0 | Compile kernels but do not execute on hardware. |
TTLANG_INITIAL_MLIR | file path | (unset) | Write the pre-optimization MLIR module to this file. |
TTLANG_FINAL_MLIR | file path | (unset) | Write the post-optimization MLIR module to this file. |
TTLANG_VERBOSE_PASSES | any value | (unset) | Print the IR after every pass in the pipeline. Output is very large; redirect to a file. |
TTLANG_DEBUG_LOCATIONS | 0/1 | 0 | Include source locations in printed MLIR (locations are always tracked internally for error messages). |
TTLANG_VERBOSE_ERRORS | 0/1 | 0 | Include raw MLIR diagnostics in error output. |
TTLANG_SIM_ONLY | 0/1 | 0 | Force import ttl to skip loading the compiled MLIR extension. Used when running the simulator from a source tree without an installed tt-lang-sim wheel (which ships the same signal as a marker module). |
Profiling-related environment variables (TTLANG_AUTO_PROFILE,
TTLANG_PERF_DUMP, TTLANG_PERF_SERV, TTLANG_SIGNPOST_PROFILE,
TTLANG_PROFILE_CSV) are documented in the
Performance Tools reference.
Other Decorator Parameters
The @ttl.operation decorator also accepts these parameters for operation structure
and layout:
| Parameter | Type | Default | Description |
|---|---|---|---|
grid | tuple or Callable | (required) | Compute grid dimensions, e.g., (2, 2) |
indexing_maps | list[Callable] | None | Lambda functions for tile indexing |
iterator_types | list[str] | None | "parallel" or "reduction" per dimension |
num_outs | int | 1 | Number of output tensor arguments |
memory_space | str | "L1" | Memory space for dataflow buffers: "L1" or "DRAM" |
tiled | bool | True | Use tiled tensor layout |
ttlang-opt Pass Reference
ttlang-opt is the standalone MLIR optimizer driver for the TTL dialect, used
primarily for compiler development and testing. It accepts all standard
mlir-opt flags (run ttlang-opt --help for the full list) plus the
TTL-specific passes and pipeline documented below.
Pipeline: ttl-to-ttkernel-pipeline
The main compilation pipeline, equivalent to what the Python API runs internally.
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{maximize-dst=true lower-to-emitc=true}'
| Option | Type | Default | Description |
|---|---|---|---|
maximize-dst | bool | true | Enable DST maximization via subblock compute and scheduling. |
enable-fpu-binary-ops | bool | true | Allow FPU strategy selection for binary add/sub/mul. |
use-block-matmul | bool | true | Lower matmul to block-level hardware calls (matmul_block). |
subblock-sync | bool | false | Refine DFB reserve/push to per-subblock granularity. |
combine-pack-tiles | bool | true | Combine consecutive pack_tile ops into pack_tile_block. |
reduce-full-fp32 | bool | true | Prefer full-fp32 reduce accumulation when supported. |
matmul-full-fp32 | bool | true | Prefer full-fp32 matmul accumulation when supported. |
strict-f32-acc | bool | false | Error if a += accumulation loop's output block exceeds f32 DST capacity. |
compiler-dfbs | bool | true | Insert compiler-allocated intermediate DFBs for DFB-only operands, source-lifetime preservation, and computed values stored by operations in multiple MLIR basic blocks. Error if disabled and any operation requires one. |
pipe-computed-addresses | bool | true | Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
pipe-capacity-sync | bool | true | Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
pipe-global-semaphores-only | bool | false | Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
pipe-batch-tiles | int64_t | 0 (auto) | Limit logical transfers per PipeTransport group. 0 selects automatically and 1 disables grouping. |
l1-budget-override | uint32_t | 0 (target default) | Override the L1 allocation budget used by DFB validation and PipeTransport selection. |
reuse-user-dfbs | bool | true | Reuse physical DFB indices for compatible logical DFBs with proven non-overlapping concurrent lifetimes. |
exact-coloring-search-limit | uint64 | 1000000 | Maximum states examined during deterministic exact DFB allocation before reporting an inconclusive result. |
specialize-cores | bool | false | Clone TTKernel functions that branch on a core coordinate once per launch coordinate (ttkernel-specialize-cores), then run canonicalize / cse. Maps from --ttl-specialize-cores. |
lower-to-emitc | bool | false | Run the TTKernel-to-EmitC backend (produces C++ source). |
The pipeline runs these passes in order:
ttl-materialize-loop-state-- replace ranked-tensor loop-carried values with compiler-created DFBsttl-insert-copy-wait-- insert missingttl.waitafterttl.copyops whose transfer handle has no wait userttl-annotate-l1-acc-loops-- detect+=accumulation loops and annotate for L1 packer accumulationttl-create-producer-compute-- create producerttl.computeoperations before intermediate materializationttl-insert-intermediate-dfbs-- materialize DFB-only operands, values that must be preserved before source release, and computed values stored by operations in multiple MLIR basic blocks; verify and error whencompiler-dfbs=falseconvert-ttl-to-compute-- lower TTL elementwise tensor ops tottl.computewith tile opsttl-insert-cb-sync-- insert missing DFB synchronizationttl-verify-pipenet-guards, thenttl-verify-pipenet-schedule-- verify PipeNet launch domains and event ordering while logical DFB identities remain distinct and before physical DFB allocationttl-form-pipe-transports-- group eligible repeated PipeNet transfers and select bounded receiver storagettl-coalesce-dfb-acquires-- coalesce compatible DFB acquiresttl-finalize-dfb-indices-- assign logical DFBs to physical indices, validate capacity, and emit runtime metadata;reuse-user-dfbscontrols user-DFB reuse andexact-coloring-search-limitbounds exhaustive fixed-limit and minimum physical-index-count queriesttl-set-compute-kernel-config-- select tile execution strategies and resolve kernel-wide DST and per-DFB unpack configurationttl-assign-dst-- DST register allocation (linear scan with copy insertion)ttl-subblock-compute-for-dst-- tilettl.computeinto DST-sized subblocks (only ifmaximize-dst=true); optionally refine reserve/push to per-subblock granularity (only ifsubblock-sync=true)ttl-lower-to-loops-- lowerttl.computetoscf.forloops; matmul computes are expanded inline viagenerateMatmulComputettl-schedule-operations-- reorder tile ops by dependency depth and kind (only ifmaximize-dst=true)ttl-annotate-cb-associations-- annotate block args with DFB indicesttl-verify-dfb-spsc-- verify per-node DFB producer/consumer uniqueness after finalizationttl-erase-pipenet-scopes-- remove verified PipeNet structural markersttl-validate-cb-budget-- verify static DFB storage fits the per-core L1 budgetconvert-ttl-to-ttkernel-- lower TTL DMA and PipeNet operations to TTKernel, selecting destination addressing, synchronization protocol, and synchronization-counter storagettkernel-insert-inits-- insert hardware init ops before compute opsttkernel-insert-l1-accumulation-- insertpack_reconfig_l1_accguards for+=and reduction loopsttkernel-combine-pack-tiles-- combine consecutivepack_tileintopack_tile_block(only ifcombine-pack-tiles=true)- Canonicalization and CSE cleanup
ttkernel-specialize-cores, thencanonicalize,cse-- per-core clone and const-fold of coordinate branches; tags clones withttl.core_coord(only ifspecialize-cores=true)- (if
lower-to-emitc=true)lower-affine,convert-ttkernel-to-emitc,emitc-form-expressions
Individual Pass Options
Each pass can also be run standalone for testing. Only passes with configurable options are listed; the remaining passes have no options.
ttl-insert-intermediate-dfbs
Insert compiler-allocated intermediate DFBs where tensor SSA values require concrete DFB storage.
| Option | Type | Default | Description |
|---|---|---|---|
enable | bool | true | Insert compiler-allocated DFBs. When false, emit an error if any operation requires one. |
ttlang-opt input.mlir -p 'func.func(ttl-insert-intermediate-dfbs{enable=false})'
ttl-finalize-dfb-indices
Assign physical indices to logical DFBs and emit the complete runtime allocation table.
| Option | Type | Default | Description |
|---|---|---|---|
reuse-user-dfbs | bool | true | Reuse a physical index when concurrent-kernel liveness proves that two compatible logical DFB lifetimes cannot overlap. When false, compact provisional user indices without introducing new user-DFB sharing and apply the same lifetime proof only to compiler-created DFBs. |
exact-coloring-search-limit | uint64 | 1000000 | Examine at most this many states during deterministic exact DFB allocation. Exhaustive search is used only when order-dependent first-fit prevents acceptance. Reaching the limit fails with an inconclusive-search diagnostic rather than a false capacity diagnostic. |
ttlang-opt input.mlir -p 'builtin.module(ttl-finalize-dfb-indices{reuse-user-dfbs=false exact-coloring-search-limit=1000000})'
ttl-set-compute-kernel-config
Resolve tile execution strategies and shared compute-kernel configuration. See Compute Kernel Configuration for the algorithm and invariants.
| Option | Type | Default | Description |
|---|---|---|---|
fp32-dest-acc-en | string | auto | Select 32-bit destination elements through the Wormhole B0/Blackhole fp32_dest_acc_en setting: auto, enabled, or disabled. |
dst-full-sync-en | string | auto | Select full DST synchronization: auto, enabled, or disabled. |
reduce-full-fp32 | bool | true | Prefer full-fp32 reduce accumulation when supported. |
matmul-full-fp32 | bool | true | Prefer full-fp32 matmul accumulation when supported. |
enable-fpu-binary-ops | bool | true | Allow eligible add/sub/mul operations to select FPU. |
ttlang-opt input.mlir -p 'func.func(ttl-set-compute-kernel-config{fp32-dest-acc-en=enabled})'
ttl-assign-dst
DST register allocator using linear scan allocation with in-place operation merging.
| Option | Type | Default | Description |
|---|---|---|---|
dst-capacity | uint32_t | 0 (auto) | Override DST register capacity. Auto-computed from fp32_dest_acc_en and dst_full_sync_en by default. Single-buffering (dst_full_sync_en=true): 32-bit elements=8, 16-bit elements=16. Double-buffering (default): 32-bit elements=4, 16-bit elements=8. |
separate-output-region | bool | false | Allocate outputs in a separate DST region (needed for reductions and some loop optimizations). |
ttlang-opt input.mlir -p 'func.func(ttl-assign-dst{dst-capacity=16})'
ttl-subblock-compute-for-dst
Partition ttl.compute into DST-sized subblocks.
| Option | Type | Default | Description |
|---|---|---|---|
subblock-sync | bool | false | Refine DFB reserve/push to per-subblock granularity, enabling pack_tile_block for contiguous subblocks. When disabled, user-placed reserve/push is preserved. |
strict-f32-acc | bool | false | Error if a += accumulation loop with non-f32 output requires subblocking. Subblocking reduces accumulation precision because bf16 L1 intermediates truncate f32 DST values. |
ttlang-opt input.mlir -p 'func.func(ttl-subblock-compute-for-dst{subblock-sync=true})'
ttl-form-pipe-transports
Group eligible repeated PipeNet transfers and select bounded receiver storage. Later PipeTransport planning replaces proven-private grouped DFB lifecycles with transport-owned scratch; scalar residuals retain the original lifecycle. Selection accounts for DFB allocation, a conservative receiver-published address table, and transport scratch.
| Option | Type | Default | Description |
|---|---|---|---|
group-size | int64_t | 0 (auto) | Limit logical transfers per group. 0 selects automatically and 1 disables grouping. |
l1-budget-override | uint32_t | 0 (target default) | Override the combined DFB and pipe scratch budget used during grouping selection. |
ttlang-opt input.mlir --ttl-form-pipe-transports='group-size=8'
convert-ttl-to-ttkernel
Lower TTL data movement and PipeNet operations to TTKernel.
| Option | Type | Default | Description |
|---|---|---|---|
reduce-full-fp32 | bool | true | Enable FP32 accumulation for reduce operations. |
pipe-computed-addresses | bool | true | Use computed receiver DFB addresses for eligible PipeNet transfers. When false, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
pipe-capacity-sync | bool | true | Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When false, computed-address transfers use receiver-post synchronization. |
pipe-global-semaphores-only | bool | false | Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage. |
ttlang-opt input.mlir -p 'builtin.module(convert-ttl-to-ttkernel{pipe-computed-addresses=true pipe-capacity-sync=false pipe-global-semaphores-only=true})'
ttl-dump-cb-flow-graph
Analyze dataflow buffer producer/consumer relationships and dump the flow graph.
| Option | Type | Default | Description |
|---|---|---|---|
output | string | "" | Path to write JSON output. Empty string prints to stderr only. |
ttlang-opt input.mlir -p 'ttl-dump-cb-flow-graph{output="/tmp/cb_graph.json"}'
ttkernel-specialize-cores
Clone TTKernel functions that branch on a core coordinate once per launch
coordinate. Requires a module-level ttl.launch_grid attribute (an i64 array
of length 2 with positive entries). Missing or malformed ttl.launch_grid is
a hard error. A valid single-core grid (product <= 1) skips specialization.
Only scf.if conditions derived from ttkernel.my_logical_x_ /
ttkernel.my_logical_y_ trigger cloning. Functions with symbol uses (for
example func.call targets) are left unspecialized with a warning so erasing
the original does not leave dangling SymbolRefAttrs; unrelated functions in
the module are still specialized. Each clone replaces coordinate reads with
arith.constants and is tagged with ttl.core_coord for runtime dispatch.
Downstream canonicalize / cse fold the now-constant branches.
This pass is off by default. Enable it through the pipeline option
specialize-cores (Python: --ttl-specialize-cores):
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{specialize-cores=true lower-to-emitc=true}'
# Or stand-alone:
ttlang-opt input.mlir -p 'builtin.module(ttkernel-specialize-cores,canonicalize,cse)'