Compiler Options
Code Generation Options
These flags control how TT-Lang compiles operations. Pass them on the command line,
or print the list with --ttl-help:
python my_kernel.py --ttl-help
python my_kernel.py --no-ttl-maximize-dst
Flag |
Default |
Description |
|---|---|---|
|
enabled |
Partition compute iteration spaces into subblocks that maximize DST register utilization, and reorder tile operations within sync regions to group by kind. Disabling falls back to per-tile synchronization. |
|
enabled |
Allow FPU strategy selection for binary add, subtract, and multiply when their operands permit it. Disabling selects SFPU. |
|
enabled |
Emit |
|
disabled |
Refine DFB reserve/push to per-subblock granularity, enabling |
|
enabled |
Combine consecutive |
|
enabled |
Prefer full-fp32 accumulation for reduce operations when supported by the target and the complete kernel configuration. |
|
enabled |
Prefer full-fp32 accumulation for matmul operations when supported by the target and the complete kernel configuration. |
|
disabled |
Error at compile time if a |
|
enabled |
Insert compiler-allocated intermediate DFBs when an operation requires DFB-attached inputs, fusion would read a source after its DFB is released, or a computed value is stored by operations in multiple MLIR basic blocks. When disabled, the compiler emits an error if materialization is required. |
|
enabled |
Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
|
enabled |
Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
|
disabled |
Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
|
|
Limit the logical transfers in one PipeTransport group. |
|
target-dependent |
Override the combined DFB, PipeNet, and synchronized-reset L1 allocation budget. |
|
enabled |
Reuse physical DFB indices when concurrent-kernel liveness proves that compatible logical DFB lifetimes do not overlap. Disabling compacts provisional user indices without introducing new user-DFB sharing. |
|
|
Examine at most |
|
disabled |
Trust explicit |
|
disabled |
Clone each TTKernel function whose control flow branches on a core coordinate once per launch coordinate ( |
Other Ways to Set These
Besides the command line, the same flags can be set through three other mechanisms. When the same flag is set in multiple places, higher-priority sources win and unmentioned flags fall through from lower levels:
Priority |
Mechanism |
Example |
|---|---|---|
1 (lowest) |
|
— |
2 |
|
|
3 |
|
|
4 (highest) |
Command-line arguments ( |
|
The options keyword can also be passed at call time to override the decorator
for a single invocation:
my_kernel(tensor_a, tensor_b, options="--no-ttl-fpu-binary-ops")
Compute Configuration
These parameters are set on the @ttl.operation decorator (not via command-line
flags) and control the TTNN compute kernel hardware configuration:
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
|
|
Constrain the Wormhole B0/Blackhole DST register-file element width: |
|
|
|
Enable full DST synchronization (single-buffering mode). Doubles DST capacity (32-bit elements: 8, 16-bit elements: 16) at the cost of a full sync between math and pack threads. |
|
|
|
Set the compute math fidelity to |
@ttl.operation(
grid=(2, 2),
fp32_dest_acc_en=True,
dst_full_sync_en=False,
math_fidelity="HiFi4",
)
def my_kernel(a, b): ...
Environment Variables
These environment variables control compilation behavior and diagnostic output. They are independent of the code generation flags above.
Variable |
Type |
Default |
Description |
|---|---|---|---|
|
|
|
Compile kernels but do not execute on hardware. |
|
file path |
(unset) |
Write the pre-optimization MLIR module to this file. |
|
file path |
(unset) |
Write the post-optimization MLIR module to this file. |
|
any value |
(unset) |
Print the IR after every pass in the pipeline. Output is very large; redirect to a file. |
|
|
|
Include source locations in printed MLIR (locations are always tracked internally for error messages). |
|
|
|
Include raw MLIR diagnostics in error output. |
|
|
|
Force |
|
any value |
(unset) |
Skip compiler verification that DFB producers, consumers, and waits execute on corresponding dynamically active launch nodes. The program must enforce those ownership and synchronization contracts. Finalized DFB preconditions, PipeNet endpoint guards, transfer correspondence, and synchronization schedules remain enabled. |
Profiling-related environment variables (TTLANG_AUTO_PROFILE,
TTLANG_PERF_DUMP, TTLANG_PERF_SERV, TTLANG_SIGNPOST_PROFILE,
TTLANG_PROFILE_CSV) are documented in the
Performance Tools reference.
Other Decorator Parameters
The @ttl.operation decorator also accepts these parameters for operation structure
and layout:
Parameter |
Type |
Default |
Description |
|---|---|---|---|
|
|
(required) |
Compute grid dimensions, e.g., |
|
|
|
Lambda functions for tile indexing |
|
|
|
|
|
|
|
Number of output tensor arguments |
|
|
|
Memory space for dataflow buffers: |
|
|
|
Use tiled tensor layout |
ttlang-opt Pass Reference
ttlang-opt is the standalone MLIR optimizer driver for the TTL dialect, used
primarily for compiler development and testing. It accepts all standard
mlir-opt flags (run ttlang-opt --help for the full list) plus the
TTL-specific passes and pipeline documented below.
Pipeline: ttl-to-ttkernel-pipeline
The main compilation pipeline, equivalent to what the Python API runs internally.
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{maximize-dst=true lower-to-emitc=true}'
Option |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Enable DST maximization via subblock compute and scheduling. |
|
bool |
|
Allow FPU strategy selection for binary add/sub/mul. |
|
bool |
|
Lower matmul to block-level hardware calls ( |
|
bool |
|
Refine DFB reserve/push to per-subblock granularity. |
|
bool |
|
Combine consecutive |
|
bool |
|
Prefer full-fp32 reduce accumulation when supported. |
|
bool |
|
Prefer full-fp32 matmul accumulation when supported. |
|
bool |
|
Error if a |
|
bool |
|
Insert compiler-allocated intermediate DFBs for DFB-only operands, source-lifetime preservation, and computed values stored by operations in multiple MLIR basic blocks. Error if disabled and any operation requires one. |
|
bool |
|
Use computed receiver DFB addresses for eligible PipeNet transfers. When disabled, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
|
bool |
|
Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When disabled, computed-address transfers use receiver-post synchronization. |
|
bool |
|
Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage, leaving local hardware semaphore ids available to the application. |
|
int64_t |
|
Limit logical transfers per PipeTransport group. |
|
uint32_t |
|
Override the combined DFB, PipeNet, and synchronized-reset L1 allocation budget. |
|
bool |
|
Reuse physical DFB indices for compatible logical DFBs with proven non-overlapping concurrent lifetimes. |
|
bool |
|
Trust explicit DFB allocation-group handoffs that lack a complete compiler proof. Automatic reuse remains proof-based. |
|
uint64 |
|
Maximum states examined during deterministic exact DFB allocation before reporting an inconclusive result. |
|
bool |
|
Clone TTKernel functions that branch on a core coordinate once per launch coordinate ( |
|
bool |
|
Run the TTKernel-to-EmitC backend (produces C++ source). |
The pipeline runs these passes in order:
ttl-materialize-loop-state– replace ranked-tensor loop-carried values with compiler-created DFBsttl-insert-copy-wait– insert missingttl.waitafterttl.copyops whose transfer handle has no wait userttl-annotate-l1-acc-loops– detect+=accumulation loops and annotate for L1 packer accumulationttl-create-producer-compute– create producerttl.computeoperations before intermediate materializationttl-insert-intermediate-dfbs– materialize DFB-only operands, values that must be preserved before source release, and computed values stored by operations in multiple MLIR basic blocks; verify and error whencompiler-dfbs=falseconvert-ttl-to-compute– lower TTL elementwise tensor ops tottl.computewith tile opsttl-insert-cb-sync– insert missing DFB synchronizationttl-verify-pipenet-guards, thenttl-verify-pipenet-schedule– verify PipeNet launch domains and event ordering while logical DFB identities remain distinct and before physical DFB allocationttl-form-pipe-transports– group eligible repeated PipeNet transfers and select bounded receiver storage while reserving synchronized-reset scratchttl-coalesce-dfb-acquires– coalesce compatible DFB acquiresttl-finalize-dfb-indices– assign logical DFBs to physical indices, validate combined DFB and synchronized-reset capacity, and emit runtime metadata;reuse-user-dfbscontrols automatic user-DFB reuse,unsafe-assume-allocation-groupstrusts only explicit unproved group handoffs, andexact-coloring-search-limitbounds exhaustive fixed-limit and minimum physical-index-count queriesttl-set-compute-kernel-config– select tile execution strategies and resolve kernel-wide DST and per-DFB unpack configurationttl-assign-dst– DST register allocation (linear scan with copy insertion)ttl-subblock-compute-for-dst– tilettl.computeinto DST-sized subblocks (only ifmaximize-dst=true); optionally refine reserve/push to per-subblock granularity (only ifsubblock-sync=true)ttl-lower-to-loops– lowerttl.computetoscf.forloops; matmul computes are expanded inline viagenerateMatmulComputettl-schedule-operations– reorder tile ops by dependency depth and kind (only ifmaximize-dst=true)ttl-annotate-cb-associations– annotate block args with DFB indicesttl-verify-dfb-spsc– verify per-node DFB producer/consumer uniqueness after finalizationttl-erase-pipenet-scopes– remove verified PipeNet structural markersttl-validate-cb-budget– verify static DFB storage plus synchronized-reset scratch fits the per-core L1 budgetconvert-ttl-to-ttkernel– lower TTL DMA, PipeNet, and synchronized-reset operations to TTKernel, select their runtime resources, and validate the exact combined L1 allocationttkernel-insert-inits– insert hardware init ops before compute opsttkernel-insert-l1-accumulation– insertpack_reconfig_l1_accguards for+=and reduction loopsttkernel-combine-pack-tiles– combine consecutivepack_tileintopack_tile_block(only ifcombine-pack-tiles=true)Canonicalization and CSE cleanup
ttkernel-specialize-cores, thencanonicalize,cse– per-core clone and const-fold of coordinate branches; tags clones withttl.core_coord(only ifspecialize-cores=true)(if
lower-to-emitc=true)lower-affine,convert-ttkernel-to-emitc,emitc-form-expressions
Individual Pass Options
Each pass can also be run standalone for testing. Only passes with configurable options are listed; the remaining passes have no options.
ttl-insert-intermediate-dfbs
Insert compiler-allocated intermediate DFBs where tensor SSA values require concrete DFB storage.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Insert compiler-allocated DFBs. When false, emit an error if any operation requires one. |
ttlang-opt input.mlir -p 'func.func(ttl-insert-intermediate-dfbs{enable=false})'
ttl-finalize-dfb-indices
Assign physical indices to logical DFBs and emit the complete runtime allocation table.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Reuse a physical index when concurrent-kernel liveness proves that two compatible logical DFB lifetimes cannot overlap. When false, compact provisional user indices without introducing new user-DFB sharing and apply the same lifetime proof only to compiler-created DFBs. |
|
uint64 |
|
Examine at most this many states during deterministic exact DFB allocation. Exhaustive search runs when order-dependent first-fit prevents acceptance or exceeds the provisional threshold after a conservative PipeNet reservation. Reaching the limit fails with an inconclusive-search diagnostic only when acceptance requires the search result; a reservation-only search may retain an authoritative-budget-valid first-fit assignment. |
|
uint32_t |
|
Override the combined DFB and synchronized-reset scratch budget used during physical allocation. |
|
bool |
|
Trust explicit DFB allocation groups when launch-domain, access-completion, pointer-handoff, or lifetime-order proof is incomplete. Emit one warning per accepted group and record the assumptions in |
ttlang-opt input.mlir -p 'builtin.module(ttl-finalize-dfb-indices{reuse-user-dfbs=true unsafe-assume-allocation-groups=false exact-coloring-search-limit=1000000})'
ttl-set-compute-kernel-config
Resolve tile execution strategies and shared compute-kernel configuration. See Compute Kernel Configuration for the algorithm and invariants.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
string |
|
Select 32-bit destination elements through the Wormhole B0/Blackhole |
|
string |
|
Select full DST synchronization: |
|
bool |
|
Prefer full-fp32 reduce accumulation when supported. |
|
bool |
|
Prefer full-fp32 matmul accumulation when supported. |
|
bool |
|
Allow eligible add/sub/mul operations to select FPU. |
ttlang-opt input.mlir -p 'ttl-set-compute-kernel-config{fp32-dest-acc-en=enabled}'
ttl-assign-dst
DST register allocator using linear scan allocation with in-place operation merging.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
uint32_t |
|
Override DST register capacity. Auto-computed from |
|
bool |
|
Allocate outputs in a separate DST region (needed for reductions and some loop optimizations). |
ttlang-opt input.mlir -p 'func.func(ttl-assign-dst{dst-capacity=16})'
ttl-subblock-compute-for-dst
Partition ttl.compute into DST-sized subblocks.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Refine DFB reserve/push to per-subblock granularity, enabling |
|
bool |
|
Error if a |
ttlang-opt input.mlir -p 'func.func(ttl-subblock-compute-for-dst{subblock-sync=true})'
ttl-form-pipe-transports
Group eligible repeated PipeNet transfers and select bounded receiver storage. Later PipeTransport planning replaces proven-private grouped DFB lifecycles with transport-owned scratch; scalar residuals retain the original lifecycle. Selection accounts for DFB allocation, a conservative receiver-published address table, transport scratch, and synchronized-reset scratch.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
int64_t |
|
Limit logical transfers per group. |
|
uint32_t |
|
Override the combined DFB, PipeNet, and synchronized-reset budget used during grouping selection. |
ttlang-opt input.mlir --ttl-form-pipe-transports='group-size=8'
ttl-validate-cb-budget
Validate finalized static DFB storage and allocator-rounded synchronized-reset
scratch against the per-core L1 budget. Exact PipeNet scratch is added by
convert-ttl-to-ttkernel after transport planning.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
uint32_t |
|
Override the combined static DFB and synchronized-reset scratch budget. |
ttlang-opt input.mlir -p 'builtin.module(ttl-validate-cb-budget{l1-budget-override=98304})'
convert-ttl-to-ttkernel
Lower TTL data movement and PipeNet operations to TTKernel.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Enable FP32 accumulation for reduce operations. |
|
bool |
|
Use computed receiver DFB addresses for eligible PipeNet transfers. When false, transfers use receiver-published destination addresses; multicast still requires proven equal runtime receiver addresses. |
|
bool |
|
Use capacity-counter synchronization when the receiver wait and pop execute on the receiver NOC thread and the computed-address transfer passes the DFB ownership and count proofs. When false, computed-address transfers use receiver-post synchronization. |
|
bool |
|
Allocate all compiler-managed PipeNet synchronization counters in GlobalSemaphore storage. |
|
uint32_t |
|
Override the budget used for final exact DFB, PipeNet scratch, and synchronized-reset scratch validation. |
ttlang-opt input.mlir -p 'builtin.module(convert-ttl-to-ttkernel{pipe-computed-addresses=true pipe-capacity-sync=false pipe-global-semaphores-only=true})'
ttl-dump-cb-flow-graph
Analyze dataflow buffer producer/consumer relationships and dump the flow graph.
Option |
Type |
Default |
Description |
|---|---|---|---|
|
string |
|
Path to write JSON output. Empty string prints to stderr only. |
ttlang-opt input.mlir -p 'ttl-dump-cb-flow-graph{output="/tmp/cb_graph.json"}'
ttkernel-specialize-cores
Clone TTKernel functions that branch on a core coordinate once per launch
coordinate. Requires a module-level ttl.launch_grid attribute (an i64 array
of length 2 with positive entries). Missing or malformed ttl.launch_grid is
a hard error. A valid single-core grid (product <= 1) skips specialization.
Only scf.if conditions derived from ttkernel.my_logical_x_ /
ttkernel.my_logical_y_ trigger cloning. Functions with symbol uses (for
example func.call targets) are left unspecialized with a warning so erasing
the original does not leave dangling SymbolRefAttrs; unrelated functions in
the module are still specialized. Each clone replaces coordinate reads with
arith.constants and is tagged with ttl.core_coord for runtime dispatch.
Downstream canonicalize / cse fold the now-constant branches.
This pass is off by default. Enable it through the pipeline option
specialize-cores (Python: --ttl-specialize-cores):
ttlang-opt input.mlir -p 'ttl-to-ttkernel-pipeline{specialize-cores=true lower-to-emitc=true}'
# Or stand-alone:
ttlang-opt input.mlir -p 'builtin.module(ttkernel-specialize-cores,canonicalize,cse)'