add_tiles

void ckernel::add_init(uint32_t icb0, uint32_t icb1, bool acc_to_dest = false, uint32_t call_line = __builtin_LINE())

Paired init for two-operand element-wise addition (add_tiles). Configures the unpacker and math pipeline so that SrcA <- icb0 and SrcB <- icb1. Call before add_tiles. The one-time hardware configuration must already have been performed via compute_kernel_hw_startup(icb0, icb1, ocb) at the start of MAIN. For general information on init functions refer to any_init. For the dest-reuse variant (one operand from DST) use add_reuse_dest_init<reuse_dest> below.

Argument

Description

Type

Valid Range

Required

icb0

CB whose tile is unpacked into SrcA (operand A)

uint32_t

0 to 31

True

icb1

CB whose tile is unpacked into SrcB (operand B)

uint32_t

0 to 31

True

acc_to_dest

If true, operation = A + B + dst_tile_idx of add_tiles

bool

0,1

False

void ckernel::add_tiles(uint32_t icb0, uint32_t icb1, uint32_t itile0, uint32_t itile1, uint32_t idst)

Performs element-wise addition C=A+B of tiles in two CBs at given indices and writes the result to the DST register at index dst_tile_index. The DST register buffer must be in acquired state via acquire_dst call. This call is blocking and is only available on the compute engine.

Return value: None

Argument

Description

Type

Valid Range

Required

in0_cb_id

The identifier of the circular buffer (CB) containing A

uint32_t

0 to 31

True

in1_cb_id

The identifier of the circular buffer (CB) containing B

uint32_t

0 to 31

True

in0_tile_index

The index of tile A within the first CB

uint32_t

Must be less than the size of the CB

True

in1_tile_index

The index of tile B within the second CB

uint32_t

Must be less than the size of the CB

True

dst_tile_index

The index of the tile in DST REG for the result C

uint32_t

Must be less than the acquired size of DST REG

True

void ckernel::add_block(uint32_t icb0, uint32_t icb1, uint32_t start_itile0, uint32_t start_itile1, uint32_t start_idst, uint32_t ntiles)

Performs element-wise addition C=A+B on ntiles consecutive tile pairs from two CBs, writing each result to a consecutive DST register slot. This is the uniform block entry point for the add op: its body is a simple loop over add_tiles, so it inherits add_tiles’s semantics and requires the same initialization (add_init) to have been called first. The DST register buffer must be in acquired state via acquire_dst call. This call is blocking and is only available on the compute engine.

NOTE: The loop implementation is transitional. In the future this for-loop must be folded into a hardware MOP / REPLAY buffer (as is being done for Quasar) so the whole block issues as a single packed op; the blocking then lives in llk-lib without changing this signature. Tracked under the Compute API Split effort (tt-metal#35739); the per-op push-down lands in tt-metal#47482.

Return value: None

Argument

Description

Type

Valid Range

Required

icb0

The identifier of the circular buffer (CB) containing A

uint32_t

0 to 31

True

icb1

The identifier of the circular buffer (CB) containing B

uint32_t

0 to 31

True

start_itile0

The index of the first tile A within the first CB

uint32_t

Must be less than the size of the CB

True

start_itile1

The index of the first tile B within the second CB

uint32_t

Must be less than the size of the CB

True

start_idst

The index of the first tile in DST REG for the result C

uint32_t

Must be less than the acquired size of DST REG

True

ntiles

The number of consecutive tile pairs to add

uint32_t

start_idst + ntiles <= acquired DST REG size

True