Open in tt-awesome →

Attention in SRAM on Tenstorrent Grayskull

community
by · Jul 18, 2024

A fused kernel for the Grayskull architecture implementing Transformer self-attention entirely within SRAM. Combines matrix multiply, attention score scaling, and Softmax without DRAM accesses, achieving significant speedups over non-fused implementations.

attention transformer sram grayskull kernel risc-v
grayskull