Attention in SRAM on Tenstorrent Grayskull
communityA fused kernel for the Grayskull architecture implementing Transformer self-attention entirely within SRAM. Combines matrix multiply, attention score scaling, and Softmax without DRAM accesses, achieving significant speedups over non-fused implementations.
Links
📄
arXiv:2407.13885
Works on
grayskull