Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent vs. NVIDIA L40S
communityShows that Text-to-Speech inference on Tenstorrent Lightning V2 achieves 4× lower cost than NVIDIA L40S. Applies BlockFloat8 (BFP8) and low-fidelity (LoFi) precision strategies to TTS despite their greater numerical fragility compared to LLMs.
Links
📄
arXiv:2604.03279
Works on
wormhole