tt-inference-server
officialProduction-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. Supports continuous batching, multiple models, and all TT hardware configurations.
Works on
wormhole
blackhole
quietbox
galaxy