Community · Open Source · Tenstorrent Ecosystem

A hidden dimension of Tenstorrent awesomeness

A curated directory of projects, tools, models, and research for Tenstorrent hardware — contributed by the community and our team. Browse by category or search across all entries.

151 Projects
13 Categories
🚀 Getting Started
The essential first steps — installer, core SDKs, and guided onboarding
tt-vscode-toolkit official
48 interactive lessons covering the full Tenstorrent developer path — from hardware detection to cus…
tt-sim Lab affiliated
A university teaching lab for TT-Metalium kernel programming on a virtual Tenstorrent chip — one-cli…
Tenstorrent Simulator Playground community
A web playground that runs real TTNN operations on the ttsim hardware simulator — no card required. …
🤖 AI & Models
Running, serving, and experimenting with AI models
TT Console official
Browser-based cloud console for exploring AI on Tenstorrent hardware. Run LLM inference, image and v…
tt-bio affiliated
Boltz-2 biomolecular model for drug discovery on Tenstorrent Blackhole. Supports single-card and mul…
koyeb/tenstorrent-examples community
Example applications and deployment configurations for running AI workloads on Tenstorrent hardware …
🕵️ AI Agents
Agentic systems and AI assistants running on TT hardware
tt-example-apps official
End-to-end AI applications running on Tenstorrent AI accelerators. Complete application examples fro…
Local AI Agents on Tenstorrent affiliated
Three agentic projects running fully on-device: local AI agents on QuietBox 2, a coding assistant po…
dstack community
Vendor-agnostic orchestration for training, inference, and agentic workloads across NVIDIA, AMD, TPU…
⚙️ Custom Kernels & Low-Level
Metalium/tt-lang kernel authoring; anything sub-compiler
tt-metal official
TT-NN operator library and TT-Metalium low-level kernel programming model. The primary SDK for devel…
Tenstorrent Cookbook: Particle Life Simulator affiliated
Particle Life simulation on Tenstorrent hardware — an emergent-behavior N-body system where simple a…
tt-tiny community
Minimal Python code to access and program the Tenstorrent Blackhole chip directly — George Hotz's ex…
🔨 Compilers & Frontends
Getting PyTorch/JAX/ONNX/CUDA models onto TT hardware
tt-forge official
Tenstorrent's MLIR-based compiler frontend. Enables running AI workloads from PyTorch, ONNX, and oth…
tt-forge-compiletron affiliated
Compile more than 100 models on tt-forge in a display format suitable for demos. Comprehensive showc…
Booth community
An open-source CUDA, HIP, and Triton compiler with no LLVM anywhere in the path. Takes the same sour…
🛠 Dev Tools & Debugging
Profiling, visualization, and debugging workloads
ttsim official
Fast full-system simulator of Tenstorrent Wormhole and Blackhole hardware. Runs TT-Metalium workload…
tensix-viz affiliated
Hardware topology visualizer for Tenstorrent chips — from individual chip to full cluster. Interacti…
nvtop community
htop-style process monitor for GPUs and AI accelerators. Supports AMD, Apple, Huawei, Intel, NVIDIA,…
🖥 Hardware & System
Drivers, firmware, monitoring, and hardware management
tt-kmd official
Tenstorrent kernel module driver. The Linux kernel module required to interface with Tenstorrent PCI…
tt-qb-lights affiliated
Sync your Tenstorrent Quietbox's RGB lighting to accelerator utilization status. Visual feedback for…
blackhole-py community
Pure Python driver for Tenstorrent Blackhole cards providing direct low-level hardware access withou…
☁️ Cloud & Orchestration
Kubernetes, cloud deployment, and multi-node infrastructure
tt-inference-server official
Production-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. Supports co…
tt-topology official
Configure Ethernet routing on multi-card Tenstorrent systems. Flash NB cards to use specific ETH rou…
vllm-tt-plugin official
Tenstorrent backend for vLLM, built on vLLM's standard plugin mechanism — install it alongside vLLM …
🔩 RISC-V & Architecture
ISA, simulation, and running Linux on TT silicon
tt-bh-linux official
Linux demo for the Tenstorrent Blackhole P100/P150 card RISC-V cores. Boot a real Linux kernel on th…
CS Fundamentals on Tenstorrent Hardware affiliated
Seven-module computer science curriculum taught on real Tenstorrent hardware. Covers RISC-V architec…
tt-sim community
Community-built Tenstorrent architecture simulator written in Python. Runs without hardware — useful…
🔬 Research & Papers
Academic papers, theses, and HPC experiments
tt-isa-documentation official
Low-level ISA and microarchitecture documentation for Tenstorrent AI architectures (Grayskull, Wormh…
polaris official
A high-level AI simulator from Tenstorrent for modeling and exploring AI accelerator and workload pe…
tt-tutorial (HPC) community
Tutorial on Tenstorrent hardware for HPC researchers from the RISC-V Testbed project at Edinburgh/EP…
🎮 Games & Demos
Creative, playful, and proof-of-concept projects
tt-animatediff official
Generates short, temporally coherent animated GIFs using the AnimateDiff model on Tenstorrent hardwa…
tt-zork-and-more affiliated
A Tenstorrent fork of Infocom's Zork I (and more!), running a Z-machine interpreter at least four di…
TT-GoL community
Conway's Game of Life implemented on Tenstorrent hardware using TT-Metal kernels.
📚 Guides, Tutorials & Education
Getting-started content, blog posts, lessons, courses
tt-installer official
Install the complete Tenstorrent software stack with one command. Handles drivers, firmware, Python …
Custom Model Training on Tenstorrent affiliated
Eight-lesson series covering the full custom training workflow on TT hardware: dataset fundamentals,…
Programming Tenstorrent Processors community
Deep-dive into the Tenstorrent architecture and Metalium programming model — circular buffers, kerne…
✍️ Blogs
Community and affiliated blogs covering Tenstorrent hardware, software, and AI
dev.to/mando222 — Tenstorrent & AI Blog affiliated
Eric Zietlow's blog covering Tenstorrent hardware, Metalium programming, and AI topics, sharing prac…
Tenstorrent Blackhole Architecture Guide community
A 6,500-word community deep dive into the Blackhole p100a architecture: the tile model (Tensix, DRAM…
A Gentle Guide: Tenstorrent Card on Arch Linux with Metalium community
Step-by-step guide to getting a Tenstorrent card running on Arch Linux with the full Metalium stack.…

Planet Tenstorrent is the ecosystem's live feed — new releases, articles, papers, talks, and community posts from across the Tenstorrent world, gathered in one place and updated daily. The latest five:

🏷 release official Aug 14, 2026
tt-kmd ttkmd-2.11.0
tenstorrent/tt-kmd

This release brings userspace access to Blackhole and Wormhole system-management firmware through a new TENSTORREAM_IOCTL_SMC_MSG interface—letting applications post requests and poll responses asynchronously via separate non-blocking calls—and forwards Blackhole firmware logs directly to the kernel log with configurable level filtering. More importantly, it squashes several nasty concurrency bugs: a deadlock between TLB allocation and process memory locking, a race where TLB windows could be reassigned mid-configuration, and device-removal races that could leave in-flight open() calls accessing freed state. The driver now also validates telemetry tables more strictly to catch corrupted firmware data and reports actual ARC errors from power-state changes instead of masking them with -EINVAL.

🏷 release official Aug 14, 2026
ttsim v1.10.1
tenstorrent/ttsim

This patch release brings ttsim closer to parity with real Wormhole and Grayskull silicon, fixing a handful of register-access bugs and relaxing overly strict alignment checks that were blocking valid instruction sequences—most notably 1-row MOVD2B and MOVB2D operations that are legitimate on hardware. Beyond those core fixes, WH simulation gains better NOC translation and atomics support, and several Tensix config registers are now modeled more completely, which matters if you're debugging data movement or synchronization issues in simulation before tape-out. The deprecation of TTSIM_SEMIHOSTING=1 signals a shift toward cleaner Ethernet firmware paths going forward.

🏷 release official Aug 14, 2026
tt-burnin v0.4.4
tenstorrent/tt-burnin

This release pulls in pyluwen 0.9.0, which brings improvements to the underlying device communication layer that tt-burnin relies on for stress testing. If you're running burnin tests on your Tenstorrent hardware, the updated pyluwen dependency should give you better stability and compatibility with the latest device firmware.

🏷 release official Aug 14, 2026
tt-smi v6.2.1
tenstorrent/tt-smi

tt-smi v6.2.1 bumps pyluwen to 0.9.0, bringing improved hardware telemetry support to the monitoring tool, alongside clearer documentation of GUI tab fields to help users interpret device metrics at a glance.

🏷 release community Aug 14, 2026
dstack 0.21.1
dstackai/dstack

Presets now support source code patching and dataset-driven benchmarking, letting you optimize serving frameworks and kernels with real workloads instead of synthetic prompts—and --previous lets you build on past optimization runs to refine configurations iteratively. Gateway fault tolerance has improved so unhealthy replicas won't block service provisioning, while in-place scaling means you can adjust replicas without redeploying, and dstack metrics now surfaces all jobs and replicas at once for easier inspection across multi-replica deployments. AWS p5en.48xlarge instances (8× H200 + EFA) are now supported, though Runpod spot offers have been removed following the platform's discontinuation notice.

🪐 Explore Planet Tenstorrent →
🚀 Getting Started
tt-metal official apt*apt*conda 1623⭐
TT-NN operator library and TT-Metalium low-level kernel programming model. The primary SDK…
tt-forge official 343⭐
Tenstorrent's MLIR-based compiler frontend. Enables running AI workloads from PyTorch, ONN…
tt-buda official 314⭐
TT-BUDA: Tenstorrent's original Python compiler and runtime for AI workloads. Legacy stack…
tt-mlir official 296⭐
Tenstorrent MLIR compiler — the core compiler infrastructure shared by tt-forge and other …
riscv-ocelot official 261⭐
The Berkeley Out-of-Order Machine with V-EXT (RISC-V Vector Extension) support. Tenstorren…
ttsim official 147⭐
Fast full-system simulator of Tenstorrent Wormhole and Blackhole hardware. Runs TT-Metaliu…
tt-isa-documentation official 125⭐
Low-level ISA and microarchitecture documentation for Tenstorrent AI architectures (Graysk…
riscv_arch_tests official 124⭐
RISC-V architectural self-checking directed tests — randomly-generated register operands a…
whisper official 96⭐
RISC-V Instruction Set Simulator (ISS) used by Tenstorrent for processor verification. Pow…
tt-xla official 74⭐
PJRT device plugin for Tenstorrent hardware. Enables JAX, PyTorch/XLA, and other XLA-based…
tt-kmd official apt* 71⭐
Tenstorrent kernel module driver. The Linux kernel module required to interface with Tenst…
RiESCUE official 69⭐
RISC-V Directed Test Framework and Compliance Suite. Comprehensive test infrastructure for…
tt-inference-server official 68⭐
Production-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. S…
tt-forge-onnx official 65⭐
ONNX graph compiler for Tenstorrent hardware. Optimizes and transforms ONNX model graphs f…
tt-buda-demos official 64⭐
Repository of model demos using TT-Buda. The largest collection of pre-compiled model exam…
tt-smi official pipapt* 62⭐
Tenstorrent System Management Interface — monitor device telemetry, issue board-level rese…
tt-lang official pippip 59⭐
Python-based DSL that sits between TT-NN and TT-Metalium — expresses custom fused kernels …
tt-bh-linux official 59⭐
Linux demo for the Tenstorrent Blackhole P100/P150 card RISC-V cores. Boot a real Linux ke…
tt-llk official 55⭐
Tenstorrent Low-Level Kernels: the C++ library that directly programs the RISC-V cores ins…
Jun 5, 2025
ttnn-visualizer official pip 54⭐
Comprehensive tool for visualizing and analyzing model execution on Tenstorrent hardware. …
WallaBMC official 52⭐
Lightweight BMC (Baseboard Management Controller) for STM32 and similar MCUs, with Web UI,…
TT-Studio official 49⭐
Web-based GUI for deploying and chatting with AI models on Tenstorrent hardware. Handles a…
tt-umd official 45⭐
User-mode driver for Tenstorrent hardware. The userspace layer that sits between the kerne…
tt-system-firmware official 42⭐
System firmware for Tenstorrent hardware. Low-level system initialization and control firm…
polaris official 39⭐
A high-level AI simulator from Tenstorrent for modeling and exploring AI accelerator and w…
luwen official cargoapt* 34⭐
Tenstorrent system interface library written in Rust. Low-level Rust bindings for communic…
tt-tvm official 31⭐
TVM for Tenstorrent ASICs. Brings the Apache TVM compiler stack to Tenstorrent hardware, e…
tensix-isa-simulator official 29⭐
ISA-level simulator for the Tensix compute engine. Simulates the matrix, vector, and scala…
tt-torch official 26⭐
Frontend integration for PyTorch with tt-mlir. Compile PyTorch models directly to Tenstorr…
tt-firmware official 24⭐
Tenstorrent firmware repository. Board management and control firmware for Tenstorrent acc…
tt-installer official 24⭐
Install the complete Tenstorrent software stack with one command. Handles drivers, firmwar…
tt-exalens official pip 21⭐
Low-level hardware debugger for Tenstorrent devices. Inspect register state, memory conten…
tt-blacksmith official 16⭐
Optimized training recipes for a variety of ML models on Tenstorrent hardware, powered by …
tt-topology official pipapt* 16⭐
Configure Ethernet routing on multi-card Tenstorrent systems. Flash NB cards to use specif…
tt-npe official 15⭐
Network-on-chip Performance Estimator for Tenstorrent Tensix-based devices. Model and esti…
tt-flash official pipapt* 14⭐
Tenstorrent firmware update utility. Flash new firmware onto Tenstorrent accelerator cards…
SFPI official apt* 14⭐
Tenstorrent SFPU programming interface — TT-enhanced RISC-V GCC and binutils plus header f…
tt-example-apps official 13⭐
End-to-end AI applications running on Tenstorrent AI accelerators. Complete application ex…
tt-forge-models official 13⭐
A shared repository of model implementations used across TT-Forge frontends — a single sou…
tt-perf-report official 11⭐
Performance report analysis tool for Tenstorrent Metal operations — analyzes perf traces t…
tt-vscode-toolkit official 8⭐
48 interactive lessons covering the full Tenstorrent developer path — from hardware detect…
Dec 18, 2025
tt-tools-common official pipapt* 7⭐
Shared helper library of common utilities used across Tenstorrent system tools such as tt-…
vllm-tt-plugin official 6⭐
Tenstorrent backend for vLLM, built on vLLM's standard plugin mechanism — install it along…
tt-toplike official apt*apt* 6⭐
A vibrant htop-style visualizer for Tenstorrent hardware written in Rust. Real-time proces…
tt-rpm official 6⭐
Cycle-level, execution-driven RISC-V CPU performance model built on Sparta (MAP) with Whis…
Documentation for the low-level layer of tt-metal: compute LLK APIs and data movement APIs…
tt-system-tools official apt* 5⭐
System setup and support utilities for Tenstorrent hardware — hugepages-setup configures t…
tt-emule official 4⭐
A C++ software emulator of the Tenstorrent device-level kernel and host APIs. Run tt-metal…
tt-CableGen official 4⭐
Network cabling visualizer for Tenstorrent scale-out deployments: describe a target topolo…
tt-local-generator official 3⭐
Generate infinite videos and images (and imaginative prompts to inspire them) on Tenstorre…
tt-kernel official 3⭐
Distributes models over the Hugging Face Hub and serves them on Tenstorrent hardware — `tt…
tt-burnin official pipapt* 3⭐
Command-line utility that runs a high power-consumption workload on Tenstorrent devices — …
tt-animatediff official
Generates short, temporally coherent animated GIFs using the AnimateDiff model on Tenstorr…
Official documentation hub for running Tenstorrent accelerators on Kubernetes. Centers on …
TT Console official
Browser-based cloud console for exploring AI on Tenstorrent hardware. Run LLM inference, i…
TenGEMM official
Tensix GEMM performance estimator and visualizer. A React app that models matrix-multiplic…
tt-cli official
Single entry point to the Tenstorrent software stack: `tt update` converges a machine onto…
Official setup and onboarding guide for the TT-QuietBox 2 — a compact, liquid-cooled AI wo…
ttsim-qemu official
Tenstorrent's fork of QEMU that provides the full-system emulation layer behind ttsim. Mod…
tt-bio affiliated 118⭐
Boltz-2 biomolecular model for drug discovery on Tenstorrent Blackhole. Supports single-ca…
· Jan 31, 2026
grayskull-attention affiliated 38⭐
FlashAttention-style attention kernel implemented entirely in on-chip SRAM on the Tenstorr…
tt-atom affiliated 26⭐
Meta's UMA interatomic potential running on Tenstorrent Blackhole — energy, forces, and st…
tt-lang-models affiliated 7⭐
A growing collection of models that use tt-lang for some or all of their implementation. R…
tt-zork-and-more affiliated 2⭐
A Tenstorrent fork of Infocom's Zork I (and more!), running a Z-machine interpreter at lea…
tt-qb-lights affiliated 2⭐
Sync your Tenstorrent Quietbox's RGB lighting to accelerator utilization status. Visual fe…
diamond affiliated 1⭐
DIAMOND: Atari game-playing agent implemented on Tenstorrent hardware via tt-lang. Diffusi…
gemma4 affiliated 1⭐
Gemma 4 language model implemented in tt-lang (e4b variant) for direct execution on Tensto…
open-oasis affiliated 1⭐
tt-lang inference script for Oasis 500M — an interactive video world model running on Tens…
tt-model-runner affiliated 1⭐
Discover, load, and benchmark models with a GUI and TUI for tt-inference-server. Makes exp…
tt-claw affiliated
A Tenstorrent-powered claw machine that rewards players with real prizes. The QuietBox 2 r…
Three agentic projects running fully on-device: local AI agents on QuietBox 2, a coding as…
dflash affiliated
DFlash: Block Diffusion for Flash Speculative Decoding on Tenstorrent hardware using tt-la…
Engram affiliated
A Tenstorrent port of the DeepSeek Engram model using tt-lang. Brings DeepSeek's memory-ef…
gsplat_tt affiliated
Port of Gaussian Splatting (3D scene reconstruction from 2D images) to Tenstorrent hardwar…
On-device image generation with Stable Diffusion XL running entirely on Tenstorrent hardwa…
Three lesson-projects covering on-device video synthesis: frame-by-frame diffusion with tt…
Eric Zietlow's blog covering Tenstorrent hardware, Metalium programming, and AI topics, sh…
Compile more than 100 models on tt-forge in a display format suitable for demos. Comprehen…
End-to-end image classification project using TT-Forge — compile and run a PyTorch classif…
tensix-viz affiliated
Hardware topology visualizer for Tenstorrent chips — from individual chip to full cluster.…
tt-warp affiliated
Warp terminal plugin for Tenstorrent — integrates hardware status, model management, and d…
Interactive browser-based visualizer of the Tenstorrent Tensix grid architecture. Explore …
TT-Metalium implementation of Conway's Game of Life as a cookbook recipe. Each generation …
Particle Life simulation on Tenstorrent hardware — an emergent-behavior N-body system wher…
tt-sim Lab affiliated
A university teaching lab for TT-Metalium kernel programming on a virtual Tenstorrent chip…
Seven-module computer science curriculum taught on real Tenstorrent hardware. Covers RISC-…
Eight-lesson series covering the full custom training workflow on TT hardware: dataset fun…
Three hands-on TT-Metalium kernel recipes: a Mandelbrot fractal explorer, real-time audio …
nvtop community 10914⭐
htop-style process monitor for GPUs and AI accelerators. Supports AMD, Apple, Huawei, Inte…
dstack community 2213⭐
Vendor-agnostic orchestration for training, inference, and agentic workloads across NVIDIA…
Booth community 1730⭐
An open-source CUDA, HIP, and Triton compiler with no LLVM anywhere in the path. Takes the…
tt-tiny community 69⭐
Minimal Python code to access and program the Tenstorrent Blackhole chip directly — George…
zyx community 63⭐
A complete ML library and compiler in Rust — "from assembly to neural networks" — with a n…
· Sep 25, 2022
tt-twitch community 29⭐
A Tenstorrent Grayskull kernel written live on Twitch by George Hotz. 120-core grid demons…
koyeb/tenstorrent-examples community 19⭐
Example applications and deployment configurations for running AI workloads on Tenstorrent…
blackhole-py community 18⭐
Pure Python driver for Tenstorrent Blackhole cards providing direct low-level hardware acc…
Tenstorrent Console Skill community 15⭐
An agent skill (SKILL.md) that teaches Claude Code, Codex, and the Agent SDK how to drive …
tenstorrent-tiny-examples community 14⭐
Simple C++ kernel experiments on a GraySkull e75 chip. Hands-on examples for learning the …
ttnn-helloworld-cpp community 14⭐
Minimal working example of using Tenstorrent TTNN in C++. The simplest possible starting p…
tt-sim community 14⭐
Community-built Tenstorrent architecture simulator written in Python. Runs without hardwar…
triton-tenstorrent community 12⭐
OpenAI Triton compiler plugin for Tenstorrent hardware. Write Triton kernels and target Te…
tt-iree community 12⭐
IREE (Intermediate Representation Execution Environment) ML compiler ported to Tenstorrent…
TT-GoL community 12⭐
Conway's Game of Life implemented on Tenstorrent hardware using TT-Metal kernels.
A web playground that runs real TTNN operations on the ttsim hardware simulator — no card …
ttMandelbrot community 8⭐
Mandelbrot Set fractal renderer running on Tenstorrent hardware. A classic demo showcasing…
tenstorrent.nix community 8⭐
Nix flake packaging the Tenstorrent software stack for NixOS and Nix users. Reproducible, …
TT-Metal Mini Template community 7⭐
Minimal working CMake project template for starting a new TT-Metal project from scratch. G…
tt-tutorial (HPC) community 7⭐
Tutorial on Tenstorrent hardware for HPC researchers from the RISC-V Testbed project at Ed…
tenstorrent-cli community 6⭐
A TypeScript/Bun terminal client for console.tenstorrent.com. Opens a chat REPL across Dee…
ttPEAK community 6⭐
clpeak-style peak-performance benchmark for Tenstorrent devices using TT-Metalium. Measure…
Unofficial documentation for the Blackhole P100A / P150, assembled from reverse-engineerin…
current community 5⭐
High-level parallel programming framework for Tenstorrent accelerators, abstracting TT-Met…
ttVecAdd community 5⭐
Minimal vector-addition example on Tenstorrent devices using TT-Metalium. A clean hello-wo…
bhx community 5⭐
Boot stock Linux cloud images on the SiFive X280 RISC-V cores inside Tenstorrent Blackhole…
ttas community 4⭐
ttas is a hacker-friendly assembler/disassembler for Tensix on Wormhole. It turns assembly…
tt-tutorial (Korean) community 4⭐
Comprehensive tutorials for the Tenstorrent software stack in Korean. Jupyter notebooks co…
ttRoPE community 4⭐
A GGML-formatted rotary positional embedding (RoPE) implementation for Tenstorrent hardwar…
Master's thesis implementing and benchmarking five allreduce algorithms (Swing, Recursive …
tt-model-bringup community 3⭐
Direct TT-Metal bringup of modern open-weight LLMs on Blackhole P150 — hand-written comput…
libtt-metal-cxx community 3⭐
Rust crate that exposes the TT-Metal host API through a C++ bridge via cxx.rs — covering d…
ttperf community pip 3⭐
A CLI wrapper that turns TT-Metal performance profiling into one command. Runs a pytest ta…
tt-monitor community 2⭐
A translucent, undecorated desktop widget showing live per-chip telemetry for Tenstorrent …
A compatibility guardrail that continuously monitors whether [tt-metal](https://github.com…
3D Gaussian Splatting rewritten to run on the matrix engine: a polynomial splat and order-…
ttWKV7 community 2⭐
A standalone tt-metal demo and test bench for the RWKV-7 (WKV7) state recurrence on Wormho…
libtt community 1⭐
A Bazel-built PJRT plugin (libtt.so) providing an XLA backend for Tenstorrent devices. Bun…
ttPseudoRowMajor community 1⭐
A small TTNN-facing C++ library (ttprm) for running view-shaped tensor work without first …
tt-tinygrad community
A Tenstorrent backend for tinygrad that targets TT-Lang rather than raw tt-metal: a Render…
A long third-party walkthrough of the Tenstorrent lineup — core architecture, per-product …
· Jun 12, 2026
Step-by-step guide to getting a Tenstorrent card running on Arch Linux with the full Metal…
· Jul 7, 2024
Honest field notes from getting a Grayskull card running and writing first Metalium kernel…
· Jun 2, 2024
Deep-dive into the Tenstorrent architecture and Metalium programming model — circular buff…
· Apr 21, 2025
Lecture 20 from William & Mary's graduate Computer Architecture course. Frames Tenstorrent…
· Oct 9, 2024
Tensix Field Guide community
A ten-chapter, plain-English tour of Tenstorrent's Tensix architecture written for someone…
Sponsored series of deep technical articles on implementing optimal SFPU kernels for the T…
· Nov 12, 2025
tt-rqm-kernels community
Structured quaternion, rotor, and phase-aware tensor kernels on ordinary floating-point te…
tt-wavelet community
One-level FP32 lifting wavelet transforms (LWT) on Wormhole and Blackhole, shipped as a TT…
A fused kernel for the Grayskull architecture implementing Transformer self-attention enti…
· Jul 18, 2024
Ports the Cooley-Tukey FFT algorithm to the Wormhole n300 RISC-V accelerator. The Wormhole…
· Jun 18, 2025
Evaluates the Tenstorrent Grayskull e75 RISC-V accelerator for matrix multiplication at re…
· May 9, 2025
Evaluates three strategies for scaling an N-body code across multiple Tenstorrent Wormhole…
· May 4, 2026
Accelerates an astrophysical N-body simulation on the Wormhole n300. Achieves 2× speedup a…
Nov 16, 2025
Implements three numerical kernels and composes them into a conjugate gradient solver on W…
Mar 24, 2026
Explores stencil computation on the Grayskull PCIe RISC-V accelerator. Early academic work…
Sep 27, 2024
Maps 2D 5-point stencil computations onto the Tenstorrent Wormhole RISC-V AI dataflow acce…
May 8, 2026
Makes multi-tenant NPU sharing practical for Blackhole-class hardware using polynomial-tim…
Apr 27, 2026
Compiler system that automatically generates efficient dataflow plans for tile-based langu…
· Dec 17, 2025
Shows that Text-to-Speech inference on Tenstorrent Lightning V2 achieves 4× lower cost tha…
· Mar 24, 2026
A 6,500-word community deep dive into the Blackhole p100a architecture: the tile model (Te…
· Feb 28, 2026
Martin Chang and Danfeng Zhang on solving real AI compute problems with open hardware, spa…
· Jan 31, 2026
Yuning Liang and Petr Penzin on closing the AI acceleration gap in the browser on RISC-V: …
· Jan 31, 2026
🏷 Recent Releases
42 releases
tt-kmd official ttkmd-2.11.0
2026-08-14T18:31:59Z
ttsim official v1.10.1
2026-08-14T18:28:50Z
tt-burnin official v0.4.4
2026-08-14T17:13:48Z
tt-smi official v6.2.1
2026-08-14T17:12:18Z
dstack community 0.21.1
2026-08-14T16:57:48Z
luwen official bh-mod-v2.1.0
2026-08-14T16:22:37Z
tt-exalens official v0.3.30
2026-08-14T14:47:59Z
tt-inference-server official v0.20.0
2026-08-14T14:11:44Z
tt-installer official v3.5.3
2026-08-13T17:40:44Z
ttnn-visualizer official v0.98.0
2026-08-12T21:18:05Z
SFPI official 7.69.0
2026-08-12T21:12:00Z
tt-system-firmware official v19.13.2
2026-08-11T21:11:11Z
tt-toplike official v0.8.0
2026-08-11T20:30:47Z
tt-metal official v0.76.0
2026-08-11T04:10:19Z
Booth community v0.5.2
2026-08-08T03:05:39Z
tt-bio affiliated v0.6.2
2026-08-07T21:33:46Z
tensix-viz affiliated v1.2.0
2026-08-05T16:35:19Z
tt-forge official 1.4.0
2026-07-30T11:17:16Z
tt-forge-onnx official 1.4.0
2026-07-30T11:12:02Z
tt-xla official 1.4.0
2026-07-30T11:06:22Z
TT-Studio official rc-v2.9.0
2026-07-28T18:09:27Z
tt-topology official v1.2.20
2026-07-23T15:35:12Z
tt-atom affiliated v0.2.1
2026-07-20T11:06:13Z
tt-vscode-toolkit official v0.1.19
2026-07-16T21:17:06Z
tt-local-generator official v0.11.0
2026-07-10T16:47:57Z
tt-umd official v0.9.9
2026-07-09T17:03:46Z
tt-flash official v3.10.0
2026-06-23T19:59:38Z
tt-animatediff official v0.9.0
2026-06-22T19:56:31Z
ttas community v0.1.0
2026-05-28T07:08:35Z
whisper official 1.861
2026-05-11T15:44:36Z
tt-sim community v1.0
2026-05-11T13:07:42Z
tt-bh-linux official v0.11
2026-04-13T15:10:59Z
tt-firmware official v19.6.0
2026-02-20T16:53:34Z
nvtop community 3.3.2
2026-02-08T17:57:16Z
tt-tools-common official v1.6.0
2025-12-23T21:02:08Z
tt-system-tools official v1.4.1
2025-12-08T17:23:48Z
RiESCUE official v1.7.0
2025-12-03T19:29:44Z
tt-torch official 0.4.0
2025-09-29T22:23:47Z
polaris official pre_perfmodel_merge
2025-09-19T18:16:27Z
riscv_arch_tests official v0.2.0+aligned-access
2025-01-23T17:16:16Z
tt-buda official v0.19.3
2024-09-24T21:01:08Z
zyx community v0.14.0
2024-09-22T13:54:32Z

Select an entry to see details

tt-metal

official
C++ · Apache-2.0 · 1623⭐ ·
tt-metal preview

TT-NN operator library and TT-Metalium low-level kernel programming model. The primary SDK for developing on Tenstorrent hardware — from high-level tensor ops to bare-metal RISC-V kernels.

metalium ttnn sdk kernels core
grayskull wormhole blackhole ttsim

tt-forge

official
Python · Apache-2.0 · 343⭐ ·
tt-forge preview

Tenstorrent's MLIR-based compiler frontend. Enables running AI workloads from PyTorch, ONNX, and other frameworks on all Tenstorrent hardware configurations through an open-source, general, and performant compiler.

LATEST 1.4.0 2026-07-30T11:17:16Z Release notes ↗
5 previous releases
1.5.0.dev20260814001246 2026-08-14T00:50:59Z
1.5.0.dev20260812001218 2026-08-12T00:51:20Z
1.5.0.dev20260811000841 2026-08-11T01:20:04Z
1.5.0.dev20260810000801 2026-08-10T00:45:45Z
1.5.0.dev20260809000535 2026-08-09T00:43:32Z
See all releases on GitHub ↗
mlir compiler pytorch onnx frontend
wormhole blackhole ttsim

tt-buda

official
Python · Apache-2.0 · 314⭐ ·

TT-BUDA: Tenstorrent's original Python compiler and runtime for AI workloads. Legacy stack — tt-forge is the recommended successor, but tt-buda has the largest model demo library.

📦 Repo
legacy compiler pytorch buda
grayskull wormhole

tt-mlir

official
C++ · Apache-2.0 · 296⭐ ·

Tenstorrent MLIR compiler — the core compiler infrastructure shared by tt-forge and other frontends. Handles graph optimization, lowering, and code generation for Tensix hardware.

5 releases
0.9.0.dev20260221pre 2026-02-21T04:31:50Z
0.9.0.dev20260220pre 2026-02-20T04:34:35Z
0.9.0.dev20260219pre 2026-02-19T04:37:24Z
0.9.0.dev20260218pre 2026-02-18T04:38:21Z
0.9.0.dev20260217pre 2026-02-17T04:37:09Z
See all releases on GitHub ↗
mlir compiler backend optimization
wormhole blackhole

riscv-ocelot

official ⑂ riscv-boom/riscv-boom
SystemVerilog · Apache-2.0 · 261⭐ ·
riscv-ocelot preview

The Berkeley Out-of-Order Machine with V-EXT (RISC-V Vector Extension) support. Tenstorrent's research-grade out-of-order RISC-V core with vector extension.

📦 Repo
risc-v out-of-order vector-extension processor-design

ttsim

official
C++ · Apache-2.0 · 147⭐ ·

Fast full-system simulator of Tenstorrent Wormhole and Blackhole hardware. Runs TT-Metalium workloads on any Linux/x86_64 system without physical silicon. Bit-exact results relative to hardware.

LATEST v1.10.1 2026-08-14T18:28:50Z Release notes ↗
4 previous releases
v1.10.0 2026-08-07T21:33:47Z
v1.9.9 2026-08-04T19:24:26Z
v1.9.8 2026-07-30T23:31:35Z
v1.9.7 2026-07-29T23:14:16Z
See all releases on GitHub ↗
simulator no-hardware bit-exact wormhole blackhole
ttsim

tt-isa-documentation

official
125⭐ ·

Low-level ISA and microarchitecture documentation for Tenstorrent AI architectures (Grayskull, Wormhole, Blackhole) — the authoritative hardware reference beneath the tt-forge / tt-metal software stack.

📦 Repo
isa architecture documentation tensix low-level
grayskull wormhole blackhole

riscv_arch_tests

official
Assembly · Apache-2.0 · 124⭐ ·

RISC-V architectural self-checking directed tests — randomly-generated register operands and data with low-level OS code for test scheduling and self-checking, runnable on a RISC-V design or an ISS such as Whisper or Spike. Generated by an internal Tenstorrent tool from the official RISC-V ISA spec.

📦 Repo
LATEST v0.2.0+aligned-access 2025-01-23T17:16:16Z Release notes ↗
2 previous releases
v0.2.0 2024-10-03T22:08:11Z
v0.1.1 2024-09-28T03:35:57Z
See all releases on GitHub ↗
riscv testing verification isa architecture

whisper

official ⑂ chipsalliance/VeeR-ISS
C++ · Apache-2.0 · 96⭐ ·

RISC-V Instruction Set Simulator (ISS) used by Tenstorrent for processor verification. Powers the co-simulation architecture checker.

📦 Repo
LATEST 1.861 2026-05-11T15:44:36Z Release notes ↗
See all releases on GitHub ↗
risc-v iss simulator verification

tt-xla

official
Python · Apache-2.0 · 74⭐ ·

PJRT device plugin for Tenstorrent hardware. Enables JAX, PyTorch/XLA, and other XLA-based frameworks to target TT accelerators.

LATEST 1.4.0 2026-07-30T11:06:22Z Release notes ↗
5 previous releases
1.5.0.dev20260814001246 2026-08-14T00:42:06Z
1.5.0.dev20260812001218 2026-08-12T00:43:01Z
1.5.0.dev20260811000841 2026-08-11T00:38:18Z
1.5.0.dev20260810000801 2026-08-10T00:37:09Z
1.5.0.dev20260809000535 2026-08-09T00:34:42Z
See all releases on GitHub ↗
xla pjrt jax pytorch
wormhole blackhole

tt-kmd

official
C · GPL-2.0 · 71⭐ ·

Tenstorrent kernel module driver. The Linux kernel module required to interface with Tenstorrent PCIe accelerator cards.

📦 Repo
LATEST ttkmd-2.11.0 2026-08-14T18:31:59Z Release notes ↗
4 previous releases
ttkmd-2.11.0-rc1pre 2026-08-11T16:26:19Z
ttkmd-2.10.0 2026-07-13T20:07:36Z
ttkmd-2.10.0-rc2pre 2026-07-10T19:55:51Z
ttkmd-2.10.0-rc1pre 2026-07-01T21:08:05Z
See all releases on GitHub ↗
kernel-module driver linux pcie
grayskull wormhole blackhole

RiESCUE

official
Python · Apache-2.0 · 69⭐ ·

RISC-V Directed Test Framework and Compliance Suite. Comprehensive test infrastructure for verifying RISC-V processor implementations against the specification.

LATEST v1.7.0 2025-12-03T19:29:44Z Release notes ↗
4 previous releases
v1.5.0 2025-11-17T21:58:14Z
v1.3.0 2025-11-06T20:12:13Z
v1.1.2 2025-10-16T17:21:43Z
v0.2.5 2025-07-10T00:59:12Z
See all releases on GitHub ↗
risc-v testing compliance verification

tt-inference-server

official
Python · Apache-2.0 · 68⭐ ·

Production-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. Supports continuous batching, multiple models, and all TT hardware configurations.

LATEST v0.20.0 2026-08-14T14:11:44Z Release notes ↗
4 previous releases
v0.19.0 2026-07-24T16:58:16Z
v0.18.0 2026-07-10T15:21:25Z
v0.17.0 2026-06-26T19:29:31Z
v0.16.0 2026-06-12T18:21:42Z
See all releases on GitHub ↗
serving openai-compatible production rest-api
wormhole blackhole quietbox galaxy

tt-forge-onnx

official
Python · Apache-2.0 · 65⭐ ·
tt-forge-onnx preview

ONNX graph compiler for Tenstorrent hardware. Optimizes and transforms ONNX model graphs for efficient execution on Tensix accelerators. Used as a backend by tt-forge for ONNX model ingestion.

📦 Repo
LATEST 1.4.0 2026-07-30T11:12:02Z Release notes ↗
5 previous releases
1.5.0.dev20260814014853 2026-08-14T02:11:01Z
1.5.0.dev20260813004712 2026-08-13T01:07:20Z
1.5.0.dev20260812005318 2026-08-12T01:41:10Z
1.5.0.dev20260811005055 2026-08-11T01:41:18Z
1.5.0.dev20260810004353 2026-08-10T01:02:45Z
See all releases on GitHub ↗
onnx compiler graph-optimization mlir
wormhole blackhole

tt-buda-demos

official
Python · Apache-2.0 · 64⭐ ·

Repository of model demos using TT-Buda. The largest collection of pre-compiled model examples for Tenstorrent hardware — BERT, ResNet, YOLO, GPT-2, Whisper, and many more.

📦 Repo
demos models bert resnet yolo gpt2
grayskull wormhole

tt-smi

official
Python · Apache-2.0 · 62⭐ ·
tt-smi preview

Tenstorrent System Management Interface — monitor device telemetry, issue board-level resets, and inspect hardware health. The nvidia-smi equivalent for Tenstorrent hardware.

📦 Repo
LATEST v6.2.1 2026-08-14T17:12:18Z Release notes ↗
4 previous releases
v6.2.0 2026-08-11T15:40:06Z
v6.1.0 2026-07-24T15:26:30Z
v6.0.0 2026-07-15T18:49:03Z
v5.3.1 2026-07-02T21:24:03Z
See all releases on GitHub ↗
# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## 3.0.26 - 29/07/25
- Added single tray galaxy reset option
- Bumped luwen from 0.7.5 -> 0.7.10
  - Chip detect now doesn't wait for eth to train for the 6U galaxy's, allowing multi tray resets to happen independently
- Updated readme with the new reset option

## 3.0.25 - 29/07/25
- Added packaging

## 3.0.24 - 04/07/25
- Now users have 2 galay reset modes available
  - glx_reset: resets the galaxy, informs users if there has been an eth failure
  - glx_reset_auto: resets the galaxy upto 3 times if eth failures are detected

## 3.0.23 - 03/07/25
- Bumped luwen 0.7.3 -> 0.7.5 to fix cargo lock compatibilty issue

## 3.0.22 - 02/07/25
- Bumped tt-tools-common 1.4.16 -> 1.4.17
- Bumped luwen 0.7.2 -> 0.7.3
- Bumped smi 3.0.21 -> 3.0.22

## 3.0.21 - 26/06/25

- Added option to not re-init chips after reset
- Updated galaxy 6u reset option from --ubb_reset to -glx_reset
- Removed the a3 arc message before doing a 6u reset, meaning we can reset even when chips are not pcie accessible
- Added eth link check and return failure if any of the eth links have a LINK_INACTIVE_FAIL_DUMMY_PACKET failure

## 3.0.20 - 04/06/25

- Chore - bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0

## 3.0.19 - 30/04/25

- Fixed an issue preventing the telemetry thread from being dispatched when the user clicked tab 2

## 3.0.18 - 22/05/25

- Added BH and WH UBB board type support
- Removed the dependency on tt-tools-common for this info

## 3.0.17 - 13/05/25

- Added proper telemetry heartbeat checks for Grayskull

## 3.0.16 - 12/05/25

- Used new ResetTypes from tools-common to simplify reset code
- Added a heartbeat spinner to the telemetry pane. We expect this spinner to update about twice per second. If the spinner is not moving, this indicates new telemetry is not being fetched.

## 3.0.15 - 24/04/25

- Patch for the ubb_reset to just discover local only post reset. Looks like eth port status 2 has been re-used to mean connected and pyluwen waits for it to clear, leading to eth timeout.

## 3.0.14 - 21/04/25

- Added wh ubb reset via command line `tt-smi --ubb_reset`. Intention is that this command line option will be removed and integrated into `tt-smi -r` after we update board detection with the correct external naming.
- Removed some unused imports and code - no functional changes

## 3.0.13 - 21/03/25

- Removed get\_sw\_versions

## 3.0.12 - 21/03/25

- Chore - bumped luwen version to include eth fw version check fix

## 3.0.11 - 13/03/25

- Chore - bumped luwen version to include enable chips with external connections but no routing

## 3.0.10 - 10/03/25

- Chore - bumped luwen version to include protoc lib detection check

## 3.0.9 - 07/03/25

- Chore - bumped luwen v
monitoring telemetry smi hardware-management
grayskull wormhole blackhole

tt-lang

official
Python · Apache-2.0 · 59⭐ ·
tt-lang preview

Python-based DSL that sits between TT-NN and TT-Metalium — expresses custom fused kernels with progressive disclosure, compiling directly to Tensix. Ships an integrated functional simulator (no hardware needed), line-by-line performance metrics, and AI-agent-friendly tooling. Two packages: tt-lang (compiler + hardware, requires ttnn) and tt-lang-sim (simulator only, works on Linux/macOS without Tenstorrent hardware).

# Changelog

All notable changes to TT-Lang will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## Version 1.1.1

### Compiler

- Fix for live-interval boundary computation (issue [#536](../../issues/536))
- Fix for all-zero results in FP32 reductions (issue # [#533](../../issues/533))
- Fix for inferred `pop` and `push` (issues [#536](../../issues/536), [#554](../../issues/554))
- Fix for write pointer tracking on pipe sender accross iterations (issue [#578](../../issues/578))
- Fix to report data type mismatch error
- Fix to report DFB over allocation error (issue [#511](../../issues/511))
- Support for pipenet predicates `is_src`, `is_dst` and `is_active` (issue [#541](../../issues/541))
- Support for `ttl.math.typecast`

### Simulator

- Support for inferred `pop`, `push` and `copy`'s transfer handle `wait`
- Support for pipenet predicates `is_src`, `is_dst` and `is_active`
- Support `all_gather`
- Support `bfloat8_b`
- Improved/actionable error messages
- Improved performance by simulating math in FP32

### Infrastructure

- TT-Lang installable with `pip install tt-lang` for full installation and `pip install tt-lang-sim` for simulator only
- [Matmul benchmarks](benchmarks/matmul/README.md)

## Version 1.0.0

### Compiler

- Support `+=` syntax in conjunction with dot product (`@`) lowered to packer L1 accumulation
- Support implicit temporary compute-kernel-local DFBs
- Support `ttl.Pipenet`
- Support implicit `ttl.Block.push` and `ttl.Block.pop`
- Support implicit `ttl.Transfer.wait`
- Support for `expm1`, `exp2`, `ceil`, `sign`, `gelu`, `silu`, `hardsigmoid`, `square`, `softsign`, `signbit`, `frac`, `trunc` in `ttl.math`

### Simulator

- Support for `ttl.GroupTransfer`
- SPMD and mesh device simulation support
- Support for `ttnn.all_reduce` CCLs
- Use tracing to report statistics with `tt-lang-sim-stats`
- Remote L1 reads/writes statistics

### Examples and documentation
- Matmul tutorial

## Version 0.1.8

### Compiler

- Support for dot product operator (`@`) with lowering to [`ckernel::matmul_block`](https://docs.tenstorrent.com/tt-metal/v0.55.0/tt-metalium/tt_metal/apis/kernel_apis/compute/matmul_block.html)
- Support for fusing matmul and certain elementwise operations
- Support lowering to `pack_tile_block`
- Support for `ttl.math.fill`, `ttl.math.reduce_sum`, `ttl.math.reduce_max`, and `ttl.math.transpose`
- Support for arbitrary sub-blocking including dot product K-dimension to allow maximizing L1 usage and reuse
- Support for `sin`, `cos`, `tan`, `asin`, `acos`, `atan` in `ttl.math`
- Support for L1 sharded tensors
- Support for tensors with BF8 data type
- SPMD support (`ttnn.open_mesh_device`)

### Simulator

- Track L1 space and number of DFBs usage and warn when exceeded
- Support for tensors with row-major layout
- Support for L1 sharded tensors

### Examples and documentat
dsl python kernels tt-lang simulator kernel-fusion
wormhole blackhole ttsim

tt-bh-linux

official★ featured
C++ · Apache-2.0 · 59⭐ ·
tt-bh-linux preview

Linux demo for the Tenstorrent Blackhole P100/P150 card RISC-V cores. Boot a real Linux kernel on the 16 high-performance RISC-V cores built into the Blackhole chip.

LATEST v0.11 2026-04-13T15:10:59Z Release notes ↗
4 previous releases
v0.10 2026-02-11T22:41:22Z
v0.9 2025-10-14T20:56:23Z
v0.5 2025-10-01T15:40:57Z
v0.4 2025-08-09T18:05:10Z
See all releases on GitHub ↗
linux risc-v blackhole bare-metal boot
blackhole

tt-llk

official
C++ · Apache-2.0 · 55⭐ · Jun 5, 2025

Tenstorrent Low-Level Kernels: the C++ library that directly programs the RISC-V cores inside each Tensix compute engine. TRISC0 (unpack), TRISC1 (math/FPU/SFPU), and TRISC2 (pack) are all programmed through this layer — it is the interface between TT-Metal kernel code and bare silicon.

tensix risc-v llk trisc brisc ncrisc low-level compute-engine
grayskull wormhole blackhole

ttnn-visualizer

official
TypeScript · Apache-2.0 · 54⭐ ·
ttnn-visualizer preview

Comprehensive tool for visualizing and analyzing model execution on Tenstorrent hardware. Interactive graphs, memory plots, tensor details, buffer overviews, operation flow graphs, and multi-instance support.

📦 Repo
LATEST v0.98.0 2026-08-12T21:18:05Z Release notes ↗
4 previous releases
v0.97.0 2026-08-05T17:46:17Z
v0.96.0 2026-07-29T21:50:25Z
v0.95.1 2026-07-24T16:36:00Z
v0.95.0 2026-07-22T19:41:33Z
See all releases on GitHub ↗
visualization profiling memory operations graphs
wormhole blackhole

WallaBMC

official
C · Apache-2.0 · 52⭐ ·
WallaBMC preview

Lightweight BMC (Baseboard Management Controller) for STM32 and similar MCUs, with Web UI, Redfish API, and HTTPS support. Built on Zephyr RTOS. Used in Tenstorrent systems.

📦 Repo
bmc stm32 redfish zephyr embedded

TT-Studio

official
TypeScript · Apache-2.0 · 49⭐ ·

Web-based GUI for deploying and chatting with AI models on Tenstorrent hardware. Handles all technical setup automatically — deploy models, run inference, and explore capabilities through a simple browser interface.

📦 Repo
LATEST rc-v2.9.0 2026-07-28T18:09:27Z Release notes ↗
4 previous releases
v2.8.0 2026-06-30T22:36:25Z
v2.7.0 2026-06-16T14:57:48Z
v2.6.0 2026-05-20T17:04:32Z
v2.5.0 2026-04-20T17:03:48Z
See all releases on GitHub ↗
web-ui gui models chat deployment
wormhole blackhole quietbox

tt-umd

official
C++ · Apache-2.0 · 45⭐ ·

User-mode driver for Tenstorrent hardware. The userspace layer that sits between the kernel module and higher-level SDKs.

📦 Repo
# Changelog

## [0.9.5] - 2026-05-12

### Changed

Hardware hang detection for NOC and PCIe.
Tracy profiler integration with instrumentation across TLB, PCIe and sysmem paths.
DeviceProtocol ported to TTDevice, including DMA migration.
SocDescriptor split into static (SocArchDescriptor) and runtime parts.
LITERAL coordinate system in CoreCoord.
Multicast to all TENSIX cores.
SMN support.
SWEmuleChip software emulation chip and Quasar simulation support (incl. 4GB TLB).
Unified UmdException/UMD_ASSERT/UMD_THROW error handling across the codebase.

## [0.9.4] - 2026-03-18

### Changed

TopologyDiscoveryOptions refactoring.
TopologyDiscoveryOption to retrain ETH links on 6u.
TLBs for TTsim.
DRAM retrain support.
DeviceProtocol changes.
Simulator in TTDevice changes.
ETH heartbeat check.

## [0.9.3] - 2026-02-24

### Changed

Sigbus safe read write API.
Remove 4U related code.
Implement BH SPI as well, so full SPI support.
P150 expects harvested cores.
TT_VISIBLE_DEVICES uses logical IDs.

## [0.9.2] - 2026-02-09

### Changed

SPI interface for Wormhole.
PCI BDF based sorting and filtering.
Multicast PCI DMA.
Support Blackhole loudbox.
Many code fixes and test enhancements.

## [0.9.1] - 2026-01-23

### Changed

Started publishing to pypi.

## [0.9.0] - 2026-01-23

### Changed

Warm reset notification and callback implementation.

## [0.8.6] - 2026-01-20

### Changed

Make predicting ETH FW from CMFW optional in TopologyDiscovery.

## [0.8.4] - 2026-01-16

### Changed

Use older manylinux image

## [0.8.3] - 2026-01-15

### Changed

Reverted remote discovery issue

## [0.8.2] - 2026-01-15

### Changed

Support warm reset without secondary bus reset.
Expose subsystem vendor id.

## [0.8.1] - 2026-01-15

### Changed

Support dma functions on TTDevice layer

## [0.8.0] - 2026-01-14

### Changed

Many functional fixes and minor changes.
Final fixes needed for integration into tt-smi.
Also contains adjustments needed for integration into exalens.

## [0.7.0] - 2025-11-29

### Changed

Changed to a more generic arc_msg API.

## [0.6.0] - 2025-11-24

### Changed

Change the usage of TLBs such that KMD is in control of TLB allocation instead of UMD.
TLBs are now allocated using KMD's dedicated API.

## [0.5.3] - 2025-11-14

### Changed

Added generation of .deb and .rpm packages.
Added three separate packages (runtime, development and python).

## [0.5.1] - 2025-11-12

### Changed

Manylinux builds and Pypi test publishing.
Many smaller fixes and improvements.

## [0.4.0] - 2025-10-18

### Changed

Removed old type names.

## [0.3.0] - 2025-10-17

### Changed

Many smaller fixes and improvements.
TTsim support improvements.
JTAG support improvement.
Fixing CMake install path.
Further work on integrating new KMD TLBs.

## [0.2.0] - 2025-09-15

### Changed

A couple of smaller fixes and improvements, including L2CPU harvesting, fixes for new FW. Better TTSim support. Further JTAG support.
Introduced new soft reset API.
Introduced lite fabric initial version.
user-mode-driver umd hardware-interface
grayskull wormhole blackhole

tt-system-firmware

official
C · Apache-2.0 · 42⭐ ·
tt-system-firmware preview

System firmware for Tenstorrent hardware. Low-level system initialization and control firmware that runs on-device.

firmware system embedded
wormhole blackhole

polaris

official
Python · Apache-2.0 · 39⭐ ·

A high-level AI simulator from Tenstorrent for modeling and exploring AI accelerator and workload performance.

📦 Repo
LATEST pre_perfmodel_merge 2025-09-19T18:16:27Z Release notes ↗
See all releases on GitHub ↗
simulator performance modeling architecture

luwen

official
Rust · Apache-2.0 · 34⭐ ·

Tenstorrent system interface library written in Rust. Low-level Rust bindings for communicating with and managing TT hardware.

📦 Repo
LATEST bh-mod-v2.1.0 2026-08-14T16:22:37Z Release notes ↗
4 previous releases
v0.9.0 2026-08-12T14:56:20Z
bh-mod-v2.0.0 2026-08-07T14:05:56Z
bh-mod-v1.1.0 2026-07-28T19:10:07Z
bh-mod-v1.0.0 2026-06-30T20:39:53Z
See all releases on GitHub ↗
rust system-interface low-level bindings
grayskull wormhole blackhole

tt-tvm

official
Python · Apache-2.0 · 31⭐ ·

TVM for Tenstorrent ASICs. Brings the Apache TVM compiler stack to Tenstorrent hardware, enabling model compilation from TensorFlow, PyTorch, ONNX, and more.

📦 Repo
tvm compiler tensorflow onnx
grayskull wormhole blackhole

tensix-isa-simulator

official
C++ · Apache-2.0 · 29⭐ ·

ISA-level simulator for the Tensix compute engine. Simulates the matrix, vector, and scalar units inside each Tensix core.

📦 Repo
tensix isa simulator compute-engine
ttsim

tt-torch

official
Python · Apache-2.0 · 26⭐ ·

Frontend integration for PyTorch with tt-mlir. Compile PyTorch models directly to Tenstorrent hardware via torch.compile integration.

LATEST 0.4.0 2025-09-29T22:23:47Z Release notes ↗
5 previous releases
0.5.0.dev20251008pre 2025-10-08T05:36:07Z
0.5.0.dev20251007pre 2025-10-07T04:22:29Z
0.5.0.dev20251006pre 2025-10-06T04:21:23Z
0.5.0.dev20251005pre 2025-10-05T04:38:19Z
0.5.0.dev20251004pre 2025-10-04T04:22:15Z
See all releases on GitHub ↗
pytorch torch-compile frontend
wormhole blackhole

tt-firmware

official
Apache-2.0 · 24⭐ ·

Tenstorrent firmware repository. Board management and control firmware for Tenstorrent accelerator cards.

📦 Repo
LATEST v19.6.0 2026-02-20T16:53:34Z Release notes ↗
4 previous releases
v19.5.0 2026-02-04T18:22:15Z
v19.4.2 2026-01-05T23:32:14Z
v19.4.1 2025-12-19T17:06:37Z
v19.4.0 2025-12-16T05:38:23Z
See all releases on GitHub ↗
firmware bmc board-management
wormhole blackhole

tt-installer

official
Shell · Apache-2.0 · 24⭐ ·

Install the complete Tenstorrent software stack with one command. Handles drivers, firmware, Python environment, and SDK setup automatically.

LATEST v3.5.3 2026-08-13T17:40:44Z Release notes ↗
4 previous releases
v3.5.2 2026-07-30T18:00:15Z
v3.5.1 2026-07-28T20:27:01Z
v3.5.0 2026-07-23T18:21:11Z
v3.4.0 2026-07-13T20:06:51Z
See all releases on GitHub ↗
installation setup one-command getting-started
wormhole blackhole

tt-exalens

official
Python · Apache-2.0 · 21⭐ ·

Low-level hardware debugger for Tenstorrent devices. Inspect register state, memory contents, and kernel execution at the hardware level.

📦 Repo
debugger low-level hardware registers
wormhole blackhole

tt-blacksmith

official
Python · Apache-2.0 · 16⭐ ·

Optimized training recipes for a variety of ML models on Tenstorrent hardware, powered by the TT-Forge compiler stack. Reference implementations for fine-tuning and training from scratch.

training fine-tuning recipes pytorch
wormhole blackhole

tt-topology

official
Python · Apache-2.0 · 16⭐ ·
tt-topology preview

Configure Ethernet routing on multi-card Tenstorrent systems. Flash NB cards to use specific ETH routing configurations for scale-out deployments.

📦 Repo
# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## 1.2.11 - 17/06/2025

### Updated

- Updated mesh coord generation to be connection type agnostic
- Added failure and exit if mesh type detected, but not enough connections
- Added warning in README about lack of supoort for BH and 6U boards

## 1.2.10 - 05/06/2025

### Updated

- Bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0

## 1.2.9 - 30/05/2025

### Updated

- Bug fix for https://github.com/tenstorrent/tt-topology/issues/39. Now the tool will use a DFS longest path to determine a linear layout if its not a fully connected graph.
- Updated initial device detection - now it needs full noc access for octopus and list options

## 1.2.8 - 08/05/2025

### Updated

- Fixed issue where tool would fail when PCI interfaces don't start from ID 0
- Now using actual PCI interface IDs from devices instead of assuming sequential numbering

## 1.2.7 - 07/05/2025

### Updated

- Use tools-common 1.4.15
- Use type checking in octopus reset

## 1.2.6 - 05/05/2025

### Updated

- Bug fix: added "ignore-eth" flag to first chip detect to avoid eth training loops forever and truly detect pcie only chips
- Chore: bumped luwen

## 1.2.5 - 15/04/2025

### Updated

- When flashing to isolated mode, we now flash the WH ethernet ports to a disabled state,
  in order to prevent their use.

## 1.2.4 - 02/04/2025

### Updated

- You can now run `tt-topology -l isolated` to flash cards to the default (non-connected) state
- Users are now warned about missing or loose cables

## 1.2.3 - 21/03/2025

### Fixed

- Bumped luwen (0.6.2 -> 0.6.3) to include eth version check bug for TG setup

## 1.2.2 - 13/03/2025

### Fixed

- Bumped luwen version to make it more robust against eth fw updates

## 1.2.1 - 13/03/2025

### Fixed

- Moved the spi reads after the reset to increase stability during M3 L2R copy
- Bumped luwen version

## 1.2.0 - 06/03/2025

### Fixed

- Updated how local eth board info is calculated to make it agnostic to eth fw version
- bumped tt-tools-common version
- Added traceback printing when catching exceptions in main.

## 1.1.5 - 14/05/2024

### Updated

- Bumped luwen (0.3.8) and tt_tools_common (1.4.3) lib versions
- Removed unused python libraries

## 1.1.4 - 25/03/2024

### Fixed
- Changed detect_chips with detect_chips_with_callback to enable detailed debug info.

## 1.1.3 - 22/03/2024

### Fixed
- Bumped tt-tools-common version to avoid pip discrepancy.

## 1.1.2 - 22/03/2024

### Fixed
- Fixed command line bug when no args are provided.

## 1.1.1 - 21/03/2024

### Fixed
- Fixed reference to pyluwen lib

## 1.1.0 - 12/03/2024

### Added
- Octopus Configuration (4 n150s connected to 1 galaxy)


## 1.0.2 - 12/03/2024

### Fixed
- Dependency bug with tt_tools
topology ethernet multi-card routing
wormhole blackhole

tt-npe

official
C++ · Apache-2.0 · 15⭐ ·
tt-npe preview

Network-on-chip Performance Estimator for Tenstorrent Tensix-based devices. Model and estimate NoC utilization before running kernels on hardware.

📦 Repo
noc performance estimator profiling
wormhole blackhole

tt-flash

official
Python · Apache-2.0 · 14⭐ ·

Tenstorrent firmware update utility. Flash new firmware onto Tenstorrent accelerator cards from the command line.

📦 Repo
LATEST v3.10.0 2026-06-23T19:59:38Z Release notes ↗
4 previous releases
v3.11.0-rc.1pre 2026-08-14T15:54:35Z
v3.9.0 2026-06-17T07:52:40Z
v3.8.0 2026-06-01T18:04:27Z
v3.7.0 2026-05-15T19:32:29Z
See all releases on GitHub ↗
# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## Unreleased

### Added

- `flash`: `--update-boot-images` writes the bundle's bootloader and recovery
  images even when the board already holds the same ones, for provisioning and
  board recovery.

### Changed

- `flash`: the boot-critical images (`cmfw`, `safeimg`, `safetail`, `failover`)
  and the ROM and failover descriptor tables are now left alone when the board
  already holds the same content, so a routine update no longer opens a
  power-loss window on the path by which a board boots at all. Whether an image
  is the same is decided by the SHA-256 and key hash that `imgtool` records in
  it, because signing is not reproducible: two builds of the same source differ
  only in the trailing signature. Pass `--update-boot-images` for the previous
  behaviour of writing them unconditionally.

### Fixed

- `flash`: a P300 chip running recovery firmware publishes no board id, so the
  pairing check filed it as not a P300, left its sibling alone in a group of
  one, and dropped both halves of the card from the flash list -- refusing the
  board because of the chip that most needed flashing. Such a chip is now
  identified by its PCI subsystem id, which carries the same UPI whatever the
  chip is running.
- `boot_fs`: `tt_boot_fs_fd.image_tag_str()` compared each `c_uint8` tag byte
  against the string `"\0"`, which never matched, so NUL padding was included in
  the decoded tag. Tags shorter than 8 bytes (e.g. `cmfw`) now compare correctly.
  This makes `read_tag` robust across the multi-table boot filesystem layout
  (ROM, failover, and mutable descriptor tables).

## 3.4.0 - 30/07/25

- Bump pyyaml 6.0.1 -> 6.0.2
- Improve error message formatting
- No longer have to use --force for flashing BH cards

## 3.3.5 - 03/07/25

- Bump luwen 0.7.3 -> 0.7.5

## 3.3.4 - 02/07/25

- Bump tt-tools-common 1.4.16 -> 1.4.17
- Bump luwen 0.6.4 -> 0.7.3

## 3.3.3 - 05/06/2025

- Bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0

## 3.3.2 - 14/05/2025

- Bump tt-tools-common version to latest

## 3.2.0 - 12/03/2025

### Updated

- luwen version bump to bring inline with tt-smi; provides stability fixes

## 3.1.3 - 06/03/2025

### Added

- luwen version bump to include bh arc init checks

## 3.1.2 - 28/02/2025

### Added

- Support for more BH cards: p100a, p150, and p150c

## 3.1.1 - 06/01/2025

### Updated

- Bumped luwen version to accomodate Maturin updates

## 3.1.0 - 29/10/2024

### Added

- Support for flashing the BH tt-boot-fs file format
- Bumped luwen version to 0.4.6 to allow resets when chip is inaccessible

## 3.0.2 - 17/10/2024

### Fixed
- Unbound variable when exception is thrown when getting current fw-version

## 3.0.1 - 16/10/2024

### Changed
- B
firmware-update flash utility
grayskull wormhole blackhole

SFPI

official
C++ · Apache-2.0 · 14⭐ ·

Tenstorrent SFPU programming interface — TT-enhanced RISC-V GCC and binutils plus header files for programming the Tensix SFPU (vector engine) from kernel code. The compiler toolchain underneath TT-Metalium's SFPU ops.

📦 Repo
LATEST 7.69.0 2026-08-12T21:12:00Z Release notes ↗
⚙ Requires PPA — setup instructions ↗
4 previous releases
7.69.0-clean-43896 2026-08-14T19:09:55Z
7.69.0-pred-43896 2026-08-13T17:34:21Z
7.69.0-lut-51346 2026-08-13T18:07:26Z
7.68.0 2026-08-07T13:57:54Z
See all releases on GitHub ↗
sfpu compiler-toolchain gcc riscv
wormhole blackhole

tt-example-apps

official
Jupyter Notebook · Apache-2.0 · 13⭐ ·

End-to-end AI applications running on Tenstorrent AI accelerators. Complete application examples from retrieval-augmented generation to image generation pipelines.

📦 Repo
rag applications end-to-end examples
wormhole blackhole

tt-forge-models

official
Python · Apache-2.0 · 13⭐ ·

A shared repository of model implementations used across TT-Forge frontends — a single source of truth for the models used in testing and benchmarking, rather than duplicating them across frontend repos.

📦 Repo
tt-forge models benchmarking testing inference

tt-perf-report

official
Python · Apache-2.0 · 11⭐ ·
tt-perf-report preview

Performance report analysis tool for Tenstorrent Metal operations — analyzes perf traces to surface throughput, bottlenecks, and optimization opportunities.

📦 Repo
performance profiling tt-metal analysis optimization

tt-vscode-toolkit

official
TypeScript · Apache-2.0 · 8⭐ · Dec 18, 2025
tt-vscode-toolkit preview

48 interactive lessons covering the full Tenstorrent developer path — from hardware detection to custom training — with click-to-run commands and hardware auto-detection. Available in VSCode and code-server.

LATEST v0.1.19 2026-07-16T21:17:06Z Release notes ↗
4 previous releases
v0.1.17 2026-07-13T15:57:03Z
v0.0.518 2026-06-30T17:57:43Z
v0.0.515 2026-06-23T21:23:15Z
v0.0.514 2026-06-23T20:15:47Z
See all releases on GitHub ↗
# Changelog

All notable changes to the TT-VSCode-Toolkit will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

---

## [0.1.19] - 2026-07-16
### Added
- **New advanced lesson "Monkeypatching TT-NN"** (order 16, after "Exploring TT-Metalium") — upgrade-safe, smallest-trace patching of TT-NN / TT-Metalium organized by developer goal (observe; fix a bug early; change a default; add something new; and a last-resort source-diff escape hatch), for the TT-QuietBox 2 case where `ttnn` is an installed package with no `~/tt-metal` source tree. Covers the env + import-order rule, the "patch to change behavior, wrap to add behavior" principle (citing Martin Chang's non-invasive `ttPseudoRowMajor` and upstream-first ggml backend), and an AI-agent verification recipe. Validated on p300c — all 19 lesson patterns exercised against real `ttnn` on a TT-QuietBox 2.
- **New reusable template `content/templates/monkeypatch/tt_patches.py`** — a dependency-free patch harness (save/restore `wrap` and `set_default`, `patched` context manager, `version_at_most` guard that zero-pads unequal-length versions, and a `verify` probe helper) with fail-loud missing-target detection, shipped with a self-contained hardware-free `test_tt_patches.py` and a usage README. The lesson embeds the full harness source in a collapsible section for transparency.
- **`tenstorrent.monkeypatch.copyHarness` command** — copies the harness folder into `~/tt-scratchpad/monkeypatch/`, replacing any prior copy (with confirmation) so it mirrors the shipped template, and offers an "Open tt_patches.py" follow-up. Wired into the lesson as a click-to-copy button.
- **`check:monkeypatch-drift` script** (wired into the pre-commit hook) fails if the `tt_patches.py` source embedded in the lesson diverges from the template file.

### Changed
- Excluded `.superpowers/` from the packaged `.vsix` via `.vscodeignore` (was shipping ~1.9 MB of session artifacts).

## [0.1.18] - 2026-07-13
### Fixed
- **PRD-246 — Jeremy's QB2 testing feedback on the first-inference lesson flow:**
  - `download-model` — fixed the broken "Step 3: Download the Model" skip link. The anchor pointed to `#step-3-download-qwen3-0-6b`, but the "Step 3: Download Qwen3-0.6B" heading slugs to `#step-3-download-qwen3-06b` (the `.` in `0.6B` is dropped, not turned into a hyphen).
  - `download-model` — consolidated the repeated, scattered Hugging Face auth flow. Removed the standalone "Already Authenticated?" pre-check and folded the `hf auth whoami` check into Step 2, so authentication reads as a single sequence (set token → check → log in) instead of appearing in multiple places.
  - `hardware-detection` — marked "Check 4: Device Reset" as optional; it is a recovery action, not part of normal detection, and a healthy device never needs it.
  - `tt-installer` — removed the redundant `tt-smi` hardwar
vscode lessons interactive getting-started code-server
wormhole blackhole quietbox ttsim

tt-tools-common

official
Python · Apache-2.0 · 7⭐ ·

Shared helper library of common utilities used across Tenstorrent system tools such as tt-smi, tt-flash, and tt-topology. A dependency rather than a standalone tool.

📦 Repo
# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## 1.4.17 - 02/07/2025

### Changed
- Loosened requirements on pyproject.toml to make it more compatible in different venvs

## 1.4.15 - 05/05/2025

### Changed
- parse\_reset\_json now returns a ResetInput with stricter typing

## 1.4.14 - 04/02/2025

### Added
- New flags in reset config file generation to disable sw\_version reporting

## 1.4.13 - 23/1/2025

- Removed nr\_hugepages count from compatibility, as hugepages allocation is tricky
  and deserves its own widget elsewhere.

## 1.4.12 - 16/1/2025

- Added TTHostCompatibilityMenu to replace Host Info and Compatibility boxes
- Added a count of nr\_hugepages to the TTHostCompatibilityMenu

## 1.4.11 - 30/12/2025

- Updated Luwen version to fix Maturin issue

## 1.4.10 - 16/12/2024

### Changed
- detect\_chips\_with\_callback now takes a print\_status arg

## 1.4.9 - 11/12/2024

### Changed
- A failed reset now results in a fail exit code on BH

## 1.4.8 - 11/10/2024

### Changed
- Updated reset completion logic to handle the case where the bmfw needs to upgrade itself

## 1.4.7 - 11/10/2024

### Added
- Implemented m3 reset option for Blackhole

### Fixed
- Fixed crash during driver version dection when the "extraversion" field is used
    - i.e. 1.28-bh

## 1.4.6 - 17/07/2024

### Added
- Reset support of Blackhole

## 1.4.5 - 11/07/2024

### Added
- Bump pyluwen library version (v0.3.8 -> v0.3.11)
- Moved pyluwen v0.3.11 to optional dependencies in pyproject.toml

## 1.4.4 - 21/06/2024

### Added
- Version bump of python dependencies in pyproject.toml (dependabot)
    - requests (2.31.0 -> 2.32.0)
    - tqdm (4.66.1 -> 4.66.3)
- Pydantic library version bump (1.* -> >=1.2) to resolve: [TT-SMI issue #27](https://github.com/tenstorrent/tt-smi/issues/27)

## 1.4.3 - 14/05/2024

### Added
- Arm platform check and warning for WH device resets in compatibility menu
- Added check for WH device init after reset and prompt user to reboot host if chips are still non recoverable
- Bumped textual (0.59.0) and luwen (0.3.8) lib versions

## 1.4.2 - 04/04/2024

### Added
- Added "silent" flag to WH and GS resets to make them more versatile for use in other tools

## 1.4.1 - 22/03/2024

### Fixed
- removed pyluwen version to avoid dependency issues in other repos

## 1.4.0 - 19/03/2024

### Added
- detect_device_fallible that will provide feedback about chip state during init

### Fixed
- Update min driver version to 1.26 to perform lds reset
- Reset config file uses dev/tenstorrent id
- Catch JSON errors in reset config parsing
- Make nested dirs when initializing reset config path

## 1.3.0 - 06/03/2024

### Added
- Migrated GS Tensix reset to tools_common
- Migrated all related GS data files
- Functions to fetch arc and eth fw versions from telemetry
library tooling shared-utilities

vllm-tt-plugin

official
Python · Apache-2.0 · 6⭐ ·

Tenstorrent backend for vLLM, built on vLLM's standard plugin mechanism — install it alongside vLLM and TT hardware registers itself as a platform whenever `ttnn` is importable. Self-contained: model registration, platform detection, scheduling, worker execution, model loading, async decode, and data-parallel/multi-lane execution all live in the plugin, so nothing Tenstorrent-specific has to land in vLLM core.

📦 Repo
vllm serving inference llm openai-compatible plugin
wormhole blackhole

tt-toplike

official
Rust · Apache-2.0 · 6⭐ ·
tt-toplike preview

A vibrant htop-style visualizer for Tenstorrent hardware written in Rust. Real-time process and utilization view for TT accelerators.

# Changelog

The **canonical, complete release log lives in [`debian/changelog`](debian/changelog)** —
that's the file the `.deb` packages are built from and where every release is
recorded in full. This file is a friendly pointer plus a summary of the most
recent releases; it deliberately does not duplicate the whole history.

To see everything:

```bash
less debian/changelog          # full history
git tag                        # released versions
```

## Recent releases

### 0.8.0
- **Fix: per-device data was shuffled on multi-card boxes.** The sysfs backend
  numbered devices in raw `readdir` order while `tt-smi` orders by PCI bus id, so
  the default (hybrid) backend's join attributed every card's SMBUS data to the
  wrong card — DDR status, GDDR temps, ECC counters, thermal trips, fan and clocks
  all landed on a neighbour. Discovery is now sorted by bus id, **and the join
  itself keys on the PCI bus id** rather than on list position, so attribution
  also holds when `tt-smi` enumerates fewer cards than hwmon does (a busy card, a
  card that failed to enumerate, `--devices` filtering, hotplug). If you ran the
  default backend on more than one card, what you saw was mixed up.
- **The safe backend got a lot less telemetry-poor.** Alongside hwmon, the sysfs
  path now reads tt-kmd's class-attribute directory
  (`/sys/class/tenstorrent/tenstorrent!N/`): clock frequencies (AICLK/AXICLK/ARCCLK),
  ARC firmware heartbeat, the real board SKU, firmware bundle version, board serial
  and thermal-trip count. `tt_card_type` **replaces** the ~1.2 s `tt-smi -s` startup
  probe on modern drivers. Needs tt-kmd ≥ 2.7; older drivers keep the old behavior.
- **Live PCIe bandwidth.** tt-kmd's `pcie_perf_counters/` are folded into in/out
  directions and differentiated between ~1 Hz samples, so Insights shows a PCIe row
  with link geometry (e.g. `Gen4 x4`) and live ▼/▲ rates. A counter set that can't
  be read at all reports nothing rather than a confident 0 B/s. Sysfs and hybrid
  backends only.
- **Better hwmon reads**: sensors are picked by their `*_label` (so the ASIC temp
  sensor is used, not whichever has the lowest index), the fan sensor is read, and
  each sensor's own `*_max` gives real per-board limits (125 W / 500 A / 90 °C on a
  p300c) instead of hardcoded 300 W / 105 °C — a limit always comes from the same
  sensor as its reading. Limits need tt-kmd ≥ 2.9. The heavier reads (class attrs,
  PCIe counters) are sampled at ~1 Hz instead of on every render frame.
- **New Insights sidebar rows**: Current (with TDC limit), Board power (tt-smi 6.x),
  PCIe, GDDR ECC (only when non-zero — uncorrectable errors in red; they were parsed
  but never shown anywhere before), and thermal trips (only when non-zero).
- **tt-smi upkeep**: full `board_info` parsing (PCIe generation/width, tolerant of
  tt-smi's number-vs-string drift); the Fan row no longer stays blank on cards with a
  spinning fan (live tt-smi emits `FAN_SPEED: "0x0"` next to a real `FAN_RPM`,
monitoring htop rust real-time
wormhole blackhole

tt-rpm

official
C++ · Apache-2.0 · 6⭐ ·

Cycle-level, execution-driven RISC-V CPU performance model built on Sparta (MAP) with Whisper supplying functional execution, so it runs real ELF binaries — CoreMark, Dhrystone — to completion. The pipeline is YAML-configurable across in-order/out-of-order execution, issue policy, execute granularity, write-port arbitration, and bypass paths, with a modeled L1 I$/D$ plus optional unified L2, per-unit logging, stats reports, and Konata pipeline visualization.

📦 Repo
risc-v performance-model simulation sparta whisper microarchitecture

tt-low-level-documentation

official
5⭐ ·

Documentation for the low-level layer of tt-metal: compute LLK APIs and data movement APIs. The data movement side covers the NOC and overlay on Wormhole and Blackhole; the compute side covers Tensix hardware and expected usage of the LLK APIs. Aimed at op and model writers who need to know what the APIs do and how the hardware behaves underneath them.

📦 Repo
llk data-movement noc tensix tt-metal documentation
wormhole blackhole

tt-system-tools

official
Shell · Apache-2.0 · 5⭐ ·

System setup and support utilities for Tenstorrent hardware — hugepages-setup configures the 1GB hugepages TT ASICs need, and tt-oops collects diagnostic data for troubleshooting. Ships as the tenstorrent-tools deb/rpm.

📦 Repo
LATEST v1.4.1 2025-12-08T17:23:48Z Release notes ↗
4 previous releases
v1.4.0 2025-09-04T20:40:58Z
v1.3.1 2025-05-02T18:17:15Z
upstream/v1.2 2025-04-04T21:10:41Z
upstream/1.1 2024-07-17T18:46:30Z
See all releases on GitHub ↗
hugepages system-setup diagnostics
grayskull wormhole blackhole

tt-emule

official
C++ · Apache-2.0 · 4⭐ ·

A C++ software emulator of the Tenstorrent device-level kernel and host APIs. Run tt-metal kernel and host code on a standard x86-64 Linux machine — no Tenstorrent hardware required.

📦 Repo
emulator tt-metal no-hardware testing kernels

tt-CableGen

official
JavaScript · Apache-2.0 · 4⭐ ·
tt-CableGen preview

Network cabling visualizer for Tenstorrent scale-out deployments: describe a target topology and it generates and renders how to physically cable multiple Wormhole or Blackhole systems together. Works in a physical-deployment mode with racking information and a logical-hierarchy mode for clustering/pod groupings, with topology import/export and a Docker deployment path.

📦 Repo
cabling scale-out topology multi-host visualization
wormhole blackhole galaxy

tt-local-generator

official
Python · Apache-2.0 · 3⭐ ·
tt-local-generator preview

Generate infinite videos and images (and imaginative prompts to inspire them) on Tenstorrent's Quietbox 2. Fully local generative media pipeline.

video-generation image-generation quietbox generative
quietbox

tt-kernel

official
Python · Apache-2.0 · 3⭐ ·

Distributes models over the Hugging Face Hub and serves them on Tenstorrent hardware — `tt-kernel serve <namespace>/<model>` pulls a bundle, registers it with the Tenstorrent vLLM plugin, and launches an OpenAI-compatible server. A vLLM bundle ships only adapter code and metadata (weights stay referenced by HF repo id), while legacy kernel-cache bundles package precompiled tt-metal kernel binaries so a model's first run is a cache hit instead of a slow JIT recompile. Explicitly experimental — the bundle format and APIs may change without notice.

📦 Repo
kernel-cache huggingface package-manager vllm tt-metal serving
wormhole blackhole

tt-burnin

official
Python · Apache-2.0 · 3⭐ ·

Command-line utility that runs a high power-consumption workload on Tenstorrent devices — used for chip testing, burn-in, and validating a system's power delivery and cooling under sustained load.

📦 Repo
# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## 0.2.2 - 31/07/2024

### Added
- Added glx reset support
- Threaded start and end of burnin to increase burnin speed
- Added prints to indicate which chip we are currently running on
- Added support for bh harvesting

## 0.2.1 - 16/01/2024

### Bug fix
- Fix for https://github.com/tenstorrent/tt-burnin/issues/6
- BH reports asic temperature as a signed 16_16 int unlike GS and WH
- Added missing support to report BH asic temperatre

## 0.2.0 - 29/10/2024

### Added
- BH burnin support

## 0.1.1 - 14/05/2024

### Updated

- Bumped luwen (0.3.8) and tt_tools_common (1.4.3) lib versions

## 0.1.0 - 04/04/2024

First release of opensource tt-burnin

### Added
- GS and WH burnin support
burn-in stress-test power hardware-validation
wormhole blackhole

tt-animatediff

official
Python · Apache-2.0 ·
tt-animatediff preview

Generates short, temporally coherent animated GIFs using the AnimateDiff model on Tenstorrent hardware. Phase 1 runs the correct SD 1.4 + MotionAdapter architecture on CPU; Phase 2 accelerates spatial denoising on Blackhole using the TTNN UNet. Produces vibrant 8-frame animations in ~15 s/frame on a P300C.

LATEST v0.9.0 2026-06-22T19:56:31Z Release notes ↗
2 previous releases
v0.6.0 2026-06-10T22:16:43Z
v0.1.0 2026-06-04T22:31:14Z
See all releases on GitHub ↗
animatediff video-generation stable-diffusion diffusion gif blackhole
blackhole

Cloud-Native Support

official

Official documentation hub for running Tenstorrent accelerators on Kubernetes. Centers on tt-operator (the umbrella Helm chart) and covers Node Feature Discovery, kernel-mode driver (tt-kmd) management, firmware flashing, Prometheus telemetry, Fabric Manager topology resolution, Dynamic Resource Allocation, and multi-node scheduling via JobSet and PMIx.

kubernetes cloud-native helm tt-operator orchestration documentation
wormhole blackhole

TT Console

official★ featured

Browser-based cloud console for exploring AI on Tenstorrent hardware. Run LLM inference, image and video generation, and browse the supported model catalog in-browser — backed by Tenstorrent accelerators. Cloud hardware access and advanced workflows (deployments, agents) available in staged rollout.

cloud console inference playground llm image-generation video-generation demo
wormhole blackhole

TenGEMM

official
TypeScript · MIT ·

Tensix GEMM performance estimator and visualizer. A React app that models matrix-multiplication workloads on a Tensix core, estimates the resulting performance, and shows how the work maps onto the hardware — useful for reasoning about a matmul's shape and fidelity choices before writing the kernel.

📦 Repo
matmul gemm visualization performance estimator tensix react

tt-cli

official
Python · Apache-2.0 ·

Single entry point to the Tenstorrent software stack: `tt update` converges a machine onto the CI-tested "golden" version set, `tt device` covers status/info/reset, and `tt model`/`tt serve` pull weights and bring up tt-inference-server. Commands either run natively or delegate to tt-smi, tt-flash, and tt-installer behind a stable interface, with `--json` output and documented exit codes on every command. Early prototype — the README labels it an internal prototype and several subcommands are still stubs.

📦 Repo
cli device-management model-serving tt-smi prototype
blackhole quietbox

TT-QuietBox 2 Guide

official

Official setup and onboarding guide for the TT-QuietBox 2 — a compact, liquid-cooled AI workstation with four Blackhole accelerators, an AMD Ryzen CPU, 256GB RAM, and 4TB NVMe. Covers hardware specs, first-boot setup, and hands-on learning paths for running pre-loaded models like Qwen3-32B and serving text, image, video, and speech models via tt-inference-server.

quietbox blackhole workstation setup getting-started documentation
blackhole quietbox

ttsim-qemu

official ⑂ qemu/qemu
C · GPL-2.0 ·

Tenstorrent's fork of QEMU that provides the full-system emulation layer behind ttsim. Models the RISC-V cores and system devices of Wormhole and Blackhole so TT-Metalium workloads can boot and run without physical silicon.

📦 Repo
simulator qemu full-system emulation no-hardware wormhole blackhole
ttsim

tt-bio

affiliated★ featured
by moritztng · Python · MIT · 118⭐ · Jan 31, 2026

Boltz-2 biomolecular model for drug discovery on Tenstorrent Blackhole. Supports single-card and multi-card configurations — QuietBox (4×) and Galaxy (32×). Approaches physics-based FEP accuracy at 1000× the speed.

LATEST v0.6.2 2026-08-07T21:33:46Z Release notes ↗
4 previous releases
v0.6.1 2026-08-07T00:28:29Z
v0.6.0 2026-08-01T07:12:48Z
v0.5.0 2026-07-27T14:14:13Z
v0.4.0 2026-07-26T18:47:55Z
See all releases on GitHub ↗
# Changelog

All notable changes to TT-Bio are recorded here. Versioning is [SemVer](https://semver.org);
releases are cut from a commit that has passed the on-hardware test suite (see `RELEASING.md`).

## [Unreleased]

### Changed

- RFdiffusion3 ships both fused bias kernels on by default (`881704d2`). The sparse
  attention bias is built in one pass instead of a poke walk (5.83x at the op,
  `703d12a1`) and the whole score+bias chain is one kernel (4.42x at the op,
  `923a9396`); both learned multiplicity batching, worth 6.26x at batch 2
  (`fa7246da`). Every step is bit-exact: the fold A/B legs land byte-identical
  designs at +12.36 %, +5.10 %, +6.83 % and +4.65 % on ms/step (`583961c4`,
  `ee4a8980`, `64a14e68`, `599d81ff`). The published throughput table in
  `docs/rfd3-design.md` was regenerated with them on (`5123065e`).
- Triangle attention runs the q-split at or below 1024 padded tokens and gates it
  off above (`063f89db`).

### Fixed

- `--trace` with Protenix-v2 or OpenDDE silently returned wrong structures for every
  target after the first when one process folded several targets of the same size: the
  captured trace was keyed on shape alone and replayed the first target's conditioning.
  The trace is now re-captured when the conditioning changes. Regression gate:
  `scripts/trace_multitarget_parity.py` (two same-size targets, one process, trace on
  vs off, byte-identical CIFs required). Boltz-2 and BoltzGen were not affected: their
  predict path resets the trace cache between targets.
- Protenix-v2 crashed on targets between 385 and 506 residues: the h=1.5 normed pair
  tensor was held past its last use (`142e0109`). ESMFold2 hit the same class at large
  targets and now frees the pair-conditioning intermediates rather than row-tiling them
  (`08565983`). OpenDDE uploads `z_struct` as one allocation into an intact hole
  (`28a91107`) and frees the expander's row chunks after the loop (`d2ad024b`).
- A clean `pip install` was missing kernel sources: the wheel and sdist now ship every
  file under `tt_bio/kernels/` through recursive globs, so a new kernel directory cannot
  drop out again (`cb3ef828`, `baa6ad0a`).
- On a tt-metal built from source, tt-bio could not find the fabric mesh-graph
  descriptor (`ee73d9a4`) or the `generic_op` kernel sources (`e2fb610a`), so a lone
  Blackhole P300 chip would not open and the fused kernels would not build.
- A worker that died of an uncaught exception reported silence (`a5921d2d`), and a
  process holding a card could outlive whoever spawned it (`3bd84f04`, `26a8c085`).
- Protenix-v2 falls back to ttnn's own matmul planner when a tuned config clashes
  instead of failing the fold (`5a207fee`).

### Performance

- OpenFold3 at 512 residues: 51.19 -> 44.535 s on one Blackhole p150a, from running
  TriangleAttention's fp32-softmax tail height-sharded in L1 rather than
  DRAM-interleaved. Bit-exact — the same CIF digest (da9b4ed68f8c0405) and plDDT as the
  control arm, and it holds at 768
drug-discovery blackhole inference biology multi-card
blackhole quietbox galaxy

grayskull-attention

affiliated
by moritztng · TeX · MIT · 38⭐ ·

FlashAttention-style attention kernel implemented entirely in on-chip SRAM on the Tenstorrent Grayskull chip using TT-Metalium. Pioneering work in low-level attention on TT hardware.

📦 Repo
attention grayskull metalium sram kernel
grayskull

tt-atom

affiliated
by moritztng · Python · MIT · 26⭐ ·
tt-atom preview

Meta's UMA interatomic potential running on Tenstorrent Blackhole — energy, forces, and stress for molecules and periodic materials behind an ASE calculator. Its per-edge Wigner rotation runs as a custom tt-metal kernel for a highest-performance uma-s build.

📦 Repo
LATEST v0.2.1 2026-07-20T11:06:13Z Release notes ↗
2 previous releases
v0.2.0 2026-07-11T07:47:26Z
v0.1.0 2026-07-08T19:08:02Z
See all releases on GitHub ↗
# Changelog

All notable changes to TT-Atom are recorded here. Versioning is [SemVer](https://semver.org);
releases are cut only from a commit that has passed the on-hardware release gate — accuracy
parity, no OOM across the supported size range, no perf or UX regression, and a clean install
smoke (see `RELEASING.md`).

## Unreleased

### Added
- **Multi-card data-parallel fan-out for `tt-atom run`**: pass several structure files with
  `--devices 0,1,...` and each card runs a full `Calculator` + relax/MD loop (or the
  single-point energy default) for its shard of structures — the high-throughput
  virtual-screening path. Per-structure results come back in input order, bit-exact vs the
  single-card path (`scripts/_multicard_sim_parity.py`); without `--devices`, multiple
  structures run one after another on one card. `--out` is a directory in batch mode and each
  written geometry carries its energy and forces. The new `tt_atom.batch.MultiCardSim` pool
  backs the CLI and is usable directly; `scripts/multicard_sim_scaling.py` measures the
  throughput scaling.
- Orb's batched path (`evaluate_batch`) now applies ZBL pair repulsion union-wide through one
  autograd pass, matching the per-system path at short contact.

### Fixed
- `tt-atom run a.xyz b.xyz --relax --devices 0` with two **different-composition** structures no
  longer crashes the worker. The multicard worker builds one UMA `Calculator` per reduced
  composition and used to call `open_device` once per `Calculator`, so a second composition opened
  the same card a second time in one process (`TT_FATAL: No MetalContext instance for context_id N`).
  The worker now opens its device once and reuses it across every `Calculator` it builds; the
  `Calculator.close()` it owns never closes a device it didn't open.
- `tt-atom run` (multicard) now exits non-zero when any structure fails. The worker has always
  caught per-structure errors and returned the other structures' results, but the CLI used to
  exit 0 regardless, so a failed structure silently dropped its output. It now reports which
  structures failed and exits non-zero while still writing the ones that succeeded.
- Built wheels now include both weight exporters, so automatic UMA and Orb cache misses work
  outside a source checkout.
- Fresh UMA and Orb cache misses can download their checkpoints again; explicit
  `HF_HUB_OFFLINE=1` still enforces offline use. Concurrent exports now use separate sidecars.
- Release mode now blocks every missing fixture, baseline, required op, and model-family OOM row.
  `--allow-gaps` remains available for development diagnostics.
- Release and UX subprocesses always open logical device 0 after `TT_VISIBLE_DEVICES` selects the
  physical card.
- Custom-op validation now rejects invalid gate modes and shapes, and program-cache keys include
  every operand layout that affects compiled accessors.
- The silicon-melt example now checks the exact-cutoff neighbour graph every step and recaptures
  only when
molecular-dynamics interatomic-potential mlip uma ase inference custom-kernel
blackhole

tt-lang-models

affiliated
by zoecarver · Python · 7⭐ ·
tt-lang-models preview

A growing collection of models that use tt-lang for some or all of their implementation. Reference implementations for bringing modern models to the tt-lang DSL.

📦 Repo
tt-lang models dsl reference

tt-zork-and-more

affiliated ⑂ historicalsource/zork1★ featured
by tsingletaryTT · Python · 2⭐ ·
tt-zork-and-more preview

A Tenstorrent fork of Infocom's Zork I (and more!), running a Z-machine interpreter at least four different ways on TT hardware. The most fun you can have with an AI accelerator.

zork z-machine interactive-fiction demo fun

tt-qb-lights

affiliated
by tsingletaryTT · Rust · 2⭐ ·

Sync your Tenstorrent Quietbox's RGB lighting to accelerator utilization status. Visual feedback for hardware activity in real time.

📦 Repo
quietbox rgb hardware fun
quietbox

diamond

affiliated ⑂ eloialonso/diamond
by zoecarver · Python · 1⭐ ·
diamond preview

DIAMOND: Atari game-playing agent implemented on Tenstorrent hardware via tt-lang. Diffusion-based world model for reinforcement learning.

atari reinforcement-learning world-model tt-lang

gemma4

affiliated
by zoecarver · Python · 1⭐ ·

Gemma 4 language model implemented in tt-lang (e4b variant) for direct execution on Tenstorrent hardware.

📦 Repo
gemma llm tt-lang inference
blackhole

open-oasis

affiliated ⑂ etched-ai/open-oasis
by zoecarver · Python · 1⭐ ·

tt-lang inference script for Oasis 500M — an interactive video world model running on Tenstorrent hardware via the tt-lang DSL.

📦 Repo
video world-model oasis tt-lang inference
blackhole

tt-model-runner

affiliated
by tsingletaryTT · Python · 1⭐ ·

Discover, load, and benchmark models with a GUI and TUI for tt-inference-server. Makes exploring available models on Tenstorrent hardware as easy as browsing a catalog.

📦 Repo
gui tui models inference benchmark
wormhole blackhole quietbox

tt-claw

affiliated
by tsingletaryTT · Shell ·

A Tenstorrent-powered claw machine that rewards players with real prizes. The QuietBox 2 runs local AI inference to act as an agent controlling the claw hardware — the OpenClaw AI assistant lesson builds directly on this project.

claw-machine agents hardware quietbox physical on-device
quietbox

Local AI Agents on Tenstorrent

affiliated★ featured
by ·

Three agentic projects running fully on-device: local AI agents on QuietBox 2, a coding assistant powered by Aider against a local inference server, and the OpenClaw AI assistant on QuietBox 2. No cloud APIs — all inference runs on TT hardware.

agents local-llm aider coding-assistant quietbox on-device
wormhole blackhole quietbox

dflash

affiliated ⑂ z-lab/dflash
by zoecarver · Python ·

DFlash: Block Diffusion for Flash Speculative Decoding on Tenstorrent hardware using tt-lang. Combines block diffusion with speculative decoding for faster inference.

speculative-decoding diffusion tt-lang inference

Engram

affiliated ⑂ deepseek-ai/Engram
by zoecarver · Python ·

A Tenstorrent port of the DeepSeek Engram model using tt-lang. Brings DeepSeek's memory-efficient architecture to TT hardware.

📦 Repo
deepseek engram tt-lang inference
blackhole

gsplat_tt

affiliated
by smartonTT · Python ·

Port of Gaussian Splatting (3D scene reconstruction from 2D images) to Tenstorrent hardware.

📦 Repo
gaussian-splatting computer-vision 3d-reconstruction blackhole
blackhole

Stable Diffusion XL on Tenstorrent

affiliated
by ·

On-device image generation with Stable Diffusion XL running entirely on Tenstorrent hardware. Full inference pipeline with no cloud dependency.

stable-diffusion sdxl image-generation diffusion on-device
wormhole blackhole

Video Generation on Tenstorrent

affiliated★ featured
by ·

Three lesson-projects covering on-device video synthesis: frame-by-frame diffusion with tt-local-generator, native AnimateDiff video animation, and video generation on QuietBox 2. All run entirely on TT hardware with no cloud dependency.

video-generation diffusion animatediff tt-local-generator quietbox on-device
wormhole blackhole quietbox

dev.to/mando222 — Tenstorrent & AI Blog

affiliated
by ·

Eric Zietlow's blog covering Tenstorrent hardware, Metalium programming, and AI topics, sharing practical experience with Blackhole and the broader TT ecosystem from our developer relations team.

blog metalium blackhole ai dev-to planet
blackhole

tt-forge-compiletron

affiliated
by tsingletaryTT · Python ·
tt-forge-compiletron preview

Compile more than 100 models on tt-forge in a display format suitable for demos. Comprehensive showcase of tt-forge model compatibility.

📦 Repo
# Changelog

All notable changes to tt-forge-compiletron are documented here.

## [Unreleased]

### Added
- `docs/kv-cache-bench.md` — teaching companion for the StaticCache KV cache
  benchmark, explaining the two-graph pattern and why static shapes matter

---

## [1.6.0] — 2026-06-30

### Added
- **StaticCache KV cache decode benchmarking** — `bench_decode.py` now compiles
  a second forge graph for the decode step using `transformers.StaticCache`.
  The StaticCache is embedded in `KVDecodeWrapper` as a submodule so forge
  traces K/V tensors as model state and emits `FillCache`/`UpdateCache` ops.
  Falls back to full-recompute for models that don't support `cache_position`.
- `_try_kv_decode()` function — detects model dtype to avoid bfloat16/float32
  mismatches, resolves tokenizer from loader or AutoTokenizer, pre-fills cache
  on CPU before forge compilation.
- Bestiary `decode_note` field now records the method used per model
  ("StaticCache KV cache" vs "no KV cache — full recompute per step").

### Changed
- Decode results updated for all 5 stages — GPT-2 2.30→5.52 tok/s, OPT
  3.98→5.05 tok/s, Phi-2 1.48 tok/s (new), Falcon 3.30 tok/s (new),
  LLaMA-LoRA 2.86 tok/s (new), Gemma-LoRA 2.40 tok/s (new), and more.

---

## [1.5.0] — 2026-06-30

### Added
- **`scripts/bench_decode.py`** — dedicated LLM decode benchmark measuring
  TTFT, prefill tok/s, and decode tok/s for all compiled causal LMs.
  Subprocess isolation + tt-smi health check prevent hardware lockups.
- **Leaderboard columns** — TTFT, Prefill tok/s, Decode tok/s, Params (M)
  replace the old Infer p50 / Throughput columns in `docs/leaderboard.html`.
- 5 benchmark stages: Stage 1 (GPT-2, OPT), Stage 2 (Phi-2, BLOOM, CodeGen),
  Stage 3 (Falcon, Allam, LLaMA-LoRA, Gemma-LoRA), Stage 4 (Qwen 2.5,
  Phi-1 LoRA), Stage 5 (DeepCogito, DeepSeek Coder, frontier models).
- `params_m` field added to all benchmarked bestiary entries.
- `hf:` loader prefix for frontier HuggingFace models loaded without a
  tt-forge-models seed loader.

### Changed
- Bestiary `throughput_unit` relabeled from generic `tok/s` → `prefill_tok/s`
  for all 54 causal LM entries to prevent confusion with decode throughput.

---

## [1.4.0] — 2026-06-29

### Added
- **`scripts/install.sh`** — turn-key smart installer: hardware pre-check,
  hugepages, disk space, forge venv, XLA venv, mesh descriptor probe,
  tt-forge-models clone, stale-shm cleanup. Outputs color-coded summary table.
- **RAM/DRAM budget calculator** — skips models whose weights exceed available
  system RAM + per-chip DRAM; prevents OOM crashes at load time.
- **`scripts/setup-venvs.sh`** — minimal venv setup script for clean Ubuntu
  24.04 installs on Tenstorrent Blackhole hardware.
- Self-contained patches directory — tt-forge-models fixes applied at
  expedition startup without modifying upstream.
- `--ephemeral` / `--evict-failures` flags — evict HF weight cache after
  each model to reclaim disk space on small-storage machines.

### Changed
tt-forge models demo compilation

Image Classification with TT-Forge

affiliated
by ·

End-to-end image classification project using TT-Forge — compile and run a PyTorch classification model on Tenstorrent hardware with no kernel authoring required.

forge image-classification pytorch compiler inference
wormhole blackhole

tensix-viz

affiliated★ featured
by tsingletaryTT · JavaScript ·
tensix-viz preview

Hardware topology visualizer for Tenstorrent chips — from individual chip to full cluster. Interactive JavaScript visualization of Tensix core layout and NoC connections.

LATEST v1.2.0 2026-08-05T16:35:19Z Release notes ↗
3 previous releases
v1.1.2 2026-06-29T17:42:46Z
v1.1.1 2026-06-26T13:32:52Z
v1.1.0 2026-06-09T22:19:42Z
See all releases on GitHub ↗
# Changelog

All notable changes to tensix-viz are documented here.

## [1.1.2] - 2026-06-29

### Fixed

- **`TensixViz.autoInit()` is idempotent for `.tensix-viz-container` elements** (`src/chip.js`)
  1.1.1 made the `[data-viz]` path idempotent but left the legacy single-chip path unguarded. When
  `autoInit()` ran twice (the bundle's self-init plus an explicit host-page call), each
  `.tensix-viz-container` canvas received a second `TensixViz` instance — two animation loops drawing
  on one canvas, which renders as a doubled/overlapping grid. `TensixViz.autoInit()` now skips any
  container already initialized (`container._tensixViz`) and records the instance on it.

### Added

- **Responsive multi-chip canvas** (`tensix-viz.css`)
  `.tv-chip-wrapper canvas { max-width: 100%; height: auto; }` — card/system canvases (created
  without the `.tensix-viz-canvas` class) now scale to fit a narrow column instead of being clipped
  by `.tv-card`'s overflow. Previously this rule had to be patched in by downstream consumers.

## [1.1.1] - 2026-06-25

### Fixed

- **Animation player accepts both script schemas** (`src/chip.js` `_execStep`)
  The player dispatched on `step.step` and read `step.cores` only, so scripts authored with the
  alternate `{ action, coords }` schema ran zero steps — the Play button (and auto-play) appeared
  dead. `_execStep` now dispatches on `step.step || step.action` and falls back `coords → cores`,
  so blocks written in either schema animate.

- **`autoInit()` is idempotent for `[data-viz]` elements** (`src/index.js`)
  `autoInit()` can run more than once (the bundle's self-init plus an explicit call). For `card`
  and `system` vizzes — which append their render into the host element — the second run appended
  a duplicate set of chips. `autoInit()` now skips any element already initialized (`el._tensixViz`).

## [1.1.0] - 2026-06-09

### Fixed

- **Heatmap: non-tensix cells no longer painted by heat overlay** (`src/chip.js` `_drawHeatmap`)
  Commit 76dca80 added `coreType !== 'tensix'` guards to the pre-built artifacts but never to
  the source. The guards are now in `src/chip.js` so the next build preserves them. Without this
  fix, DRAM (col 5 on Wormhole), ETH (row 6 on Wormhole), and PCIe (col 8 on Blackhole) cells
  were colored by the heatmap overlay and could inflate `maxVal`, compressing the visible range
  for all tensix cells.

- **Memory overlay: stale phase not rendered after `reset()` on `showMemory: true` instances**
  (`src/chip.js` `reset()` and constructor)
  After calling `viz.activate(mode)` followed by `viz.reset()` on a canvas created with
  `showMemory: true`, `_memPhase` retained the frozen `_mem` object from the animation closure.
  `reset()` calls `render()` at the end, which caused `_drawMemoryLayer()` to run with stale data,
  producing a faint DRAM glow and L1 fill bars on an otherwise blank chip. `reset()` now sets
  `this._memPhase = null`; the field is also explicitly initialized to `null` in t
visualization topology noc hardware
wormhole blackhole
Blackhole · P100 / P150 / P300c · 140 Tensix cores
Wormhole · N150 / N300 · 64 Tensix cores
mode

tt-warp

affiliated
by tsingletaryTT · Python ·

Warp terminal plugin for Tenstorrent — integrates hardware status, model management, and developer workflows directly into the Warp terminal.

📦 Repo
warp terminal plugin developer-experience

Tensix Grid Playground

affiliated
by ·

Interactive browser-based visualizer of the Tenstorrent Tensix grid architecture. Explore the NoC, core layout, and dataflow patterns without hardware — a great companion for learning kernel programming.

visualization interactive noc tensix browser architecture

Tenstorrent Cookbook: Conway's Game of Life

affiliated
by ·

TT-Metalium implementation of Conway's Game of Life as a cookbook recipe. Each generation is a full parallel kernel dispatch over the grid — a clean introduction to stateful compute on Tensix cores.

game-of-life demo cookbook parallel metalium
wormhole blackhole

Tenstorrent Cookbook: Particle Life Simulator

affiliated★ featured
by ·

Particle Life simulation on Tenstorrent hardware — an emergent-behavior N-body system where simple attraction/repulsion rules between species produce complex lifelike patterns. Cookbook recipe demonstrating parallel N-body compute on Tensix.

particle-life n-body simulation emergent cookbook demo
wormhole blackhole

tt-sim Lab

affiliated
by Pragnajit Datta Roy · Shell ·

A university teaching lab for TT-Metalium kernel programming on a virtual Tenstorrent chip — one-click GitHub Codespace, no silicon and nothing installed locally. The primary track (labs 00-06) points tt-metal straight at libttsim via TT_METAL_SIMULATOR and walks from elementwise add through NoC multicast to multi-core and multicast matmul, backed by a source-level matmul guide. An optional advanced track (labs 10-16) boots an Ubuntu guest under ttsim-qemu, loads tt-kmd, surfaces /dev/tenstorrent/0, and runs tt-metal through the full PCIe path.

ttsim simulator qemu codespaces metalium matmul labs curriculum academia education
ttsim wormhole blackhole

CS Fundamentals on Tenstorrent Hardware

affiliated★ featured
by ·

Seven-module computer science curriculum taught on real Tenstorrent hardware. Covers RISC-V architecture, memory hierarchy, parallel computing, networks and NoC, synchronization, abstraction layers, and computational complexity — all grounded in what is physically happening on the chip.

computer-science curriculum risc-v parallelism memory noc education
wormhole blackhole

Custom Model Training on Tenstorrent

affiliated
by ·

Eight-lesson series covering the full custom training workflow on TT hardware: dataset fundamentals, configuration patterns, fine-tuning, multi-device distributed training, experiment tracking, model architecture basics, and training from scratch.

training fine-tuning multi-device distributed experiment-tracking curriculum
wormhole blackhole

Tenstorrent Cookbook: Core Recipes

affiliated
by ·

Three hands-on TT-Metalium kernel recipes: a Mandelbrot fractal explorer, real-time audio signal processing pipeline, and custom image filter stack. Each recipe is a complete kernel project with full source in the lesson.

cookbook mandelbrot audio image-processing metalium demo
wormhole blackhole

nvtop

community★ featured
by Syllo · C · GPL-3.0 · 10914⭐ ·
nvtop preview

htop-style process monitor for GPUs and AI accelerators. Supports AMD, Apple, Huawei, Intel, NVIDIA, Qualcomm — and Tenstorrent. Real-time utilization, memory, and process info in a terminal UI.

📦 Repo
LATEST 3.3.2 2026-02-08T17:57:16Z Release notes ↗
4 previous releases
3.3.1 2026-01-18T13:12:34Z
3.3.0 2026-01-16T13:28:09Z
3.2.0 2025-03-29T11:26:44Z
3.1.0 2024-02-23T15:04:44Z
See all releases on GitHub ↗
monitoring tui htop process-monitor terminal
wormhole blackhole

dstack

community★ featured
by dstackai · Python · MPL-2.0 · 2213⭐ ·
dstack preview

Vendor-agnostic orchestration for training, inference, and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

LATEST 0.21.1 2026-08-14T16:57:48Z Release notes ↗
4 previous releases
0.21.0 2026-08-06T13:31:14Z
0.20.29 2026-07-24T15:55:07Z
0.20.28 2026-07-16T14:58:47Z
0.20.27 2026-07-09T13:10:13Z
See all releases on GitHub ↗
orchestration kubernetes cloud multi-vendor

Booth

community★ featured
by Zaneham · C · Apache-2.0 · 1730⭐ ·

An open-source CUDA, HIP, and Triton compiler with no LLVM anywhere in the path. Takes the same sources you would hand to nvcc, ROCm, or Triton's JIT and emits AMD RDNA 2/3/4 binaries, NVIDIA PTX, Tenstorrent Metalium C++, native RV32IM, or plain x86-64 — so a Triton matmul can run on a laptop that has never seen a GPU. Also reads a deliberate subset of MLIR (`func.func` plus the `arith` dialect) and Fortran `do concurrent` kernels via LFortran. Formerly BarraCUDA; renamed to honour Kathleen Booth.

📦 Repo
LATEST v0.5.2 2026-08-08T03:05:39Z Release notes ↗
2 previous releases
v5.01 2026-07-14T08:02:45Z
v0.5.0 2026-05-29T04:30:28Z
See all releases on GitHub ↗
Booth — Changelog
=================

## Unreleased

### Frontend

- `kath --mlir` reads MLIR text, no LLVM in the path. Čertík's pure-C
  reader vendored under `src/mlir/vendor` (mlir 826b69c9, corec a160199d),
  reached only through `src/mlir/mlir_fe.c` (Zane Hambly, 2026-08-11)

- `src/mlir/lower.c` walks the parsed module into BIR: `func.func`, `return`,
  `arith.constant` and every arith binop, compare and conversion the reader
  classifies. From there it is the pipeline CUDA and Triton already use, and
  MLIR reaches all four backends. `--mlir --pp` reprints instead
  (Zane Hambly, 2026-08-11)

- an op outside the subset stops the lowering and names itself. Skipping it
  would leave a function that compiles and computes something else
  (Zane Hambly, 2026-08-11)

- five fixes to the vendored reader, all worth upstreaming, and four of them
  are `func.func` being unfinished where `tt.func` is not: `parser_init`
  renamed off Booth's own, `parser_error`'s `exit(1)` replaced by a
  `mlir_parse_fail()` the linker supplies, `func.func` binding its arguments
  before parsing the body rather than after, `func.func` accepting the
  `attributes` clause where MLIR actually writes it, and `arith.xori`,
  `shli` and `shrsi` added to `op_string_to_type`, which the printer could
  already write but the parser could not read back
  (Zane Hambly, 2026-08-11)

- `ml_parse` resets the reader's process-wide type interning, which upstream
  assumes one context per process. Without it a closed context left the next
  parse in freed memory (Zane Hambly, 2026-08-11)

- the Triton lowering records pool overflow through `bir_pfull`, which the C99
  one already did and it never has. It answered a full block pool with index 0,
  a live block, so `bir_pchk` could not see a Triton arena exhaustion at all
  (Zane Hambly, 2026-08-11)

- Triton blocks are named. String offset 0 is a live string, so a nameless
  block printed as whatever went into the table first, and all four blocks of
  a loop kernel were labelled with the kernel's own name
  (Zane Hambly, 2026-08-11)

### Architecture

- BIR arena writers record a `pool_full` bit rather than returning index 0,
  which is a live entry and not a sentinel. A full pool emitted wrong
  immediates under exit 0; `bir_pchk` now refuses (Zane Hambly, 2026-08-11)

- #160: DCE and mem2reg move instructions without moving `inst_lines[]`
  with them, so every line number past the first deleted instruction
  pointed at the wrong source. Four sites fixed
  (Zane Hambly, 2026-08-09)

### Build

- #160: vendor Kauri (MIT) as `src/kauri.h`, included from `barracuda.h`,
  so `KA_GUARD`, `KA_CHK` and `KA_PNEW` are available tree-wide
  (Zane Hambly, 2026-08-09)

### CI and tests

- `make mutate` bends one line of Booth at a time in a scratch copy and checks
  the suite notices, from a table in `tests/mutants.tbl`. Ported from Kahu's
  (Zane Hambly, 2026-08-12)

- six tests that were not testing what they looked like they were. The `cfd`
  f
cuda hip triton fortran mlir compiler cross-platform metalium rv32im no-llvm
blackhole

tt-tiny

community★ featured
by geohot · Python · 69⭐ ·

Minimal Python code to access and program the Tenstorrent Blackhole chip directly — George Hotz's exploration of TT hardware programmability with pointed commentary on the architecture.

📦 Repo
blackhole low-level exploration
blackhole

zyx

community
by zk4x · Rust · LGPL-3.0 · 63⭐ · Sep 25, 2022
zyx preview

A complete ML library and compiler in Rust — "from assembly to neural networks" — with a native Tenstorrent backend (src/backend/tenstorrent), autograd, custom kernels, multi-backend support, and Python bindings.

LATEST v0.14.0 2024-09-22T13:54:32Z Release notes ↗
4 previous releases
v0.12.0 2024-03-10T09:41:35Z
v0.11.3 2023-10-08T12:12:13Z
v0.11.2 2023-10-08T11:55:14Z
v0.11.1 2023-09-28T14:06:33Z
See all releases on GitHub ↗
rust ml-compiler tensor-library autograd backend

tt-twitch

community
by geohot · C++ · 29⭐ ·

A Tenstorrent Grayskull kernel written live on Twitch by George Hotz. 120-core grid demonstration of live kernel programming.

📦 Repo
grayskull kernel live-coding demo
grayskull

koyeb/tenstorrent-examples

community
by koyeb · Dockerfile · 19⭐ ·

Example applications and deployment configurations for running AI workloads on Tenstorrent hardware via Koyeb's cloud platform.

cloud koyeb deployment examples

blackhole-py

community
by boopdotpng · Python · MIT · 18⭐ ·

Pure Python driver for Tenstorrent Blackhole cards providing direct low-level hardware access without going through the full TT-Metal stack.

📦 Repo
driver python blackhole low-level hardware-access
blackhole

Tenstorrent Console Skill

community
by aldegad · Shell · MIT · 15⭐ ·

An agent skill (SKILL.md) that teaches Claude Code, Codex, and the Agent SDK how to drive the console.tenstorrent.com inference API: OpenAI-compatible chat with DeepSeek-R1 and Qwen3, async image jobs, and Wan 2.2 text-to-video. Ships runnable curl examples and a mock-curl test harness; documentation is bilingual Korean/English.

📦 Repo
agent-skill claude-code codex agent-sdk tenstorrent-console inference-api text-to-video wan2-2 korean

tenstorrent-tiny-examples

community
by jaebaek · C++ · 14⭐ ·

Simple C++ kernel experiments on a GraySkull e75 chip. Hands-on examples for learning the TT-Metal programming model at the metal level.

📦 Repo
examples grayskull cpp learning
grayskull

ttnn-helloworld-cpp

community
by marty1885 · C++ · 14⭐ ·

Minimal working example of using Tenstorrent TTNN in C++. The simplest possible starting point for C++ developers targeting TT hardware with TTNN.

📦 Repo
c++ ttnn hello-world template
wormhole blackhole

tt-sim

community★ featured
by mesham · Python · 14⭐ ·

Community-built Tenstorrent architecture simulator written in Python. Runs without hardware — useful for researchers and developers exploring the Tensix architecture offline.

📦 Repo
LATEST v1.0 2026-05-11T13:07:42Z Release notes ↗
See all releases on GitHub ↗
simulator architecture no-hardware research

triton-tenstorrent

community★ featured
by kernelize-ai · C++ · 12⭐ ·

OpenAI Triton compiler plugin for Tenstorrent hardware. Write Triton kernels and target Tensix cores — brings the Triton ML kernel ecosystem to TT devices.

📦 Repo
triton openai-triton compiler kernels
wormhole blackhole

tt-iree

community★ featured
by swote-git · C++ · Apache-2.0 · 12⭐ ·

IREE (Intermediate Representation Execution Environment) ML compiler ported to Tenstorrent AI accelerators. Brings the IREE compiler ecosystem to TT hardware.

📦 Repo
iree compiler mlir inference
wormhole blackhole

TT-GoL

community
by JushBJJ · C++ · 12⭐ ·
TT-GoL preview

Conway's Game of Life implemented on Tenstorrent hardware using TT-Metal kernels.

📦 Repo
game-of-life demo kernels

Tenstorrent Simulator Playground

community
by thatdspguy · TypeScript · MIT · 8⭐ ·
Tenstorrent Simulator Playground preview

A web playground that runs real TTNN operations on the ttsim hardware simulator — no card required. Switch between Wormhole and Blackhole, run elementwise/activation/matmul ops or a small MLP, draw a digit and classify it with a trained MNIST net, then sweep parameters in 1D or 2D and read latency, throughput, and memory back as line charts, 3D surfaces, and heatmaps.

📦 Repo
# Changelog

All notable changes to the Tenstorrent Simulator Playground will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [1.4.0] - 2026-01-27

### Added

- **Updated Demo Assets**
  - New digit recognition demo GIF with improved visualization

### Fixed

- **Docker Desktop Compatibility**
  - Fixed 500 Internal Server Error when running digit recognition in Docker containers
  - Implemented native Docker inference that runs Python directly instead of attempting WSL calls
  - Proper environment variable handling for containerized ttsim execution
  - Automatic SOC descriptor selection for Wormhole/Blackhole chips in Docker

## [1.3.0] - 2026-01-27

### Added

- **Enhanced Network Visualization**
  - 28×28 pixel grid visualization for input layer displaying the actual drawn digit
  - Vertical output column showing all 10 digit classes (0-9) with labels
  - Green ring highlight indicator for the predicted digit
  - Pixel data pass-through from drawing canvas to network visualization

### Changed

- **Performance Statistics Display**
  - Average throughput now displays with 2 decimal places for precision

## [1.2.0] - 2026-01-27

### Added

- **Digit Recognition Page**
  - Interactive drawing canvas for handwritten digit input (0-9)
  - Real-time MNIST neural network inference on Tenstorrent simulator
  - Network architecture visualization showing layer activations and connections
  - Performance statistics tracking (average latency and throughput)
  - Confidence display with horizontal bar chart for all 10 digit classes
  - Trained 2-layer MLP model (784→128→10) achieving ~98% accuracy
  - New digit recognition API endpoints (`/api/digit/predict`, `/api/digit/model-info`)

- **Simple 2-Layer MLP Operation**
  - Configurable neural network architecture editor
  - Interactive network diagram showing input, hidden, and output layers
  - Configurable sizes: input (32/64), hidden (32/64), output (16/32/64), batch (16/32/64)
  - Parameter count display with real-time updates
  - Visual representation of layer connections and ReLU activation

- **Matrix Multiplication Operation**
  - Support for 32×32 and 64×64 matrix operations
  - Optimized for TTNN performance benchmarking
  - Added matrix multiply icon to operation selector

- **React Router Integration**
  - Separate pages for Digit Recognition (/) and Mathematical Operations (/math-operations)
  - Clean URL structure with browser navigation support
  - Navigation sidebar with page links

### Changed

- **UI Redesign**
  - Non-collapsible sidebar with fixed width (256px) for better text display
  - Simplified branding: "TTSim Playground" with chip icon
  - Mathematical Operations page header now matches Digit Recognition style
  - Operation selector reorganized into 4×3 grid by category
  - Updated icons for subtract (circle with line) and matrix
ttsim simulator playground ttnn visualization mnist parameter-sweep react
ttsim wormhole blackhole

ttMandelbrot

community
by marty1885 · C · 0BSD · 8⭐ ·

Mandelbrot Set fractal renderer running on Tenstorrent hardware. A classic demo showcasing parallel compute on Tensix cores.

📦 Repo
mandelbrot demo fractals parallel

tenstorrent.nix

community
by RossComputerGuy · Nix · LGPL-2.1 · 8⭐ ·

Nix flake packaging the Tenstorrent software stack for NixOS and Nix users. Reproducible, declarative installation of TT drivers and tools.

📦 Repo
nix nixos packaging flake reproducible

TT-Metal Mini Template

community
by JushBJJ · C++ · 7⭐ ·

Minimal working CMake project template for starting a new TT-Metal project from scratch. Good starting point for community kernel development.

📦 Repo
template cmake starter boilerplate

tt-tutorial (HPC)

community
by RISCVtestbed · C++ · BSD-3-Clause · 7⭐ ·

Tutorial on Tenstorrent hardware for HPC researchers from the RISC-V Testbed project at Edinburgh/EPCC. Covers Wormhole from an HPC parallel-computing perspective.

📦 Repo
tutorial hpc epcc edinburgh wormhole
wormhole

tenstorrent-cli

community
by aldegad · TypeScript · MIT · 6⭐ ·

A TypeScript/Bun terminal client for console.tenstorrent.com. Opens a chat REPL across DeepSeek-R1, Qwen3-32B, Qwen3-VL, and Gemma, with slash commands that submit image, Wan 2.2 video, TTS, and STT jobs, poll them, and save the results under ./output. Reads its API key only from the TENSTORRENT_KEY environment variable — never from disk.

📦 Repo
cli repl typescript bun tenstorrent-console inference-api text-to-video

ttPEAK

community
by TT-Bounty-Hunters · C++ · ISC · 6⭐ ·

clpeak-style peak-performance benchmark for Tenstorrent devices using TT-Metalium. Measures theoretical peak throughput across operations — useful for hardware characterization.

📦 Repo
benchmark performance clpeak metalium
wormhole blackhole

Unofficial Blackhole Documentation

community
by boopdotpng · Python · MIT · 6⭐ ·

Unofficial documentation for the Blackhole P100A / P150, assembled from reverse-engineering, disassembly, and hands-on experiment. Walks from a self-contained intro through chip architecture (NoC, Tensix tiles, RISC-V cores, L1, memory map), the 3-kernel matmul model, SFPI kernel writing, circular-buffer dataflow, the JIT build and dispatch pipeline, firmware boot sequence, and multi-host scaling. The author notes most pages were drafted by coding agents, with a `human/` folder that is explicitly hand-written.

📦 Repo
documentation blackhole reverse-engineering noc sfpi circular-buffers dispatch firmware
blackhole

current

community
by seansiddens · C++ · 5⭐ ·

High-level parallel programming framework for Tenstorrent accelerators, abstracting TT-Metal into a research-oriented programming model for parallel computation.

📦 Repo
framework parallel abstraction research
wormhole blackhole

ttVecAdd

community
by marty1885 · C++ · ISC · 5⭐ ·

Minimal vector-addition example on Tenstorrent devices using TT-Metalium. A clean hello-world for the TT-Metal kernel programming model in C++.

📦 Repo
vector-add example metalium hello-world

bhx

community★ featured
by olofj · Rust · 5⭐ ·

Boot stock Linux cloud images on the SiFive X280 RISC-V cores inside Tenstorrent Blackhole AI accelerators. Per-card Rust daemon with virtio-mmio block/net/console and U-Boot/EFI support.

📦 Repo
# Changelog

Notable changes per release. Format loosely follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
this project does not yet promise SemVer compatibility on the RPC
wire format or library API surface (we're not 1.0).

## Unreleased

V2 virtio-dispatch redesign. The kick ring + completion ring + host-
side throttle that grew up around #184 are gone; in their place is a
per-(slot, queue) dirty bitmap in BRISC L1. The bitmap is level-
sensitive — guest QUEUE_NOTIFY storms coalesce into a single set
byte, so the dispatch path can't fall behind under any burst. Wire
incompatible with 0.9.0; `TENSIX_PROTOCOL_VERSION` bumped 4 → 5.

### Added

- **V2 dirty-bitmap dispatch** (`#187` / `#188` / `#189`). BRISC
  writes 1 to `CTRL_OFF_DIRTY[slot][queue]` on every guest
  QUEUE_NOTIFY; the daemon's `Dispatcher` clears the byte and
  dispatches each pass. Replaces V1's 2048-entry kick ring +
  daemon-side `consume_kick_ring_pass` consumer.
- **V2 processed-cursor table** at `CTRL_OFF_PROCESSED`. Daemon
  publishes `used.idx` after each successful dispatch so
  warm-resume reads cursors directly without re-probing guest
  DRAM.
- **`bhx_notify_events_total`, `bhx_dispatch_passes_total`,
  `bhx_dispatch_queues_drained`** Prometheus counters surface the
  new dispatch path. The burst regression test (`scripts/
  soak_virtio_burst.py`) asserts `dispatch_passes_total > 0` to
  confirm the workload reached the new path.
- **`scripts/soak_virtio_burst.py`** — multi-queue burst regression
  test. Sustains 16-job direct=1 fio randwrite + a tight
  `printf` loop to `/dev/console`, samples `/metrics` every 1 s,
  and verifies the daemon log contains zero
  `kick.*drop|rescue|throttle.*ENGAGE` matches.
- **`DaemonState.chip_reset_this_session`** flag — gates
  `maybe_opportunistic_reset_board` so 4-way parallel cold boots
  reset the chip exactly once, not once per L2CPU. Without this
  the second-and-later resets blip the chip while earlier-booted
  L2CPUs hold mmap pages, SIGBUSing their workers.
- **`Dispatcher` (was `KickPoller`)** with documented testability
  seam (`CtrlL1Access` trait); `drain_dirty_bitmap` is unit-tested
  against an in-memory L1 fake covering all five visit/clear
  semantics cases plus the address-formula pins.

### Changed

- **`KickPoller` → `Dispatcher`**, plus `kick_poller` → `dispatcher`
  field on `DaemonState`, `tensix-kick-poller` → `tensix-dispatcher`
  thread name, `[kick-poller]` → `[dispatcher]` log tag,
  `kicks_consumed` → `dispatches_total`,
  `last_kick_slot_queue` → `last_dispatch_slot_queue`. Pure
  rename; no behavior change. V1 vocabulary scrubbed throughout
  the codebase (firmware, daemon, scripts, docs).
- **`CTRL_SIZE` shrinks 36 KiB → 4 KiB**. V2 footprint is ~1.5 KiB;
  the rest is reserved for future fields.
- **Stats-page offsets repacked** — V1 `STATS_OFF_KICK_DROPS`,
  `STATS_OFF_COMPL_EVENTS`, `STATS_OFF_LAST_COMPL` retired with
  V1 (#190); deprecated PRECAP / BLINDCAP / POSTCAP slots dropp
blackhole risc-v linux boot virtio
blackhole

ttas

community
by Zaneham · C · Apache-2.0 · 4⭐ ·

ttas is a hacker-friendly assembler/disassembler for Tensix on Wormhole. It turns assembly into the exact 32-bit words the hardware runs, and turns binaries back into readable instructions using the same shared instruction table.

📦 Repo
LATEST v0.1.0 2026-05-28T07:08:35Z Release notes ↗
1 previous release
v0.0.1 2026-05-27T15:19:11Z
See all releases on GitHub ↗
assembler
wormhole

tt-tutorial (Korean)

community
by changh95 · Jupyter Notebook · 4⭐ ·

Comprehensive tutorials for the Tenstorrent software stack in Korean. Jupyter notebooks covering the full developer path from hardware setup to model inference.

📦 Repo
tutorial korean jupyter getting-started
wormhole

ttRoPE

community
by Martin Chang · C++ · 0BSD · 4⭐ ·

A GGML-formatted rotary positional embedding (RoPE) implementation for Tenstorrent hardware — one of the operator building blocks behind the community effort to give llama.cpp a Metalium backend.

📦 Repo
rope operator ggml llama-cpp tt-metal cpp
wormhole

Collective Operations on Wormhole n150 (Sapienza University of Rome)

community

Master's thesis implementing and benchmarking five allreduce algorithms (Swing, Recursive Doubling, Bandwidth Optimal, Latency Optimal, Shared Memory) on the Wormhole n150. Bandwidth Optimal achieved best performance, approaching within 2× of theoretical optimal.

📦 Repo
allreduce collective-ops wormhole mpi bandwidth
wormhole

tt-model-bringup

community
by aweditya · 3⭐ ·

Direct TT-Metal bringup of modern open-weight LLMs on Blackhole P150 — hand-written compute graphs with no PJRT and no JAX. Covers Qwen3.6-27B, Qwen3.6-35B-A3B MoE, Gemma 4 12B, and Nemotron-3 Nano 30B-A3B, plus a zoo of single-chip Llama / Qwen2.5 / SmolLM ports, backed by custom fused `owned_*` kernels, a continuous-batching engine, an OpenAI-compatible HTTP server, and a wiki documenting each design decision.

llm model-bringup qwen gemma nemotron moe continuous-batching openai-compatible custom-kernels
blackhole quietbox

libtt-metal-cxx

community
by Knight-Ops · Rust · 3⭐ ·

Rust crate that exposes the TT-Metal host API through a C++ bridge via cxx.rs — covering device management, program/kernel creation (from source file or inline string), circular buffers, semaphores, runtime arguments, sharded buffers, and MeshDevice workflows, with hardware-backed integration tests.

📦 Repo
rust bindings cxx tt-metal ffi host-api
wormhole blackhole

ttperf

community
by Aswincloud · Python · MIT · 3⭐ ·

A CLI wrapper that turns TT-Metal performance profiling into one command. Runs a pytest target under Tenstorrent's profiler, streams progress live, then parses the resulting CSV and reports total device kernel duration. Supports profiling by operation name (`ttperf add`) as well as by test path, and installs from PyPI.

# Changelog

All notable changes to this project will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [0.1.6] - 2025-01-14

### Added
- **Memory Configuration Support**: New command-line options for tensor memory configuration
- `--memory-config CONFIG`: General option with choices `[dram, l1]`
- `--dram`: Shortcut flag for DRAM memory (default)
- `--l1`: Shortcut flag for L1 memory
- Memory configuration extraction from CSV profiler output
- Memory config display in test result summaries

### Changed
- **Default tensor shape reduced from `[1, 1, 1024, 1024]` to `[1, 1, 32, 32]` for better performance**
- Enhanced `create_test_tensor()` function to accept memory_config parameter
- Updated all `ttnn.from_torch()` calls to use memory_config parameter
- Improved CSV extraction to read memory configuration from profiler output
- Enhanced debug output to show memory configuration

### Technical
- Added `validate_memory_config()` function with alias support
- Extended environment variable system with `TTPERF_CUSTOM_MEMORY_CONFIG`
- Updated operation_configs.json to include memory_config field
- Enhanced test file configuration parsing to handle memory settings
- Improved result reporting to include memory configuration details

## [0.1.4] - 2025-01-14

### Changed
- **Major Improvement**: Configuration extraction now reads from CSV profiler output instead of parsing text with regex
- Replaced 50+ complex regex patterns with structured CSV data parsing
- Enhanced `extract_test_config_and_status()` function to prioritize CSV data over text parsing
- Added new `extract_config_from_csv()` function for reliable configuration extraction

### Fixed
- More accurate shape, dtype, and layout detection from profiler results
- Improved reliability of configuration reporting in test summaries
- Better handling of tensor dimension parsing (e.g., "32[32]" format)

### Technical
- CSV-based extraction provides structured, consistent data vs. unreliable text parsing
- Maintains backward compatibility with text parsing as fallback
- Cleaner, more maintainable codebase with reduced complexity

## [0.1.0] - 2025-07-14

### Added
- Initial release of ttperf CLI tool
- Support for profiling TT-Metal tests with pytest
- Automatic CSV path extraction from profiler output
- Device kernel duration calculation
- Real-time output streaming
- Flexible command-line argument parsing
- Support for named profiles
- Comprehensive error handling

### Features
- Simple CLI interface: `ttperf [name] [pytest] <test_path>`
- Automatic detection of test files and paths
- Integration with TT-Metal profiler tools
- CSV parsing for performance metrics
- Real-time progress monitoring

### Dependencies
- pandas for CSV processing
- Python 3.7+ support
- TT-Metal development environment

## [Unreleased]

### Planned
- Enhanced error messages

##
profiling performance cli tt-metal pytest python

tt-monitor

community
by antonibertel · Rust · MIT · 2⭐ ·
tt-monitor preview

A translucent, undecorated desktop widget showing live per-chip telemetry for Tenstorrent accelerators: temperature and power sparklines against the card's thermal and TDP limits, AI clock, voltage, current, DRAM channel training and ECC error counts, PCIe link generation/width, and board identity. Reads hardware directly through luwen over /dev/tenstorrent — no Python, no tt-smi subprocess, and no root.

📦 Repo
telemetry monitoring widget luwen gtk4 rust ecc desktop
wormhole blackhole

tetsuh/tt-metal-community-distro-matrix

community
by tetsuh · Python · Apache-2.0 · 2⭐ ·
tetsuh/tt-metal-community-distro-matrix preview

A compatibility guardrail that continuously monitors whether [tt-metal](https://github.com/tenstorrent/tt-metal) and the official [tt-installer](https://github.com/tenstorrent/tt-installer) build successfully on community Linux distributions that are not part of Tenstorrent's official CI.

📦 Repo

tt-splat — matrix-native 3D Gaussian Splatting on Blackhole

community
by kinginu · Python · Apache-2.0 · 2⭐ ·
tt-splat — matrix-native 3D Gaussian Splatting on Blackhole preview

3D Gaussian Splatting rewritten to run on the matrix engine: a polynomial splat and order-independent weighted-sum blending replace exp and depth-sorted alpha, so the pipeline becomes GEMM → activation → GEMM. Renderer + trainer, trained device-resident on a Blackhole p150a.

📦 Repo
3d-gaussian-splatting 3dgs rendering matrix-engine weighted-sum-rendering
blackhole ttsim

ttWKV7

community
by Martin Chang · C++ · 2⭐ ·

A standalone tt-metal demo and test bench for the RWKV-7 (WKV7) state recurrence on Wormhole, built with a GGML backend in mind. Two compute kernels cover the domain: a chunked-parallel DPLR matmul path for any sequence length with on-chip chunk carry, and a sequential per-token decode path for L <= 32 that is faster for large-batch token generation. The host runner validates both against a CPU oracle by PCC/NMSE and benchmarks them over a sequence/batch grid.

📦 Repo
rwkv wkv7 operator tt-metal ggml recurrent cpp
wormhole

libtt

community
by Philipp Moritz · 1⭐ ·

A Bazel-built PJRT plugin (libtt.so) providing an XLA backend for Tenstorrent devices. Bundles the tt-xla PJRT implementation with tt-mlir and tt-metal into a single shared object so JAX code runs on Tenstorrent hardware, with patches so sglang-jax works out of the box.

📦 Repo
xla pjrt jax bazel sglang

ttPseudoRowMajor

community
by Martin Chang · C++ · Apache-2.0 · 1⭐ ·

A small TTNN-facing C++ library (ttprm) for running view-shaped tensor work without first materializing the view in DRAM. Targets Tenstorrent TILE tensors and uses cached device operations to gather/scatter through layout views.

📦 Repo
ttnn tensor tile layout cpp

tt-tinygrad

community
by Gogopex · Python ·

A Tenstorrent backend for tinygrad that targets TT-Lang rather than raw tt-metal: a Renderer classifies each UOp kernel graph as matmul, reduce, or elementwise and emits ttl.math.* Python source, and a Compiled device parses the rendered kernel's contract, materializes ttnn tensors from host buffers, and calls it in-process. Proof of concept — 110 pass / 13 xfail across 125 cases on a QuietBox, covering fused matmul, reductions, softmax, layernorm, and attention chains, on top of a three-line patch to upstream tinygrad.

📦 Repo
tinygrad tt-lang ttnn backend codegen renderer python
wormhole quietbox

Inside the Tenstorrent Chips: Grayskull, Wormhole, Blackhole, and Galaxy

community
by · Jun 12, 2026

A long third-party walkthrough of the Tenstorrent lineup — core architecture, per-product specs and pricing, and what the published benchmarks against NVIDIA and AMD actually support. Notable for its candour: it states plainly that independent third-party benchmarks remain sparse and flags firmware changes that reduced earlier performance claims. Part 4 of a six-part series on inference hardware.

architecture benchmarks grayskull wormhole blackhole galaxy blog comparison
grayskull wormhole blackhole galaxy

A Gentle Guide: Tenstorrent Card on Arch Linux with Metalium

community
by · Jul 7, 2024

Step-by-step guide to getting a Tenstorrent card running on Arch Linux with the full Metalium stack. Practical troubleshooting from someone who did it the hard way first.

arch-linux metalium installation blog getting-started
grayskull wormhole

Thoughts and Logs After Messing with Tenstorrent Grayskull

community
by · Jun 2, 2024

Honest field notes from getting a Grayskull card running and writing first Metalium kernels. Covers setup pitfalls, processor hangs, memory protection quirks, and what makes Metalium compelling despite early rough edges.

grayskull metalium getting-started blog honest-review
grayskull

Programming Tenstorrent Processors

community★ featured
by · Apr 21, 2025

Deep-dive into the Tenstorrent architecture and Metalium programming model — circular buffers, kernel synchronization, NoC routing, and where the footguns are. The honest guide to thinking in Tensix.

metalium programming-model tensix noc circular-buffers blog
wormhole blackhole

Tenstorrent Architecture — W&M CSCI654 Advanced Computer Architecture

community
by · Oct 9, 2024

Lecture 20 from William & Mary's graduate Computer Architecture course. Frames Tenstorrent in the landscape between GPUs and TPUs, draws comparisons to Cerebras and SambaNova, then dives deep into the Wormhole chip and Tensix core: the 5 RISC-V core design, SFPU, NoC, and dataflow execution model.

lecture architecture wormhole tensix risc-v sfpu noc academia
wormhole

Tensix Field Guide

community
by Tanay Anand · CC-BY-4.0 ·

A ten-chapter, plain-English tour of Tenstorrent's Tensix architecture written for someone who knows what a CPU and a GPU are and nothing else: the chip-level grid and NoC, the five RISC-V baby cores, the matrix engine and its LoFi/HiFi fidelity trade-off, the SFPU, L1 and circular buffers, and why everything is 32x32 tile-shaped. The goal is to get a newcomer to the point of reading tt-metal kernel code in one sitting. Self-described draft; every claim traces back to tt-metal tech reports, tt-llk docs, or METALIUM_GUIDE.

📦 Repo
tensix architecture sfpu noc circular-buffers tile-format risc-v education book
wormhole blackhole

Tenstorrent SFPU Kernel Series — Jason Davies

community★ featured
by jasondavies · Nov 12, 2025

Sponsored series of deep technical articles on implementing optimal SFPU kernels for the Tenstorrent Wormhole and Blackhole vector units. Covers where, typecasting, 16/32-bit integer multiplication, cube root, and accurate sin/cos/tan — with cycle counts, assembly walkthroughs, and Blackhole vs Wormhole comparisons throughout.

sfpu assembly vector-unit cycle-counting wormhole blackhole optimization sponsored
wormhole blackhole

tt-rqm-kernels

community
by RQM-Technologies-dev · Python · Apache-2.0 ·

Structured quaternion, rotor, and phase-aware tensor kernels on ordinary floating-point tensors, plus StructuredBench. Includes CPU/PyTorch references, simulator and emulator paths, and reproducible Wormhole/N300 evidence for quaternion multiply (`qmul`), fused SU(2) composition, and H2A Hamiltonian lowering.

structured-tensors quaternion rotor tt-metalium tt-lang simulator wormhole n300 benchmarks pytorch custom-kernels
wormhole ttsim

tt-wavelet

community
by ke1rro · C++ · MIT ·

One-level FP32 lifting wavelet transforms (LWT) on Wormhole and Blackhole, shipped as a TTNN-linked op library plus standalone `lwt`, `ilwt`, `lwt_2d`, and `ilwt_2d` binaries and a benchmark harness. Builds the whole local stack — TT-Metal, the TTNN Python bindings, and ttnn-wavelet — against the TT-Metal revision pinned in its submodule.

📦 Repo
wavelet lifting-transform signal-processing ttnn tt-metal cpp
wormhole blackhole

Attention in SRAM on Tenstorrent Grayskull

community
by · Jul 18, 2024

A fused kernel for the Grayskull architecture implementing Transformer self-attention entirely within SRAM. Combines matrix multiply, attention score scaling, and Softmax without DRAM accesses, achieving significant speedups over non-fused implementations.

attention transformer sram grayskull kernel risc-v
grayskull

Exploring Fast Fourier Transforms on the Tenstorrent Wormhole

community
by · Jun 18, 2025

Ports the Cooley-Tukey FFT algorithm to the Wormhole n300 RISC-V accelerator. The Wormhole draws 8× less power and consumes 2.8× less energy than a 24-core Xeon Platinum for a 2D FFT. ISC 2025.

fft wormhole hpc risc-v energy-efficiency epcc
wormhole

Assessing Tenstorrent Grayskull RISC-V MatMul Acceleration for LLMs

community
by · May 9, 2025

Evaluates the Tenstorrent Grayskull e75 RISC-V accelerator for matrix multiplication at reduced numerical precision (BFP8 and LoFi), a fundamental kernel in LLM inference computation.

matmul grayskull risc-v bfp8 lofi llm precision
grayskull

Porting Strategies for Gravitational N-Body Simulations on Tenstorrent Wormhole

community
by · May 4, 2026

Evaluates three strategies for scaling an N-body code across multiple Tenstorrent Wormhole accelerators. Builds on the established performance of single-card N-body work to explore parallelism via the on-chip NoC and multi-accelerator configurations.

n-body astrophysics hpc wormhole risc-v multi-accelerator simulation
wormhole

Accelerating Gravitational N-Body Simulations on Tenstorrent Wormhole

community
Nov 16, 2025

Accelerates an astrophysical N-body simulation on the Wormhole n300. Achieves 2× speedup and 2× energy savings over a highly optimized CPU implementation. SC '25 Workshop.

n-body astrophysics hpc wormhole risc-v simulation
wormhole

Numerical Kernels on a Spatial Accelerator: Tenstorrent Wormhole

community
Mar 24, 2026

Implements three numerical kernels and composes them into a conjugate gradient solver on Wormhole. Demonstrates AI accelerators merit consideration for HPC workloads traditionally dominated by CPUs and GPUs. 2026.

numerical-methods hpc conjugate-gradient wormhole sparse
wormhole

Accelerating Stencils on the Tenstorrent Grayskull RISC-V Accelerator

community
Sep 27, 2024

Explores stencil computation on the Grayskull PCIe RISC-V accelerator. Early academic work examining TT hardware for HPC stencil workloads. 2024.

stencil hpc grayskull risc-v
grayskull

Stencil Computations on Tenstorrent Wormhole

community
May 8, 2026

Maps 2D 5-point stencil computations onto the Tenstorrent Wormhole RISC-V AI dataflow accelerator via two implementations: element-wise decomposition (Axpy) and matrix-multiplication reformulation (MatMul). Profiling shows the isolated Wormhole kernel is competitive with CPU execution, with PCIe transfers and initialization driving end-to-end overhead; Axpy achieves lower energy than the CPU baseline at large scales. Identifies architectural and software directions for making AI accelerators viable for HPC stencil workloads. 2025.

stencil hpc wormhole risc-v energy-efficiency benchmarks dataflow
wormhole

SwiftNPU: Scalable Shape-Flexible Allocation for Inter-Core Connected NPUs

community
Apr 27, 2026

Makes multi-tenant NPU sharing practical for Blackhole-class hardware using polynomial-time allocation algorithms. Delivers up to 1.37× higher utilization and 1.14× faster workload completion. Up to 890,000× faster than NP-hard baselines.

multi-tenant allocation blackhole npu scheduling
blackhole

TileLoom: Automatic Dataflow Planning for Spatial Dataflow Accelerators

community
by · Dec 17, 2025

Compiler system that automatically generates efficient dataflow plans for tile-based languages on spatial accelerators including Tenstorrent Wormhole. Exploits on-chip network forwarding between processing elements to reduce DRAM pressure.

compiler dataflow spatial-accelerator tile-based on-chip-network wormhole
wormhole

Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent vs. NVIDIA L40S

community
by · Mar 24, 2026

Shows that Text-to-Speech inference on Tenstorrent Lightning V2 achieves 4× lower cost than NVIDIA L40S. Applies BlockFloat8 (BFP8) and low-fidelity (LoFi) precision strategies to TTS despite their greater numerical fragility compared to LLMs.

tts text-to-speech inference bfp8 lofi cost-efficiency precision
wormhole

Tenstorrent Blackhole Architecture Guide

community★ featured
by · Feb 28, 2026

A 6,500-word community deep dive into the Blackhole p100a architecture: the tile model (Tensix, DRAM, SiFive x280 L2CPU, Ethernet, PCIe, NoC arc), firmware startup sequence, MOP micro-op processor, replay buffer, FPU/SFPU sync, and the anatomy of a kernel. From the author of blackhole-py.

blackhole architecture tensix noc sifive-x280 firmware mop sfpu deep-dive blog
blackhole

All in RISC-V, RISC-V All in AI — FOSDEM 2026

community
by · Jan 31, 2026

Martin Chang and Danfeng Zhang on solving real AI compute problems with open hardware, spanning AI PC / edge devices and AI servers: pairing high-performance RISC-V CPUs with NPUs, and Tenstorrent's RISC-V cores and scalable mesh for AI workloads. The speaker's companion write-up covers the Tensix programming model in depth — the five RISC-V cores per tile, Dst register double-buffering, and the macro-recording hardware that lets control cores run ahead of the math engine.

fosdem talk risc-v tensix deepcomputing edge programming-model

WebNN and WebLLM on RISC-V — FOSDEM 2026

community
by · Jan 31, 2026

Yuning Liang and Petr Penzin on closing the AI acceleration gap in the browser on RISC-V: how WebNN and WebLLM can reach efficient on-device inference using the RVV 1.0 variable-length vector ISA and Tenstorrent hardware underneath.

fosdem talk risc-v rvv webnn webllm browser on-device