A curated directory of projects, tools, models, and research for Tenstorrent hardware — contributed by the community and our team. Browse by category or search across all entries.
133Projects
12Categories
Browse by category
🚀 Getting Started
The essential first steps — installer, core SDKs, and guided onboarding
tt-metal
official
TT-NN operator library and TT-Metalium low-level kernel programming model. The primary SDK for devel…
tt-forge
official
Tenstorrent's MLIR-based compiler frontend. Enables running AI workloads from PyTorch, ONNX, and oth…
tt-vscode-toolkit
official
48 interactive lessons covering the full Tenstorrent developer path — from hardware detection to cus…
🤖 AI & Models
Running, serving, and experimenting with AI models
TT Console
official
Browser-based cloud console for exploring AI on Tenstorrent hardware. Run LLM inference, image and v…
tt-bio
affiliated
Boltz-2 biomolecular model for drug discovery on Tenstorrent Blackhole. Supports single-card and mul…
koyeb/tenstorrent-examples
community
Example applications and deployment configurations for running AI workloads on Tenstorrent hardware …
🕵️ AI Agents
Agentic systems and AI assistants running on TT hardware
tt-example-apps
official
End-to-end AI applications running on Tenstorrent AI accelerators. Complete application examples fro…
Local AI Agents on Tenstorrent
affiliated
Three agentic projects running fully on-device: local AI agents on QuietBox 2, a coding assistant po…
dstack
community
Vendor-agnostic orchestration for training, inference, and agentic workloads across NVIDIA, AMD, TPU…
Tenstorrent MLIR compiler — the core compiler infrastructure shared by tt-forge and other frontends.…
tt-forge-compiletron
affiliated
Compile more than 100 models on tt-forge in a display format suitable for demos. Comprehensive showc…
BarraCUDA
community
Open-source CUDA compiler targeting multiple GPU architectures including Tenstorrent. Compiles .cu f…
🛠 Dev Tools & Debugging
Profiling, visualization, and debugging workloads
ttsim
official
Fast full-system simulator of Tenstorrent Wormhole and Blackhole hardware. Runs TT-Metalium workload…
tensix-viz
affiliated
Hardware topology visualizer for Tenstorrent chips — from individual chip to full cluster. Interacti…
nvtop
community
htop-style process monitor for GPUs and AI accelerators. Supports AMD, Apple, Huawei, Intel, NVIDIA,…
🖥 Hardware & System
Drivers, firmware, monitoring, and hardware management
tt-kmd
official
Tenstorrent kernel module driver. The Linux kernel module required to interface with Tenstorrent PCI…
tt-qb-lights
affiliated
Sync your Tenstorrent Quietbox's RGB lighting to accelerator utilization status. Visual feedback for…
blackhole-py
community
Pure Python driver for Tenstorrent Blackhole cards providing direct low-level hardware access withou…
☁️ Cloud & Orchestration
Kubernetes, cloud deployment, and multi-node infrastructure
tt-inference-server
official
Production-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. Supports co…
tt-topology
official
Configure Ethernet routing on multi-card Tenstorrent systems. Flash NB cards to use specific ETH rou…
vllm-tt-plugin
official
Tenstorrent backend for vLLM, built on vLLM's standard plugin mechanism — install it alongside vLLM …
🔩 RISC-V & Architecture
ISA, simulation, and running Linux on TT silicon
tt-bh-linux
official
Linux demo for the Tenstorrent Blackhole P100/P150 card RISC-V cores. Boot a real Linux kernel on th…
CS Fundamentals on Tenstorrent Hardware
affiliated
Seven-module computer science curriculum taught on real Tenstorrent hardware. Covers RISC-V architec…
tt-sim
community
Community-built Tenstorrent architecture simulator written in Python. Runs without hardware — useful…
🔬 Research & Papers
Academic papers, theses, and HPC experiments
tt-isa-documentation
official
Low-level ISA and microarchitecture documentation for Tenstorrent AI architectures (Grayskull, Wormh…
polaris
official
A high-level AI simulator from Tenstorrent for modeling and exploring AI accelerator and workload pe…
tt-tutorial (HPC)
community
Tutorial on Tenstorrent hardware for HPC researchers from the RISC-V Testbed project at Edinburgh/EP…
🎮 Games & Demos
Creative, playful, and proof-of-concept projects
tt-animatediff
official
Generates short, temporally coherent animated GIFs using the AnimateDiff model on Tenstorrent hardwa…
tt-zork-and-more
affiliated
A Tenstorrent fork of Infocom's Zork I (and more!), running a Z-machine interpreter at least four di…
TT-GoL
community
Conway's Game of Life implemented on Tenstorrent hardware using TT-Metal kernels.
📚 Guides, Tutorials & Education
Getting-started content, blog posts, lessons, courses
tt-installer
official
Install the complete Tenstorrent software stack with one command. Handles drivers, firmware, Python …
Custom Model Training on Tenstorrent
affiliated
Eight-lesson series covering the full custom training workflow on TT hardware: dataset fundamentals,…
Programming Tenstorrent Processors
community
Deep-dive into the Tenstorrent architecture and Metalium programming model — circular buffers, kerne…
Fresh from Planet Tenstorrent
Planet Tenstorrent is the ecosystem's live feed — new releases,
articles, papers, talks, and community posts from across the Tenstorrent world,
gathered in one place and updated daily. The latest five:
RFD3 multi-card design now scales from 0.68x to 3.48x a single card after fixing host thread pool oversubscription, and ESMC embeddings gained a 1.47x speedup through ttnn trace capture—both changes that matter if you're doing large batches or iterative design on Tenstorrent hardware. The release also unifies the design CLI around tt-bio design INPUT --model boltzgen|rfd3, mirroring the predict interface, with --devices N now properly honored on paths that were silently ignoring it; tt-bio gen still works as a hidden alias but will disappear in a future release. A slate of fixes restored bf16 defaults that were causing ESMFold2 NaN confidence, fixed trace forwarding for Protenix v2, and cleaned up worker lifecycle issues when dispatchers die, all validated against per-model accuracy floors and a parity gate covering 23 test legs with no observed drift.
The jump to Pydantic v2 unblocks Python 3.14 support and opens the door to newer ecosystem libraries, while bringing faster validation and CLI startup — but if you're using the Python API or have dstack plugins, you'll need to verify they're compatible with Pydantic v2 models before upgrading. Beyond the foundational work, this release brings several UX wins: dstack metrics now shows CPU, memory, and GPU utilization as sparklines over the last hour so you can spot trends at a glance; AWS gateways with ACM certificates can now run multiple replicas with HTTPS handled directly by dstack instead of requiring an external load balancer; and the preset system gained baseline trials, per-trial learnings, and a richer CLI view to make benchmark results more interpretable. You'll also see GPU driver detection in dstack fleet -v and new --full-offers and --unallocated flags for better capacity discovery on Kubernetes and Slurm. Note that servers and CLIs must be upgraded together — new CLIs won't work with 0.20.x servers.
The visualizer now renders cluster topology and placement directly from the cluster descriptor without requiring a baked SoC descriptor, making it easier to inspect multi-chip systems on the fly—and when you do have per-chip arch data, it's applied as optional enrichment rather than a hard blocker. Multi-host profiling is cleaner too: remote sync can now discover and download performance reports per rank, and reads default to rank 0 to eliminate the visual collision of overlapping data from every rank. A few quality-of-life fixes round out the release—version checks no longer flag patch updates, SSH config is read to auto-populate remote connections, and performance charts got a tidier layout with grouped headings.
tensix-viz v1.2.0 sharpens memory stats configuration for hardware-powered visuals, making it easier to instrument and monitor Tensix performance directly from visualization tools. This matters when you're profiling real workloads on silicon and need reliable memory telemetry baked into your dashboards without extra plumbing. Check the full changelog for implementation details.
This site is two old traditions sharing one orbit: an
awesome list and a planet. Open source is an
ecosystem the way space is — moonshot projects igniting into galaxies
of forks and stars, and planet sites keeping the whole universe in
view. (Around here the metaphor is load-bearing: we ship hardware
called Galaxy.) Both traditions are gifts from decades of that
culture, and both deserve some tribute.
🕶 The awesome list
Humans curating links for other humans is the oldest genre on the web —
Yahoo! began life in 1994 as "Jerry and David's Guide to the World
Wide Web", and the volunteer-run
DMOZ / Open Directory Project
kept hand-sorted order for two decades. In 2014, Sindre Sorhus distilled
that instinct into a GitHub-native microformat with
sindresorhus/awesome:
one README, a ruthless curation bar ("only awesome things"), and pull
requests as the editorial process. The
awesome manifesto
turned list-making into a commons — thousands of lists,
lists of lists,
and giants like
awesome-python,
awesome-selfhosted, and
awesome-go.
Synth heads are gloriously covered too:
awesome-musicdsp,
awesome-audio-dsp,
awesome-webaudio, and
awesome-supercollider.
Because the format is halfway to being a database, people have long
rendered lists into websites — tt-awesome just commits to the bit:
every entry is a JSON file, and the README, this site,
data.json, and the feeds are all built from the same source.
🪐 The planet
In the early 2000s, Jeff Waugh and Scott James Remnant wrote
Planet,
a little Python feed aggregator that river-merged a community's blogs
into one page — and free software communities never looked back.
Planet GNOME,
Planet KDE,
Planet Debian,
Planet Gentoo,
Planet Ubuntu,
Planet Fedora (the Red Hat family), and
Planet Mozilla
are all still ticking decades later, many having passed through Sam
Ruby's Planet Venus rewrite along the way. The lineage runs
deeper still: before planets there were blogrolls, and before blogrolls,
web rings —
WebRing was built in 1995 by a teenaged Sage Weil, who grew up to create
Ceph, which is about as open-source-full-circle as a story gets.
Planet Tenstorrent carries that
torch for this ecosystem.
TT-NN operator library and TT-Metalium low-level kernel programming model. The primary SDK for developing on Tenstorrent hardware — from high-level tensor ops to bare-metal RISC-V kernels.
Tenstorrent's MLIR-based compiler frontend. Enables running AI workloads from PyTorch, ONNX, and other frameworks on all Tenstorrent hardware configurations through an open-source, general, and performant compiler.
TT-BUDA: Tenstorrent's original Python compiler and runtime for AI workloads. Legacy stack — tt-forge is the recommended successor, but tt-buda has the largest model demo library.
Tenstorrent MLIR compiler — the core compiler infrastructure shared by tt-forge and other frontends. Handles graph optimization, lowering, and code generation for Tensix hardware.
Fast full-system simulator of Tenstorrent Wormhole and Blackhole hardware. Runs TT-Metalium workloads on any Linux/x86_64 system without physical silicon. Bit-exact results relative to hardware.
RISC-V architectural self-checking directed tests — randomly-generated register operands and data with low-level OS code for test scheduling and self-checking, runnable on a RISC-V design or an ISS such as Whisper or Spike. Generated by an internal Tenstorrent tool from the official RISC-V ISA spec.
Low-level ISA and microarchitecture documentation for Tenstorrent AI architectures (Grayskull, Wormhole, Blackhole) — the authoritative hardware reference beneath the tt-forge / tt-metal software stack.
RISC-V Directed Test Framework and Compliance Suite. Comprehensive test infrastructure for verifying RISC-V processor implementations against the specification.
Production-ready model serving for Tenstorrent hardware with OpenAI-compatible REST API. Supports continuous batching, multiple models, and all TT hardware configurations.
ONNX graph compiler for Tenstorrent hardware. Optimizes and transforms ONNX model graphs for efficient execution on Tensix accelerators. Used as a backend by tt-forge for ONNX model ingestion.
Repository of model demos using TT-Buda. The largest collection of pre-compiled model examples for Tenstorrent hardware — BERT, ResNet, YOLO, GPT-2, Whisper, and many more.
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## 3.0.26 - 29/07/25
- Added single tray galaxy reset option
- Bumped luwen from 0.7.5 -> 0.7.10
- Chip detect now doesn't wait for eth to train for the 6U galaxy's, allowing multi tray resets to happen independently
- Updated readme with the new reset option
## 3.0.25 - 29/07/25
- Added packaging
## 3.0.24 - 04/07/25
- Now users have 2 galay reset modes available
- glx_reset: resets the galaxy, informs users if there has been an eth failure
- glx_reset_auto: resets the galaxy upto 3 times if eth failures are detected
## 3.0.23 - 03/07/25
- Bumped luwen 0.7.3 -> 0.7.5 to fix cargo lock compatibilty issue
## 3.0.22 - 02/07/25
- Bumped tt-tools-common 1.4.16 -> 1.4.17
- Bumped luwen 0.7.2 -> 0.7.3
- Bumped smi 3.0.21 -> 3.0.22
## 3.0.21 - 26/06/25
- Added option to not re-init chips after reset
- Updated galaxy 6u reset option from --ubb_reset to -glx_reset
- Removed the a3 arc message before doing a 6u reset, meaning we can reset even when chips are not pcie accessible
- Added eth link check and return failure if any of the eth links have a LINK_INACTIVE_FAIL_DUMMY_PACKET failure
## 3.0.20 - 04/06/25
- Chore - bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0
## 3.0.19 - 30/04/25
- Fixed an issue preventing the telemetry thread from being dispatched when the user clicked tab 2
## 3.0.18 - 22/05/25
- Added BH and WH UBB board type support
- Removed the dependency on tt-tools-common for this info
## 3.0.17 - 13/05/25
- Added proper telemetry heartbeat checks for Grayskull
## 3.0.16 - 12/05/25
- Used new ResetTypes from tools-common to simplify reset code
- Added a heartbeat spinner to the telemetry pane. We expect this spinner to update about twice per second. If the spinner is not moving, this indicates new telemetry is not being fetched.
## 3.0.15 - 24/04/25
- Patch for the ubb_reset to just discover local only post reset. Looks like eth port status 2 has been re-used to mean connected and pyluwen waits for it to clear, leading to eth timeout.
## 3.0.14 - 21/04/25
- Added wh ubb reset via command line `tt-smi --ubb_reset`. Intention is that this command line option will be removed and integrated into `tt-smi -r` after we update board detection with the correct external naming.
- Removed some unused imports and code - no functional changes
## 3.0.13 - 21/03/25
- Removed get\_sw\_versions
## 3.0.12 - 21/03/25
- Chore - bumped luwen version to include eth fw version check fix
## 3.0.11 - 13/03/25
- Chore - bumped luwen version to include enable chips with external connections but no routing
## 3.0.10 - 10/03/25
- Chore - bumped luwen version to include protoc lib detection check
## 3.0.9 - 07/03/25
- Chore - bumped luwen v
monitoringtelemetrysmihardware-management
Works on
grayskullwormholeblackhole
tt-lang
official
Python · Apache-2.0 · 59⭐ ·
Python-based DSL that sits between TT-NN and TT-Metalium — expresses custom fused kernels with progressive disclosure, compiling directly to Tensix. Ships an integrated functional simulator (no hardware needed), line-by-line performance metrics, and AI-agent-friendly tooling. Two packages: tt-lang (compiler + hardware, requires ttnn) and tt-lang-sim (simulator only, works on Linux/macOS without Tenstorrent hardware).
# Changelog
All notable changes to TT-Lang will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## Version 1.1.1
### Compiler
- Fix for live-interval boundary computation (issue [#536](../../issues/536))
- Fix for all-zero results in FP32 reductions (issue # [#533](../../issues/533))
- Fix for inferred `pop` and `push` (issues [#536](../../issues/536), [#554](../../issues/554))
- Fix for write pointer tracking on pipe sender accross iterations (issue [#578](../../issues/578))
- Fix to report data type mismatch error
- Fix to report DFB over allocation error (issue [#511](../../issues/511))
- Support for pipenet predicates `is_src`, `is_dst` and `is_active` (issue [#541](../../issues/541))
- Support for `ttl.math.typecast`
### Simulator
- Support for inferred `pop`, `push` and `copy`'s transfer handle `wait`
- Support for pipenet predicates `is_src`, `is_dst` and `is_active`
- Support `all_gather`
- Support `bfloat8_b`
- Improved/actionable error messages
- Improved performance by simulating math in FP32
### Infrastructure
- TT-Lang installable with `pip install tt-lang` for full installation and `pip install tt-lang-sim` for simulator only
- [Matmul benchmarks](benchmarks/matmul/README.md)
## Version 1.0.0
### Compiler
- Support `+=` syntax in conjunction with dot product (`@`) lowered to packer L1 accumulation
- Support implicit temporary compute-kernel-local DFBs
- Support `ttl.Pipenet`
- Support implicit `ttl.Block.push` and `ttl.Block.pop`
- Support implicit `ttl.Transfer.wait`
- Support for `expm1`, `exp2`, `ceil`, `sign`, `gelu`, `silu`, `hardsigmoid`, `square`, `softsign`, `signbit`, `frac`, `trunc` in `ttl.math`
### Simulator
- Support for `ttl.GroupTransfer`
- SPMD and mesh device simulation support
- Support for `ttnn.all_reduce` CCLs
- Use tracing to report statistics with `tt-lang-sim-stats`
- Remote L1 reads/writes statistics
### Examples and documentation
- Matmul tutorial
## Version 0.1.8
### Compiler
- Support for dot product operator (`@`) with lowering to [`ckernel::matmul_block`](https://docs.tenstorrent.com/tt-metal/v0.55.0/tt-metalium/tt_metal/apis/kernel_apis/compute/matmul_block.html)
- Support for fusing matmul and certain elementwise operations
- Support lowering to `pack_tile_block`
- Support for `ttl.math.fill`, `ttl.math.reduce_sum`, `ttl.math.reduce_max`, and `ttl.math.transpose`
- Support for arbitrary sub-blocking including dot product K-dimension to allow maximizing L1 usage and reuse
- Support for `sin`, `cos`, `tan`, `asin`, `acos`, `atan` in `ttl.math`
- Support for L1 sharded tensors
- Support for tensors with BF8 data type
- SPMD support (`ttnn.open_mesh_device`)
### Simulator
- Track L1 space and number of DFBs usage and warn when exceeded
- Support for tensors with row-major layout
- Support for L1 sharded tensors
### Examples and documentat
Linux demo for the Tenstorrent Blackhole P100/P150 card RISC-V cores. Boot a real Linux kernel on the 16 high-performance RISC-V cores built into the Blackhole chip.
Comprehensive tool for visualizing and analyzing model execution on Tenstorrent hardware. Interactive graphs, memory plots, tensor details, buffer overviews, operation flow graphs, and multi-instance support.
Tenstorrent Low-Level Kernels: the C++ library that directly programs the RISC-V cores inside each Tensix compute engine. TRISC0 (unpack), TRISC1 (math/FPU/SFPU), and TRISC2 (pack) are all programmed through this layer — it is the interface between TT-Metal kernel code and bare silicon.
Lightweight BMC (Baseboard Management Controller) for STM32 and similar MCUs, with Web UI, Redfish API, and HTTPS support. Built on Zephyr RTOS. Used in Tenstorrent systems.
Web-based GUI for deploying and chatting with AI models on Tenstorrent hardware. Handles all technical setup automatically — deploy models, run inference, and explore capabilities through a simple browser interface.
# Changelog
## [0.9.5] - 2026-05-12
### Changed
Hardware hang detection for NOC and PCIe.
Tracy profiler integration with instrumentation across TLB, PCIe and sysmem paths.
DeviceProtocol ported to TTDevice, including DMA migration.
SocDescriptor split into static (SocArchDescriptor) and runtime parts.
LITERAL coordinate system in CoreCoord.
Multicast to all TENSIX cores.
SMN support.
SWEmuleChip software emulation chip and Quasar simulation support (incl. 4GB TLB).
Unified UmdException/UMD_ASSERT/UMD_THROW error handling across the codebase.
## [0.9.4] - 2026-03-18
### Changed
TopologyDiscoveryOptions refactoring.
TopologyDiscoveryOption to retrain ETH links on 6u.
TLBs for TTsim.
DRAM retrain support.
DeviceProtocol changes.
Simulator in TTDevice changes.
ETH heartbeat check.
## [0.9.3] - 2026-02-24
### Changed
Sigbus safe read write API.
Remove 4U related code.
Implement BH SPI as well, so full SPI support.
P150 expects harvested cores.
TT_VISIBLE_DEVICES uses logical IDs.
## [0.9.2] - 2026-02-09
### Changed
SPI interface for Wormhole.
PCI BDF based sorting and filtering.
Multicast PCI DMA.
Support Blackhole loudbox.
Many code fixes and test enhancements.
## [0.9.1] - 2026-01-23
### Changed
Started publishing to pypi.
## [0.9.0] - 2026-01-23
### Changed
Warm reset notification and callback implementation.
## [0.8.6] - 2026-01-20
### Changed
Make predicting ETH FW from CMFW optional in TopologyDiscovery.
## [0.8.4] - 2026-01-16
### Changed
Use older manylinux image
## [0.8.3] - 2026-01-15
### Changed
Reverted remote discovery issue
## [0.8.2] - 2026-01-15
### Changed
Support warm reset without secondary bus reset.
Expose subsystem vendor id.
## [0.8.1] - 2026-01-15
### Changed
Support dma functions on TTDevice layer
## [0.8.0] - 2026-01-14
### Changed
Many functional fixes and minor changes.
Final fixes needed for integration into tt-smi.
Also contains adjustments needed for integration into exalens.
## [0.7.0] - 2025-11-29
### Changed
Changed to a more generic arc_msg API.
## [0.6.0] - 2025-11-24
### Changed
Change the usage of TLBs such that KMD is in control of TLB allocation instead of UMD.
TLBs are now allocated using KMD's dedicated API.
## [0.5.3] - 2025-11-14
### Changed
Added generation of .deb and .rpm packages.
Added three separate packages (runtime, development and python).
## [0.5.1] - 2025-11-12
### Changed
Manylinux builds and Pypi test publishing.
Many smaller fixes and improvements.
## [0.4.0] - 2025-10-18
### Changed
Removed old type names.
## [0.3.0] - 2025-10-17
### Changed
Many smaller fixes and improvements.
TTsim support improvements.
JTAG support improvement.
Fixing CMake install path.
Further work on integrating new KMD TLBs.
## [0.2.0] - 2025-09-15
### Changed
A couple of smaller fixes and improvements, including L2CPU harvesting, fixes for new FW. Better TTSim support. Further JTAG support.
Introduced new soft reset API.
Introduced lite fabric initial version.
user-mode-driverumdhardware-interface
Works on
grayskullwormholeblackhole
tt-system-firmware
official
C · Apache-2.0 · 41⭐ ·
System firmware for Tenstorrent hardware. Low-level system initialization and control firmware that runs on-device.
TVM for Tenstorrent ASICs. Brings the Apache TVM compiler stack to Tenstorrent hardware, enabling model compilation from TensorFlow, PyTorch, ONNX, and more.
Optimized training recipes for a variety of ML models on Tenstorrent hardware, powered by the TT-Forge compiler stack. Reference implementations for fine-tuning and training from scratch.
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## 1.2.11 - 17/06/2025
### Updated
- Updated mesh coord generation to be connection type agnostic
- Added failure and exit if mesh type detected, but not enough connections
- Added warning in README about lack of supoort for BH and 6U boards
## 1.2.10 - 05/06/2025
### Updated
- Bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0
## 1.2.9 - 30/05/2025
### Updated
- Bug fix for https://github.com/tenstorrent/tt-topology/issues/39. Now the tool will use a DFS longest path to determine a linear layout if its not a fully connected graph.
- Updated initial device detection - now it needs full noc access for octopus and list options
## 1.2.8 - 08/05/2025
### Updated
- Fixed issue where tool would fail when PCI interfaces don't start from ID 0
- Now using actual PCI interface IDs from devices instead of assuming sequential numbering
## 1.2.7 - 07/05/2025
### Updated
- Use tools-common 1.4.15
- Use type checking in octopus reset
## 1.2.6 - 05/05/2025
### Updated
- Bug fix: added "ignore-eth" flag to first chip detect to avoid eth training loops forever and truly detect pcie only chips
- Chore: bumped luwen
## 1.2.5 - 15/04/2025
### Updated
- When flashing to isolated mode, we now flash the WH ethernet ports to a disabled state,
in order to prevent their use.
## 1.2.4 - 02/04/2025
### Updated
- You can now run `tt-topology -l isolated` to flash cards to the default (non-connected) state
- Users are now warned about missing or loose cables
## 1.2.3 - 21/03/2025
### Fixed
- Bumped luwen (0.6.2 -> 0.6.3) to include eth version check bug for TG setup
## 1.2.2 - 13/03/2025
### Fixed
- Bumped luwen version to make it more robust against eth fw updates
## 1.2.1 - 13/03/2025
### Fixed
- Moved the spi reads after the reset to increase stability during M3 L2R copy
- Bumped luwen version
## 1.2.0 - 06/03/2025
### Fixed
- Updated how local eth board info is calculated to make it agnostic to eth fw version
- bumped tt-tools-common version
- Added traceback printing when catching exceptions in main.
## 1.1.5 - 14/05/2024
### Updated
- Bumped luwen (0.3.8) and tt_tools_common (1.4.3) lib versions
- Removed unused python libraries
## 1.1.4 - 25/03/2024
### Fixed
- Changed detect_chips with detect_chips_with_callback to enable detailed debug info.
## 1.1.3 - 22/03/2024
### Fixed
- Bumped tt-tools-common version to avoid pip discrepancy.
## 1.1.2 - 22/03/2024
### Fixed
- Fixed command line bug when no args are provided.
## 1.1.1 - 21/03/2024
### Fixed
- Fixed reference to pyluwen lib
## 1.1.0 - 12/03/2024
### Added
- Octopus Configuration (4 n150s connected to 1 galaxy)
## 1.0.2 - 12/03/2024
### Fixed
- Dependency bug with tt_tools
topologyethernetmulti-cardrouting
Works on
wormholeblackhole
tt-npe
official
C++ · Apache-2.0 · 15⭐ ·
Network-on-chip Performance Estimator for Tenstorrent Tensix-based devices. Model and estimate NoC utilization before running kernels on hardware.
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## 3.4.0 - 30/07/25
- Bump pyyaml 6.0.1 -> 6.0.2
- Improve error message formatting
- No longer have to use --force for flashing BH cards
## 3.3.5 - 03/07/25
- Bump luwen 0.7.3 -> 0.7.5
## 3.3.4 - 02/07/25
- Bump tt-tools-common 1.4.16 -> 1.4.17
- Bump luwen 0.6.4 -> 0.7.3
## 3.3.3 - 05/06/2025
- Bumped tt-tools-common version to fix driver version check for compatability with tt-kmd 2.0.0
## 3.3.2 - 14/05/2025
- Bump tt-tools-common version to latest
## 3.2.0 - 12/03/2025
### Updated
- luwen version bump to bring inline with tt-smi; provides stability fixes
## 3.1.3 - 06/03/2025
### Added
- luwen version bump to include bh arc init checks
## 3.1.2 - 28/02/2025
### Added
- Support for more BH cards: p100a, p150, and p150c
## 3.1.1 - 06/01/2025
### Updated
- Bumped luwen version to accomodate Maturin updates
## 3.1.0 - 29/10/2024
### Added
- Support for flashing the BH tt-boot-fs file format
- Bumped luwen version to 0.4.6 to allow resets when chip is inaccessible
## 3.0.2 - 17/10/2024
### Fixed
- Unbound variable when exception is thrown when getting current fw-version
## 3.0.1 - 16/10/2024
### Changed
- Bumped luwen version to 0.4.5 to resolve false positives on bad chip detection
## 3.0.0 - 23/08/2024
- NO BREAKING CHANGES! Major version bump to signify new generation of product.
- Added support for p100
## 2.2.0 - 19/07/2024
### Updated
- Added support for an alternative spi flash configuration via a new version of luwen
## 2.0.8 - 14/05/2024
### Updated
- Bumped luwen (0.3.8) and tt_tools_common (1.4.3) lib versions
## 2.0.1 - 2.0.7
- Dependency updates
## 2.0.0
- WH flash release
## 1.0.0
- GS flash release
firmware-updateflashutility
Works on
grayskullwormholeblackhole
SFPI
official
C++ · Apache-2.0 · 14⭐ ·
Tenstorrent SFPU programming interface — TT-enhanced RISC-V GCC and binutils plus header files for programming the Tensix SFPU (vector engine) from kernel code. The compiler toolchain underneath TT-Metalium's SFPU ops.
End-to-end AI applications running on Tenstorrent AI accelerators. Complete application examples from retrieval-augmented generation to image generation pipelines.
A shared repository of model implementations used across TT-Forge frontends — a single source of truth for the models used in testing and benchmarking, rather than duplicating them across frontend repos.
Performance report analysis tool for Tenstorrent Metal operations — analyzes perf traces to surface throughput, bottlenecks, and optimization opportunities.
48 interactive lessons covering the full Tenstorrent developer path — from hardware detection to custom training — with click-to-run commands and hardware auto-detection. Available in VSCode and code-server.
# Changelog
All notable changes to the TT-VSCode-Toolkit will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
---
## [0.1.19] - 2026-07-16
### Added
- **New advanced lesson "Monkeypatching TT-NN"** (order 16, after "Exploring TT-Metalium") — upgrade-safe, smallest-trace patching of TT-NN / TT-Metalium organized by developer goal (observe; fix a bug early; change a default; add something new; and a last-resort source-diff escape hatch), for the TT-QuietBox 2 case where `ttnn` is an installed package with no `~/tt-metal` source tree. Covers the env + import-order rule, the "patch to change behavior, wrap to add behavior" principle (citing Martin Chang's non-invasive `ttPseudoRowMajor` and upstream-first ggml backend), and an AI-agent verification recipe. Validated on p300c — all 19 lesson patterns exercised against real `ttnn` on a TT-QuietBox 2.
- **New reusable template `content/templates/monkeypatch/tt_patches.py`** — a dependency-free patch harness (save/restore `wrap` and `set_default`, `patched` context manager, `version_at_most` guard that zero-pads unequal-length versions, and a `verify` probe helper) with fail-loud missing-target detection, shipped with a self-contained hardware-free `test_tt_patches.py` and a usage README. The lesson embeds the full harness source in a collapsible section for transparency.
- **`tenstorrent.monkeypatch.copyHarness` command** — copies the harness folder into `~/tt-scratchpad/monkeypatch/`, replacing any prior copy (with confirmation) so it mirrors the shipped template, and offers an "Open tt_patches.py" follow-up. Wired into the lesson as a click-to-copy button.
- **`check:monkeypatch-drift` script** (wired into the pre-commit hook) fails if the `tt_patches.py` source embedded in the lesson diverges from the template file.
### Changed
- Excluded `.superpowers/` from the packaged `.vsix` via `.vscodeignore` (was shipping ~1.9 MB of session artifacts).
## [0.1.18] - 2026-07-13
### Fixed
- **PRD-246 — Jeremy's QB2 testing feedback on the first-inference lesson flow:**
- `download-model` — fixed the broken "Step 3: Download the Model" skip link. The anchor pointed to `#step-3-download-qwen3-0-6b`, but the "Step 3: Download Qwen3-0.6B" heading slugs to `#step-3-download-qwen3-06b` (the `.` in `0.6B` is dropped, not turned into a hyphen).
- `download-model` — consolidated the repeated, scattered Hugging Face auth flow. Removed the standalone "Already Authenticated?" pre-check and folded the `hf auth whoami` check into Step 2, so authentication reads as a single sequence (set token → check → log in) instead of appearing in multiple places.
- `hardware-detection` — marked "Check 4: Device Reset" as optional; it is a recovery action, not part of normal detection, and a healthy device never needs it.
- `tt-installer` — removed the redundant `tt-smi` hardwar
Shared helper library of common utilities used across Tenstorrent system tools such as tt-smi, tt-flash, and tt-topology. A dependency rather than a standalone tool.
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## 1.4.17 - 02/07/2025
### Changed
- Loosened requirements on pyproject.toml to make it more compatible in different venvs
## 1.4.15 - 05/05/2025
### Changed
- parse\_reset\_json now returns a ResetInput with stricter typing
## 1.4.14 - 04/02/2025
### Added
- New flags in reset config file generation to disable sw\_version reporting
## 1.4.13 - 23/1/2025
- Removed nr\_hugepages count from compatibility, as hugepages allocation is tricky
and deserves its own widget elsewhere.
## 1.4.12 - 16/1/2025
- Added TTHostCompatibilityMenu to replace Host Info and Compatibility boxes
- Added a count of nr\_hugepages to the TTHostCompatibilityMenu
## 1.4.11 - 30/12/2025
- Updated Luwen version to fix Maturin issue
## 1.4.10 - 16/12/2024
### Changed
- detect\_chips\_with\_callback now takes a print\_status arg
## 1.4.9 - 11/12/2024
### Changed
- A failed reset now results in a fail exit code on BH
## 1.4.8 - 11/10/2024
### Changed
- Updated reset completion logic to handle the case where the bmfw needs to upgrade itself
## 1.4.7 - 11/10/2024
### Added
- Implemented m3 reset option for Blackhole
### Fixed
- Fixed crash during driver version dection when the "extraversion" field is used
- i.e. 1.28-bh
## 1.4.6 - 17/07/2024
### Added
- Reset support of Blackhole
## 1.4.5 - 11/07/2024
### Added
- Bump pyluwen library version (v0.3.8 -> v0.3.11)
- Moved pyluwen v0.3.11 to optional dependencies in pyproject.toml
## 1.4.4 - 21/06/2024
### Added
- Version bump of python dependencies in pyproject.toml (dependabot)
- requests (2.31.0 -> 2.32.0)
- tqdm (4.66.1 -> 4.66.3)
- Pydantic library version bump (1.* -> >=1.2) to resolve: [TT-SMI issue #27](https://github.com/tenstorrent/tt-smi/issues/27)
## 1.4.3 - 14/05/2024
### Added
- Arm platform check and warning for WH device resets in compatibility menu
- Added check for WH device init after reset and prompt user to reboot host if chips are still non recoverable
- Bumped textual (0.59.0) and luwen (0.3.8) lib versions
## 1.4.2 - 04/04/2024
### Added
- Added "silent" flag to WH and GS resets to make them more versatile for use in other tools
## 1.4.1 - 22/03/2024
### Fixed
- removed pyluwen version to avoid dependency issues in other repos
## 1.4.0 - 19/03/2024
### Added
- detect_device_fallible that will provide feedback about chip state during init
### Fixed
- Update min driver version to 1.26 to perform lds reset
- Reset config file uses dev/tenstorrent id
- Catch JSON errors in reset config parsing
- Make nested dirs when initializing reset config path
## 1.3.0 - 06/03/2024
### Added
- Migrated GS Tensix reset to tools_common
- Migrated all related GS data files
- Functions to fetch arc and eth fw versions from telemetry
librarytoolingshared-utilities
tt-toplike
official
Rust · Apache-2.0 · 6⭐ ·
A vibrant htop-style visualizer for Tenstorrent hardware written in Rust. Real-time process and utilization view for TT accelerators.
# Changelog
The **canonical, complete release log lives in [`debian/changelog`](debian/changelog)** —
that's the file the `.deb` packages are built from and where every release is
recorded in full. This file is a friendly pointer plus a summary of the most
recent releases; it deliberately does not duplicate the whole history.
To see everything:
```bash
less debian/changelog # full history
git tag # released versions
```
## Recent releases
### 0.7.33
- **[i] media/diffusion monitoring** — SkyReels / SDXL / z-image servers
(tt-media-inference-server) now show live telemetry instead of a blank panel.
They expose a `tt_media_server_*` Prometheus namespace (not `vllm:`), so a
dedicated parser reads completed generations, the **in-flight `jobs_in_progress`
gauge**, and per-generation timing. The Feeding snake is reused: headline is
generations/min + seconds-per-gen, the body tracks in-flight jobs, and the
panel shows in-flight/done + stage times (no tok/s). Verified against a live
0.15.0 SkyReels box; dedupes the duplicate series that build emits.
- Fix: the legend / help / explain overlay panel truncated its own text (fixed
42-col width vs 50–66-col content, and `Paragraph` clips instead of wrapping).
It now measures its widest line and sizes to fit, clamped to the terminal.
### 0.7.18
- **`--remote <host[:port]>`** — watch a remote QuietBox's telemetry over a
WebSocket stream (plaintext, unauthed: trusted-LAN only). Every visualization
runs against the remote chips; the process panel and `[i]` inference monitor
still describe the **local** machine.
- Remote hardening: the process panel no longer TT-filters local processes under
`--remote`, the Arcade `⚔` duel is suppressed under `--remote`, and backend
status reports last-frame age / flags a stale stream.
- Packaging: WS support is a default-on `remote` cargo feature (opt out with
`--no-default-features`).
### 0.7.17
- **Arcade duel** — the hero now duels the inference snake: a telemetry-true
tug-of-war strip when a model is serving, the `⚔` marker sliding toward
whichever side dominates (chip power/util vs tokens/s + queue depth).
- Per-device power/temp now shows once as a shared strip instead of per section.
- Memory Castle gains a compact 8-column tier before the fleet-grid fallback.
- 1990s BBS/demoscene ANSI chrome (`╔══[ SECTION ]══▓▒░`), themed under grayskull.
### 0.7.16
- **`/theme grayskull`** — an app-wide grayscale palette (a thousand shades of
grey, cyan/purple accents, hot pink as the only hot color). `/theme default`
restores full color; bare `/theme` toggles.
monitoringhtoprustreal-time
Works on
wormholeblackhole
tt-rpm
official
C++ · Apache-2.0 · 6⭐ ·
Cycle-level, execution-driven RISC-V CPU performance model built on Sparta (MAP) with Whisper supplying functional execution, so it runs real ELF binaries — CoreMark, Dhrystone — to completion. The pipeline is YAML-configurable across in-order/out-of-order execution, issue policy, execute granularity, write-port arbitration, and bypass paths, with a modeled L1 I$/D$ plus optional unified L2, per-unit logging, stats reports, and Konata pipeline visualization.
Documentation for the low-level layer of tt-metal: compute LLK APIs and data movement APIs. The data movement side covers the NOC and overlay on Wormhole and Blackhole; the compute side covers Tensix hardware and expected usage of the LLK APIs. Aimed at op and model writers who need to know what the APIs do and how the hardware behaves underneath them.
System setup and support utilities for Tenstorrent hardware — hugepages-setup configures the 1GB hugepages TT ASICs need, and tt-oops collects diagnostic data for troubleshooting. Ships as the tenstorrent-tools deb/rpm.
Tenstorrent backend for vLLM, built on vLLM's standard plugin mechanism — install it alongside vLLM and TT hardware registers itself as a platform whenever `ttnn` is importable. Self-contained: model registration, platform detection, scheduling, worker execution, model loading, async decode, and data-parallel/multi-lane execution all live in the plugin, so nothing Tenstorrent-specific has to land in vLLM core.
A C++ software emulator of the Tenstorrent device-level kernel and host APIs. Run tt-metal kernel and host code on a standard x86-64 Linux machine — no Tenstorrent hardware required.
Network cabling visualizer for Tenstorrent scale-out deployments: describe a target topology and it generates and renders how to physically cable multiple Wormhole or Blackhole systems together. Works in a physical-deployment mode with racking information and a logical-hierarchy mode for clustering/pod groupings, with topology import/export and a Docker deployment path.
Distributes models over the Hugging Face Hub and serves them on Tenstorrent hardware — `tt-kernel serve <namespace>/<model>` pulls a bundle, registers it with the Tenstorrent vLLM plugin, and launches an OpenAI-compatible server. A vLLM bundle ships only adapter code and metadata (weights stay referenced by HF repo id), while legacy kernel-cache bundles package precompiled tt-metal kernel binaries so a model's first run is a cache hit instead of a slow JIT recompile. Explicitly experimental — the bundle format and APIs may change without notice.
Command-line utility that runs a high power-consumption workload on Tenstorrent devices — used for chip testing, burn-in, and validating a system's power delivery and cooling under sustained load.
# Changelog
All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## 0.2.2 - 31/07/2024
### Added
- Added glx reset support
- Threaded start and end of burnin to increase burnin speed
- Added prints to indicate which chip we are currently running on
- Added support for bh harvesting
## 0.2.1 - 16/01/2024
### Bug fix
- Fix for https://github.com/tenstorrent/tt-burnin/issues/6
- BH reports asic temperature as a signed 16_16 int unlike GS and WH
- Added missing support to report BH asic temperatre
## 0.2.0 - 29/10/2024
### Added
- BH burnin support
## 0.1.1 - 14/05/2024
### Updated
- Bumped luwen (0.3.8) and tt_tools_common (1.4.3) lib versions
## 0.1.0 - 04/04/2024
First release of opensource tt-burnin
### Added
- GS and WH burnin support
burn-instress-testpowerhardware-validation
Works on
wormholeblackhole
tt-animatediff
official
Python · Apache-2.0 ·
Generates short, temporally coherent animated GIFs using the AnimateDiff model on Tenstorrent hardware. Phase 1 runs the correct SD 1.4 + MotionAdapter architecture on CPU; Phase 2 accelerates spatial denoising on Blackhole using the TTNN UNet. Produces vibrant 8-frame animations in ~15 s/frame on a P300C.
Official documentation hub for running Tenstorrent accelerators on Kubernetes. Centers on tt-operator (the umbrella Helm chart) and covers Node Feature Discovery, kernel-mode driver (tt-kmd) management, firmware flashing, Prometheus telemetry, Fabric Manager topology resolution, Dynamic Resource Allocation, and multi-node scheduling via JobSet and PMIx.
Browser-based cloud console for exploring AI on Tenstorrent hardware. Run LLM inference, image and video generation, and browse the supported model catalog in-browser — backed by Tenstorrent accelerators. Cloud hardware access and advanced workflows (deployments, agents) available in staged rollout.
Single entry point to the Tenstorrent software stack: `tt update` converges a machine onto the CI-tested "golden" version set, `tt device` covers status/info/reset, and `tt model`/`tt serve` pull weights and bring up tt-inference-server. Commands either run natively or delegate to tt-smi, tt-flash, and tt-installer behind a stable interface, with `--json` output and documented exit codes on every command. Early prototype — the README labels it an internal prototype and several subcommands are still stubs.
Official setup and onboarding guide for the TT-QuietBox 2 — a compact, liquid-cooled AI workstation with four Blackhole accelerators, an AMD Ryzen CPU, 256GB RAM, and 4TB NVMe. Covers hardware specs, first-boot setup, and hands-on learning paths for running pre-loaded models like Qwen3-32B and serving text, image, video, and speech models via tt-inference-server.
Tenstorrent's fork of QEMU that provides the full-system emulation layer behind ttsim. Models the RISC-V cores and system devices of Wormhole and Blackhole so TT-Metalium workloads can boot and run without physical silicon.
Boltz-2 biomolecular model for drug discovery on Tenstorrent Blackhole. Supports single-card and multi-card configurations — QuietBox (4×) and Galaxy (32×). Approaches physics-based FEP accuracy at 1000× the speed.
# Changelog
All notable changes to TT-Bio are recorded here. Versioning is [SemVer](https://semver.org);
releases are cut from a commit that has passed the on-hardware test suite (see `RELEASING.md`).
## [Unreleased]
## [0.6.1] - 2026-08-07
Design gets one verb: `tt-bio design INPUT --model boltzgen|rfd3`, mirroring `tt-bio predict
--model ...`. `tt-bio gen` still works unchanged and is now a hidden deprecated alias for
`tt-bio design --model boltzgen`. ESMC-300M/600M single-sequence embedding is trace-captured,
and RFD3 multi-card design no longer strangles itself on host threads: four cards aggregated
0.68x of a single card before, 3.48x after. `--devices N` is honoured on the single-card embed
and RFD3 design paths, where it had been ignored, and `predict` now says so out loud when a run
finishes on fewer cards than you asked for instead of quietly returning a slower result.
This release also merges back the v0.6.0 release commits, which were tagged but never landed on
`main`. For five days `main` reported version 0.5.0, shipped no 0.6.0 changelog entry, and
carried a `tests/` file that aborted pytest collection so the host suite never ran.
### Changed
- **Unified design CLI** — `tt-bio design INPUT --model boltzgen|rfd3` is now the single design
command, mirroring `tt-bio predict --model ...`. BoltzGen pipeline options (`--steps`,
`--config STEP key=val`, `--num_designs`, `--budget`, `--devices`, `--out_dir`) moved onto the
shared command as boltzgen-scoped flags (`boltzgen` is the default model, so existing
single-model invocations need only swap the verb). The RFD3 checkpoint flag is renamed
`--golden_dir` to `--checkpoint`, with the old spelling kept as a hidden deprecated alias for
one release. The RFD3 engine modules moved from flat `tt_bio/rfd3*.py` files into a
self-contained `tt_bio/rfd3/` package mirroring `tt_bio/boltzgen/`; `import tt_bio.rfd3`
resolves to the package. Numerics are bit-identical: this is a CLI, layout and docs change
only. BoltzGen user documentation moved from the README to `docs/boltzgen-design.md`, and the
README now has one Design section covering both models. (`26cf293a`)
- `fa77e884` predict: warn loudly when a run completes on fewer cards than requested, rather
than silently returning the slower result.
- `7bee7292` embed and `ef0265ef` rfd3: honour `--devices N` on the single-card path.
### Deprecated
- **`tt-bio gen`** — hidden from `--help` and prints a deprecation warning on stderr, then
forwards every argument to BoltzGen unchanged. It still works; it will be removed in a future
release. Use `tt-bio design INPUT --model boltzgen` instead (`gen run X --output out` becomes
`design X --model boltzgen --out_dir out`).
### Fixed
- `280de387` tenstorrent: restore the bf16 construction defaults in fast mode. This was the root
cause of ESMFold2 returning NaN confidence.
- `a657a636` boltz2: use one sample-chunk width for the whole diffusion trajectory and pad the
short
FlashAttention-style attention kernel implemented entirely in on-chip SRAM on the Tenstorrent Grayskull chip using TT-Metalium. Pioneering work in low-level attention on TT hardware.
Meta's UMA interatomic potential running on Tenstorrent Blackhole — energy, forces, and stress for molecules and periodic materials behind an ASE calculator. Its per-edge Wigner rotation runs as a custom tt-metal kernel for a highest-performance uma-s build.
# Changelog
All notable changes to TT-Atom are recorded here. Versioning is [SemVer](https://semver.org);
releases are cut only from a commit that has passed the on-hardware release gate — accuracy
parity, no OOM across the supported size range, no perf or UX regression, and a clean install
smoke (see `RELEASING.md`).
## Unreleased
### Added
- **Multi-card data-parallel fan-out for `tt-atom run`**: pass several structure files with
`--devices 0,1,...` and each card runs a full `Calculator` + relax/MD loop (or the
single-point energy default) for its shard of structures — the high-throughput
virtual-screening path. Per-structure results come back in input order, bit-exact vs the
single-card path (`scripts/_multicard_sim_parity.py`); without `--devices`, multiple
structures run one after another on one card. `--out` is a directory in batch mode and each
written geometry carries its energy and forces. The new `tt_atom.batch.MultiCardSim` pool
backs the CLI and is usable directly; `scripts/multicard_sim_scaling.py` measures the
throughput scaling.
- Orb's batched path (`evaluate_batch`) now applies ZBL pair repulsion union-wide through one
autograd pass, matching the per-system path at short contact.
### Fixed
- `tt-atom run a.xyz b.xyz --relax --devices 0` with two **different-composition** structures no
longer crashes the worker. The multicard worker builds one UMA `Calculator` per reduced
composition and used to call `open_device` once per `Calculator`, so a second composition opened
the same card a second time in one process (`TT_FATAL: No MetalContext instance for context_id N`).
The worker now opens its device once and reuses it across every `Calculator` it builds; the
`Calculator.close()` it owns never closes a device it didn't open.
- `tt-atom run` (multicard) now exits non-zero when any structure fails. The worker has always
caught per-structure errors and returned the other structures' results, but the CLI used to
exit 0 regardless, so a failed structure silently dropped its output. It now reports which
structures failed and exits non-zero while still writing the ones that succeeded.
- Built wheels now include both weight exporters, so automatic UMA and Orb cache misses work
outside a source checkout.
- Fresh UMA and Orb cache misses can download their checkpoints again; explicit
`HF_HUB_OFFLINE=1` still enforces offline use. Concurrent exports now use separate sidecars.
- Release mode now blocks every missing fixture, baseline, required op, and model-family OOM row.
`--allow-gaps` remains available for development diagnostics.
- Release and UX subprocesses always open logical device 0 after `TT_VISIBLE_DEVICES` selects the
physical card.
- Custom-op validation now rejects invalid gate modes and shapes, and program-cache keys include
every operand layout that affects compiled accessors.
- The silicon-melt example now checks the exact-cutoff neighbour graph every step and recaptures
only when
A growing collection of models that use tt-lang for some or all of their implementation. Reference implementations for bringing modern models to the tt-lang DSL.
A Tenstorrent fork of Infocom's Zork I (and more!), running a Z-machine interpreter at least four different ways on TT hardware. The most fun you can have with an AI accelerator.
Discover, load, and benchmark models with a GUI and TUI for tt-inference-server. Makes exploring available models on Tenstorrent hardware as easy as browsing a catalog.
A Tenstorrent-powered claw machine that rewards players with real prizes. The QuietBox 2 runs local AI inference to act as an agent controlling the claw hardware — the OpenClaw AI assistant lesson builds directly on this project.
Three agentic projects running fully on-device: local AI agents on QuietBox 2, a coding assistant powered by Aider against a local inference server, and the OpenClaw AI assistant on QuietBox 2. No cloud APIs — all inference runs on TT hardware.
DFlash: Block Diffusion for Flash Speculative Decoding on Tenstorrent hardware using tt-lang. Combines block diffusion with speculative decoding for faster inference.
Three lesson-projects covering on-device video synthesis: frame-by-frame diffusion with tt-local-generator, native AnimateDiff video animation, and video generation on QuietBox 2. All run entirely on TT hardware with no cloud dependency.
# Changelog
All notable changes to tt-forge-compiletron are documented here.
## [Unreleased]
### Added
- `docs/kv-cache-bench.md` — teaching companion for the StaticCache KV cache
benchmark, explaining the two-graph pattern and why static shapes matter
---
## [1.6.0] — 2026-06-30
### Added
- **StaticCache KV cache decode benchmarking** — `bench_decode.py` now compiles
a second forge graph for the decode step using `transformers.StaticCache`.
The StaticCache is embedded in `KVDecodeWrapper` as a submodule so forge
traces K/V tensors as model state and emits `FillCache`/`UpdateCache` ops.
Falls back to full-recompute for models that don't support `cache_position`.
- `_try_kv_decode()` function — detects model dtype to avoid bfloat16/float32
mismatches, resolves tokenizer from loader or AutoTokenizer, pre-fills cache
on CPU before forge compilation.
- Bestiary `decode_note` field now records the method used per model
("StaticCache KV cache" vs "no KV cache — full recompute per step").
### Changed
- Decode results updated for all 5 stages — GPT-2 2.30→5.52 tok/s, OPT
3.98→5.05 tok/s, Phi-2 1.48 tok/s (new), Falcon 3.30 tok/s (new),
LLaMA-LoRA 2.86 tok/s (new), Gemma-LoRA 2.40 tok/s (new), and more.
---
## [1.5.0] — 2026-06-30
### Added
- **`scripts/bench_decode.py`** — dedicated LLM decode benchmark measuring
TTFT, prefill tok/s, and decode tok/s for all compiled causal LMs.
Subprocess isolation + tt-smi health check prevent hardware lockups.
- **Leaderboard columns** — TTFT, Prefill tok/s, Decode tok/s, Params (M)
replace the old Infer p50 / Throughput columns in `docs/leaderboard.html`.
- 5 benchmark stages: Stage 1 (GPT-2, OPT), Stage 2 (Phi-2, BLOOM, CodeGen),
Stage 3 (Falcon, Allam, LLaMA-LoRA, Gemma-LoRA), Stage 4 (Qwen 2.5,
Phi-1 LoRA), Stage 5 (DeepCogito, DeepSeek Coder, frontier models).
- `params_m` field added to all benchmarked bestiary entries.
- `hf:` loader prefix for frontier HuggingFace models loaded without a
tt-forge-models seed loader.
### Changed
- Bestiary `throughput_unit` relabeled from generic `tok/s` → `prefill_tok/s`
for all 54 causal LM entries to prevent confusion with decode throughput.
---
## [1.4.0] — 2026-06-29
### Added
- **`scripts/install.sh`** — turn-key smart installer: hardware pre-check,
hugepages, disk space, forge venv, XLA venv, mesh descriptor probe,
tt-forge-models clone, stale-shm cleanup. Outputs color-coded summary table.
- **RAM/DRAM budget calculator** — skips models whose weights exceed available
system RAM + per-chip DRAM; prevents OOM crashes at load time.
- **`scripts/setup-venvs.sh`** — minimal venv setup script for clean Ubuntu
24.04 installs on Tenstorrent Blackhole hardware.
- Self-contained patches directory — tt-forge-models fixes applied at
expedition startup without modifying upstream.
- `--ephemeral` / `--evict-failures` flags — evict HF weight cache after
each model to reclaim disk space on small-storage machines.
### Changed
tt-forgemodelsdemocompilation
Image Classification with TT-Forge
affiliated
by
·
End-to-end image classification project using TT-Forge — compile and run a PyTorch classification model on Tenstorrent hardware with no kernel authoring required.
Hardware topology visualizer for Tenstorrent chips — from individual chip to full cluster. Interactive JavaScript visualization of Tensix core layout and NoC connections.
# Changelog
All notable changes to tensix-viz are documented here.
## [1.1.2] - 2026-06-29
### Fixed
- **`TensixViz.autoInit()` is idempotent for `.tensix-viz-container` elements** (`src/chip.js`)
1.1.1 made the `[data-viz]` path idempotent but left the legacy single-chip path unguarded. When
`autoInit()` ran twice (the bundle's self-init plus an explicit host-page call), each
`.tensix-viz-container` canvas received a second `TensixViz` instance — two animation loops drawing
on one canvas, which renders as a doubled/overlapping grid. `TensixViz.autoInit()` now skips any
container already initialized (`container._tensixViz`) and records the instance on it.
### Added
- **Responsive multi-chip canvas** (`tensix-viz.css`)
`.tv-chip-wrapper canvas { max-width: 100%; height: auto; }` — card/system canvases (created
without the `.tensix-viz-canvas` class) now scale to fit a narrow column instead of being clipped
by `.tv-card`'s overflow. Previously this rule had to be patched in by downstream consumers.
## [1.1.1] - 2026-06-25
### Fixed
- **Animation player accepts both script schemas** (`src/chip.js` `_execStep`)
The player dispatched on `step.step` and read `step.cores` only, so scripts authored with the
alternate `{ action, coords }` schema ran zero steps — the Play button (and auto-play) appeared
dead. `_execStep` now dispatches on `step.step || step.action` and falls back `coords → cores`,
so blocks written in either schema animate.
- **`autoInit()` is idempotent for `[data-viz]` elements** (`src/index.js`)
`autoInit()` can run more than once (the bundle's self-init plus an explicit call). For `card`
and `system` vizzes — which append their render into the host element — the second run appended
a duplicate set of chips. `autoInit()` now skips any element already initialized (`el._tensixViz`).
## [1.1.0] - 2026-06-09
### Fixed
- **Heatmap: non-tensix cells no longer painted by heat overlay** (`src/chip.js` `_drawHeatmap`)
Commit 76dca80 added `coreType !== 'tensix'` guards to the pre-built artifacts but never to
the source. The guards are now in `src/chip.js` so the next build preserves them. Without this
fix, DRAM (col 5 on Wormhole), ETH (row 6 on Wormhole), and PCIe (col 8 on Blackhole) cells
were colored by the heatmap overlay and could inflate `maxVal`, compressing the visible range
for all tensix cells.
- **Memory overlay: stale phase not rendered after `reset()` on `showMemory: true` instances**
(`src/chip.js` `reset()` and constructor)
After calling `viz.activate(mode)` followed by `viz.reset()` on a canvas created with
`showMemory: true`, `_memPhase` retained the frozen `_mem` object from the animation closure.
`reset()` calls `render()` at the end, which caused `_drawMemoryLayer()` to run with stale data,
producing a faint DRAM glow and L1 fill bars on an otherwise blank chip. `reset()` now sets
`this._memPhase = null`; the field is also explicitly initialized to `null` in t
Interactive browser-based visualizer of the Tenstorrent Tensix grid architecture. Explore the NoC, core layout, and dataflow patterns without hardware — a great companion for learning kernel programming.
TT-Metalium implementation of Conway's Game of Life as a cookbook recipe. Each generation is a full parallel kernel dispatch over the grid — a clean introduction to stateful compute on Tensix cores.
Particle Life simulation on Tenstorrent hardware — an emergent-behavior N-body system where simple attraction/repulsion rules between species produce complex lifelike patterns. Cookbook recipe demonstrating parallel N-body compute on Tensix.
Seven-module computer science curriculum taught on real Tenstorrent hardware. Covers RISC-V architecture, memory hierarchy, parallel computing, networks and NoC, synchronization, abstraction layers, and computational complexity — all grounded in what is physically happening on the chip.
Eight-lesson series covering the full custom training workflow on TT hardware: dataset fundamentals, configuration patterns, fine-tuning, multi-device distributed training, experiment tracking, model architecture basics, and training from scratch.
Three hands-on TT-Metalium kernel recipes: a Mandelbrot fractal explorer, real-time audio signal processing pipeline, and custom image filter stack. Each recipe is a complete kernel project with full source in the lesson.
htop-style process monitor for GPUs and AI accelerators. Supports AMD, Apple, Huawei, Intel, NVIDIA, Qualcomm — and Tenstorrent. Real-time utilization, memory, and process info in a terminal UI.
Vendor-agnostic orchestration for training, inference, and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Open-source CUDA compiler targeting multiple GPU architectures including Tenstorrent. Compiles .cu files to run on AMD and Tenstorrent hardware without modification.
Booth — Changelog
=================
## Unreleased
### Frontend
- double-precision `fmax`, `fmin` and `fmod`. The ocean kernels in
[Jorge Galvez](https://github.com/JorgeG94)'s do-concurrent benchmarks call
them, and only the `f`-suffixed single-precision forms were recognised
(Zane Hambly, 2026-08-06)
- raise the cap on arguments in one call to 64, and say so when a call goes
past it. Sema stopped counting at 16 and then reported an arity mismatch
against the count it had stopped at, so a correct 23-argument call in the
same ocean benchmarks was rejected and told the wrong number
(Zane Hambly, 2026-08-06)
- #142: parse function pointer declarators, and constructors and destructors
(Zane Hambly, 2026-07-27)
- #144: move the Triton errors into the shared error catalogue
(Zane Hambly, 2026-07-28)
- keep block comment state across lines, so a macro quoted in prose is no
longer expanded
(Zane Hambly, 2026-07-22)
- stop dereferencing anonymous-struct name sentinels as source offsets
(Zane Hambly, 2026-07-22)
### Triton
- let `--cpu` and `--rv64` through the mode gate, and fail with a nonzero
status rather than a silent zero
(Zane Hambly, 2026-07-16)
### HIP
- #137: support bare convergent warp and lane intrinsics
(Maou, 2026-07-27)
### Architecture
- one target per run. Several backends at once all wrote to the same
`-o` path, so you got whichever came last in the registry under the
name you asked for, and a zero exit
- backend contract (`be_desc_t`) with static registration; every
existing backend sits behind the same shape and the driver iterates
`be_list` instead of the copy-pasted if-chain. A backend also owns its
own command line now, so adding a target means one file and one line in
the list rather than editing a shared config struct and the driver's
argument loop. Skeleton in `src/backend/skeleton/` and
`docs/backends.md` for anyone adding a target
(Zane Hambly, 2026-08-02)
### Backends
- seed atomic RMW as divergent in the AMD divergence analysis, so a GEP off
an `atomicAdd` result no longer takes the scalar path and emits
`s_add_u32` with a VGPR source; unblocks per-thread SYSPRINT on AMD
(Zane Hambly, 2026-08-01)
- #138: real high-half multiply on x86-64 and RV64, and an honest refusal
where it cannot be done
(Zane Hambly, 2026-07-26)
- metal: refuse a kernel that uses `double` rather than narrowing it to
`float`. Apple GPUs have no fp64, and quietly halving the precision the
source asked for is worse than saying so
(Zane Hambly, 2026-08-06)
- a divergent return masks lanes instead of ending the wave, so AMD kernels
no longer lose the lanes that did not take the branch
(Zane Hambly, 2026-07-25)
- `--amdgpu` honours `-o`, and the register-plan line goes to stderr rather
than into the middle of the assembly on stdout, where it stopped the result
assembling
(Zane Hambly, 2026-08-06)
### Tensix
- `--tt-chip` selects wormhole or blackhole, so L1 and te
Minimal Python code to access and program the Tenstorrent Blackhole chip directly — George Hotz's exploration of TT hardware programmability with pointed commentary on the architecture.
A complete ML library and compiler in Rust — "from assembly to neural networks" — with a native Tenstorrent backend (src/backend/tenstorrent), autograd, custom kernels, multi-backend support, and Python bindings.
Community-built Tenstorrent architecture simulator written in Python. Runs without hardware — useful for researchers and developers exploring the Tensix architecture offline.
OpenAI Triton compiler plugin for Tenstorrent hardware. Write Triton kernels and target Tensix cores — brings the Triton ML kernel ecosystem to TT devices.
IREE (Intermediate Representation Execution Environment) ML compiler ported to Tenstorrent AI accelerators. Brings the IREE compiler ecosystem to TT hardware.
Tutorial on Tenstorrent hardware for HPC researchers from the RISC-V Testbed project at Edinburgh/EPCC. Covers Wormhole from an HPC parallel-computing perspective.
clpeak-style peak-performance benchmark for Tenstorrent devices using TT-Metalium. Measures theoretical peak throughput across operations — useful for hardware characterization.
High-level parallel programming framework for Tenstorrent accelerators, abstracting TT-Metal into a research-oriented programming model for parallel computation.
Boot stock Linux cloud images on the SiFive X280 RISC-V cores inside Tenstorrent Blackhole AI accelerators. Per-card Rust daemon with virtio-mmio block/net/console and U-Boot/EFI support.
# Changelog
Notable changes per release. Format loosely follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/);
this project does not yet promise SemVer compatibility on the RPC
wire format or library API surface (we're not 1.0).
## Unreleased
V2 virtio-dispatch redesign. The kick ring + completion ring + host-
side throttle that grew up around #184 are gone; in their place is a
per-(slot, queue) dirty bitmap in BRISC L1. The bitmap is level-
sensitive — guest QUEUE_NOTIFY storms coalesce into a single set
byte, so the dispatch path can't fall behind under any burst. Wire
incompatible with 0.9.0; `TENSIX_PROTOCOL_VERSION` bumped 4 → 5.
### Added
- **V2 dirty-bitmap dispatch** (`#187` / `#188` / `#189`). BRISC
writes 1 to `CTRL_OFF_DIRTY[slot][queue]` on every guest
QUEUE_NOTIFY; the daemon's `Dispatcher` clears the byte and
dispatches each pass. Replaces V1's 2048-entry kick ring +
daemon-side `consume_kick_ring_pass` consumer.
- **V2 processed-cursor table** at `CTRL_OFF_PROCESSED`. Daemon
publishes `used.idx` after each successful dispatch so
warm-resume reads cursors directly without re-probing guest
DRAM.
- **`bhx_notify_events_total`, `bhx_dispatch_passes_total`,
`bhx_dispatch_queues_drained`** Prometheus counters surface the
new dispatch path. The burst regression test (`scripts/
soak_virtio_burst.py`) asserts `dispatch_passes_total > 0` to
confirm the workload reached the new path.
- **`scripts/soak_virtio_burst.py`** — multi-queue burst regression
test. Sustains 16-job direct=1 fio randwrite + a tight
`printf` loop to `/dev/console`, samples `/metrics` every 1 s,
and verifies the daemon log contains zero
`kick.*drop|rescue|throttle.*ENGAGE` matches.
- **`DaemonState.chip_reset_this_session`** flag — gates
`maybe_opportunistic_reset_board` so 4-way parallel cold boots
reset the chip exactly once, not once per L2CPU. Without this
the second-and-later resets blip the chip while earlier-booted
L2CPUs hold mmap pages, SIGBUSing their workers.
- **`Dispatcher` (was `KickPoller`)** with documented testability
seam (`CtrlL1Access` trait); `drain_dirty_bitmap` is unit-tested
against an in-memory L1 fake covering all five visit/clear
semantics cases plus the address-formula pins.
### Changed
- **`KickPoller` → `Dispatcher`**, plus `kick_poller` → `dispatcher`
field on `DaemonState`, `tensix-kick-poller` → `tensix-dispatcher`
thread name, `[kick-poller]` → `[dispatcher]` log tag,
`kicks_consumed` → `dispatches_total`,
`last_kick_slot_queue` → `last_dispatch_slot_queue`. Pure
rename; no behavior change. V1 vocabulary scrubbed throughout
the codebase (firmware, daemon, scripts, docs).
- **`CTRL_SIZE` shrinks 36 KiB → 4 KiB**. V2 footprint is ~1.5 KiB;
the rest is reserved for future fields.
- **Stats-page offsets repacked** — V1 `STATS_OFF_KICK_DROPS`,
`STATS_OFF_COMPL_EVENTS`, `STATS_OFF_LAST_COMPL` retired with
V1 (#190); deprecated PRECAP / BLINDCAP / POSTCAP slots dropp
ttas is a hacker-friendly assembler/disassembler for Tensix on Wormhole. It turns assembly into the exact 32-bit words the hardware runs, and turns binaries back into readable instructions using the same shared instruction table.
Comprehensive tutorials for the Tenstorrent software stack in Korean. Jupyter notebooks covering the full developer path from hardware setup to model inference.
Master's thesis implementing and benchmarking five allreduce algorithms (Swing, Recursive Doubling, Bandwidth Optimal, Latency Optimal, Shared Memory) on the Wormhole n150. Bandwidth Optimal achieved best performance, approaching within 2× of theoretical optimal.
Rust crate that exposes the TT-Metal host API through a C++ bridge via cxx.rs — covering device management, program/kernel creation (from source file or inline string), circular buffers, semaphores, runtime arguments, sharded buffers, and MeshDevice workflows, with hardware-backed integration tests.
A compatibility guardrail that continuously monitors whether [tt-metal](https://github.com/tenstorrent/tt-metal) and the official [tt-installer](https://github.com/tenstorrent/tt-installer) build successfully on community Linux distributions that are not part of Tenstorrent's official CI.
A Bazel-built PJRT plugin (libtt.so) providing an XLA backend for Tenstorrent devices. Bundles the tt-xla PJRT implementation with tt-mlir and tt-metal into a single shared object so JAX code runs on Tenstorrent hardware, with patches so sglang-jax works out of the box.
3D Gaussian Splatting rewritten to run on the matrix engine: a polynomial splat and order-independent weighted-sum blending replace exp and depth-sorted alpha, so the pipeline becomes GEMM → activation → GEMM. Renderer + trainer, trained device-resident on a Blackhole p150a.
A small TTNN-facing C++ library (ttprm) for running view-shaped tensor work without first materializing the view in DRAM. Targets Tenstorrent TILE tensors and uses cached device operations to gather/scatter through layout views.
A Gentle Guide: Tenstorrent Card on Arch Linux with Metalium
community
by
· Jul 7, 2024
Step-by-step guide to getting a Tenstorrent card running on Arch Linux with the full Metalium stack. Practical troubleshooting from someone who did it the hard way first.
Thoughts and Logs After Messing with Tenstorrent Grayskull
community
by
· Jun 2, 2024
Honest field notes from getting a Grayskull card running and writing first Metalium kernels. Covers setup pitfalls, processor hangs, memory protection quirks, and what makes Metalium compelling despite early rough edges.
Deep-dive into the Tenstorrent architecture and Metalium programming model — circular buffers, kernel synchronization, NoC routing, and where the footguns are. The honest guide to thinking in Tensix.
Lecture 20 from William & Mary's graduate Computer Architecture course. Frames Tenstorrent in the landscape between GPUs and TPUs, draws comparisons to Cerebras and SambaNova, then dives deep into the Wormhole chip and Tensix core: the 5 RISC-V core design, SFPU, NoC, and dataflow execution model.
Sponsored series of deep technical articles on implementing optimal SFPU kernels for the Tenstorrent Wormhole and Blackhole vector units. Covers where, typecasting, 16/32-bit integer multiplication, cube root, and accurate sin/cos/tan — with cycle counts, assembly walkthroughs, and Blackhole vs Wormhole comparisons throughout.
Structured quaternion, rotor, and phase-aware tensor kernels on ordinary floating-point tensors, plus StructuredBench. Includes CPU/PyTorch references, simulator and emulator paths, and reproducible Wormhole/N300 evidence for quaternion multiply (`qmul`), fused SU(2) composition, and H2A Hamiltonian lowering.
A fused kernel for the Grayskull architecture implementing Transformer self-attention entirely within SRAM. Combines matrix multiply, attention score scaling, and Softmax without DRAM accesses, achieving significant speedups over non-fused implementations.
Exploring Fast Fourier Transforms on the Tenstorrent Wormhole
community
by
· Jun 18, 2025
Ports the Cooley-Tukey FFT algorithm to the Wormhole n300 RISC-V accelerator. The Wormhole draws 8× less power and consumes 2.8× less energy than a 24-core Xeon Platinum for a 2D FFT. ISC 2025.
Assessing Tenstorrent Grayskull RISC-V MatMul Acceleration for LLMs
community
by
· May 9, 2025
Evaluates the Tenstorrent Grayskull e75 RISC-V accelerator for matrix multiplication at reduced numerical precision (BFP8 and LoFi), a fundamental kernel in LLM inference computation.
Porting Strategies for Gravitational N-Body Simulations on Tenstorrent Wormhole
community
by
· May 4, 2026
Evaluates three strategies for scaling an N-body code across multiple Tenstorrent Wormhole accelerators. Builds on the established performance of single-card N-body work to explore parallelism via the on-chip NoC and multi-accelerator configurations.
Accelerating Gravitational N-Body Simulations on Tenstorrent Wormhole
community
Nov 16, 2025
Accelerates an astrophysical N-body simulation on the Wormhole n300. Achieves 2× speedup and 2× energy savings over a highly optimized CPU implementation. SC '25 Workshop.
Numerical Kernels on a Spatial Accelerator: Tenstorrent Wormhole
community
Mar 24, 2026
Implements three numerical kernels and composes them into a conjugate gradient solver on Wormhole. Demonstrates AI accelerators merit consideration for HPC workloads traditionally dominated by CPUs and GPUs. 2026.
Maps 2D 5-point stencil computations onto the Tenstorrent Wormhole RISC-V AI dataflow accelerator via two implementations: element-wise decomposition (Axpy) and matrix-multiplication reformulation (MatMul). Profiling shows the isolated Wormhole kernel is competitive with CPU execution, with PCIe transfers and initialization driving end-to-end overhead; Axpy achieves lower energy than the CPU baseline at large scales. Identifies architectural and software directions for making AI accelerators viable for HPC stencil workloads. 2025.
SwiftNPU: Scalable Shape-Flexible Allocation for Inter-Core Connected NPUs
community
Apr 27, 2026
Makes multi-tenant NPU sharing practical for Blackhole-class hardware using polynomial-time allocation algorithms. Delivers up to 1.37× higher utilization and 1.14× faster workload completion. Up to 890,000× faster than NP-hard baselines.
TileLoom: Automatic Dataflow Planning for Spatial Dataflow Accelerators
community
by
· Dec 17, 2025
Compiler system that automatically generates efficient dataflow plans for tile-based languages on spatial accelerators including Tenstorrent Wormhole. Exploits on-chip network forwarding between processing elements to reduce DRAM pressure.
Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent vs. NVIDIA L40S
community
by
· Mar 24, 2026
Shows that Text-to-Speech inference on Tenstorrent Lightning V2 achieves 4× lower cost than NVIDIA L40S. Applies BlockFloat8 (BFP8) and low-fidelity (LoFi) precision strategies to TTS despite their greater numerical fragility compared to LLMs.
A 6,500-word community deep dive into the Blackhole p100a architecture: the tile model (Tensix, DRAM, SiFive x280 L2CPU, Ethernet, PCIe, NoC arc), firmware startup sequence, MOP micro-op processor, replay buffer, FPU/SFPU sync, and the anatomy of a kernel. From the author of blackhole-py.