Agent tools (MCP)
TT-NN Visualizer ships a Model Context Protocol server so a coding agent can ask a report the questions a person asks it during model bring-up: which operations dominate, where the time went inside them, whether a change helped, and what the run allocated.
It is a read-only surface over the same report readers the web application uses. It starts
no web server, opens no port, and writes nothing to a report. By default it appends usage
events to ~/.ttnn-visualizer/usage/events.log (tool name and outcome only). Set
USAGE_RECORDING_DISABLED=true or create the marker file under that directory to opt out;
see the event logging reference.
Running it
ttnn-visualizer-mcp
The server speaks newline-delimited JSON-RPC on stdin and stdout, which is what an MCP client expects from a stdio server. Register it with your client the way you would any other stdio MCP server — for Claude Code:
claude mcp add ttnn-visualizer -- ttnn-visualizer-mcp
Reports are addressed by path, not by the instanceId the web application uses, so the
server needs no database and no running application.
The tools
Tool |
Answers |
|---|---|
|
What this report contains, and which of the tools below apply to it. Returns a handle the others take. |
|
The costliest operations by device time, op-to-op gap, total percentage, FLOPS, DRAM bandwidth or core count. Each row carries the profiler |
|
Per-zone, per-RISC totals from |
|
Per-operation-code deltas between two reports, largest movement first. |
|
Operations matching a name substring or a call-stack substring, with the ids |
|
One operation: its input and output tensors with shape, dtype, layout and memory config, and what it had allocated. With a linked performance report, also the performance rows it launched and their device time. |
|
Memory footprint per operation, keyed by buffer type and ranked within each, with each type’s largest footprint and the device’s L1 geometry. A floor rather than the peak — see below. |
|
Which operation produced a tensor and which ones consumed it. |
|
What one operation was called with, and where in the model code it came from. |
top_ops, zone_timings and diff_reports read the performance CSVs; the five
operation tools — find_operations, operation_detail, memory_profile,
tensor_flow and operation_provenance — read the profiler report’s SQLite database;
load_report opens both, since its whole job is saying which of the others apply. A capture can carry either, both, or — as far as these tools are concerned —
neither, which is what makes load_report’s answer worth reading first.
Call load_report first. It reports what is answerable rather than making you discover it
one failed call at a time, because the report kinds are independent: a performance-only
capture has no operation graph and no tensor data, and a report with no device profiler log
cannot answer zone_timings. Each database tool is checked against the tables it actually
reads, so a truncated capture that holds operations but not buffers is reported as
answering find_operations and nothing else, rather than advertising every tool and
failing most of them.
What the answers mean
Five properties are deliberate, and worth knowing before you act on a number.
Every tool returns an aggregate or a bounded slice, never the table. A performance report runs to tens of thousands of rows of about thirty-five fields. Limits are capped server-side, so asking for more returns the cap rather than the report.
No rows are hidden from you. The web application’s performance view hides host operations by default and can narrow to a signpost range. Those are display choices, and an agent handed them unannounced would reason confidently about a partial set. The tools read the report with host operations included, signpost markers kept, and no signpost range.
One filter is applied, deliberately: per-device rows are merged, so an operation that ran on eight devices is one row rather than eight. That is the unit an operation is reported in rather than a subset of the report — but it is a choice, so every response repeats the projection it used and you can see it.
Allocation figures are per bank, and nothing adds one memory type to another.
memory_profile and the allocations in operation_detail come from the report’s
max_size_per_bank column. That column is divided by the bank count of its own memory
type, and those counts differ — so a DRAM figure plus an L1 figure is two denominators in
one integer rather than a quantity. memory_profile therefore keys everything by buffer
type and ranks operations within a type, never across types.
A per-bank L1 figure is comparable to the l1_bank_size returned beside it, which is a
conservative bound. A device-wide total cannot be derived from the response, and the
response says so rather than offering a formula: multiplying by the bank count is only
right for a buffer interleaved across every bank, and most are not — on a local resnet50
capture the operation holding the largest L1 footprint uses 56 of 64 banks, so multiplying overstates
it by 14%, and 16-bank operations in the same report by 4x. How many banks a buffer
actually occupies lives in page-level data that no tool exposes. The report carries no
DRAM capacity at all, which is likewise stated rather than left as a gap.
What operation_detail does give you is each tensor’s memory_config, whose
memory_layout says whether a tensor is interleaved or sharded, and whose shard_spec
carries the grid and shape when it is — so you can tell why a per-bank figure does not
scale by the device’s bank count, rather than only being told that it doesn’t.
A tensor’s size is a whole-tensor byte count — where the report carries that column.
Where it does not, which is the common case, the report’s own query substitutes the
per-bank allocation figure, and the response labels it bytes_per_bank accordingly. So
read tensor_size_unit rather than assuming the two figures are in different units.
memory_profile reports a floor, not the peak. It reads the buffers table, which
records tensor allocations only. A circular buffer never appears there, and neither does a
tensor allocated and freed inside a single operation. Across the local captures circular
buffers are 52–100% of the real L1 peak wherever a peak has any, and intra-op tensors a
further 24–38%; on resnet50_may08_1841 the tool reports 487,424 — a figure eighteen
operations share — where replaying the captured graph gives 1,159,840 at operation 29,
2.4x out and not the same operations. Two of the local reports hold almost no L1 tensors at
all, so the tool reports a figure near zero for a model that is saturating L1. Do not use
this figure to answer an out-of-memory question; neither the true peak nor the operation
holding it can be derived from it.
memory_profile groups by operation because the report records what was live at each
operation rather than what that operation allocated. A per-operation sum is therefore the
footprint at that point in the run, and the largest of them is that type’s largest
footprint. It is a maximum and never a sum across the
run: buffers persist across operations, so adding them would count one allocation once per
operation it stayed live through. Because resident memory barely moves, that figure is
usually shared by many operations, and operations_at_peak says how many — the difference between
a single operation you can go and fix and a plateau across the whole run.
A multi-host report is read one rank at a time, and refuses rather than guess. Operation ids restart at 1 per rank, so
reading every rank at once would collide operations that merely share an id. The database
tools default to rank 0, take a rank argument, and name the rank they read in every
response along with a caveat saying the figures describe that rank rather than the job. A
rank the report has no rows for is refused rather than answered empty, since “no
operations at rank 9” otherwise reads as a fact about the run. And where a report carries
rank on its operations but not on the table holding the figures — a schema mix that
would silently union every rank and attribute the total to whichever rank you asked for —
the call is refused instead, because a caveat naming a rank the numbers do not describe is
worse than no answer.
A total on a partitioned run carries a caveat. When a report spans more than one
sub-device, operations on different sub-devices can run concurrently, so summed device
time, total percentages and op-to-op gaps overstate elapsed time. Responses that include a
total say so, and name the sub-devices involved — top_ops and diff_reports alike, since
a delta between two unsound totals is unsound the same way.
Totals are only reported for metrics a sum means something for: device time, op-to-op gap and total percentage. DRAM bandwidth, FLOPS and core count are per-operation figures, and adding them across a report would produce a number that looks authoritative and is not.
A null bound does not mean an operation is fine. Every row from top_ops and every
perf_rows entry from operation_detail carries bound_analysis, which says which roofline
model tt-perf-report ran: full (matmuls) derives DRAM, FLOPs and the bound, or leaves the
bound null when the trace lacked the inputs the model needs; flops_only (convolutions)
derives FLOPs alone, so DRAM and the bound are never set; none (everything else) derives
nothing, so a blank there says nothing about whether the operation is a bottleneck. A SLOW
bound means the full model ran and neither DRAM nor FLOPs reached 65% of peak. Rows also carry
op_category, which tt-perf-report 1.4.0 sets to Compute, CCL, DM, TM, Host or
Other and leaves null on signposts. The list can grow in later releases (CCL was added
recently), so treat an unfamiliar value as a category, not an error. Other is time no
category explains.
zone_timings carries a caveat of its own: cycles are summed across every core that ran
the zone, so they measure occupancy rather than wall-clock duration. A core is counted per
device — the captures we test against span 8 and 32 PCIe slots, and coordinates alone
repeat on each — so a zone can report thousands of cores on a multi-device run. Durations
need the device log’s type column to pair zone starts with ends; a capture without it
reports occurrence counts only, and a capture that stopped mid-zone reports how many starts
and ends failed to pair so a partial total does not read as a complete one.
A performance row and a profiler operation are linked by order, not by a shared id.
Neither report records the other’s id, so when a handle is loaded with both, the tools
align the device operations each profiler operation launched (from its captured graph)
against the performance rows, in order — the same match the application uses to link a
memory report to a performance report. top_ops then gives each row an operation_id,
and operation_detail lists an operation’s perf_rows with device time in microseconds;
its own duration is host seconds.
A device operation whose launch threw — an allocation failure, for one — never reached the
device and has no performance row, so the match leaves it out. Left in, it would put a name in
the order that no row answers, and a report with a failure would usually not link at all. Newer captures mark its scope aborted; older ones leave it
unclosed. The operation keeps any earlier device operations that did run.
Every response that carries the link says how it went in operation_link:
linkednames the rank the ids belong to — the CSV records no rank, and rank 0 is what the application links against — andmatched_rowssays how many rows linked.operation_detailread at any other rank reportsunavailableand lists noperf_rows, because ids restart per rank.unlinkedmeans the two sequences do not line up — usually two different runs, or a performance capture that stopped early — and everyoperation_idis then null rather than guessed.unavailablemeans the match was never tried: one half is missing or cannot be read, the profiler report records no captured graph, one of its captured graphs is not readable, or a multi-host report carries norankon the graph or device tables. A link that skipped an unreadable graph could still align and pair rows with the wrong operations, so none is attempted.
On a linked pair an operation_id is still null for a row with no profiler operation:
a host op, a signpost, or a row past the end of a profiler capture that stopped before
the performance one did. The match tolerates trailing rows, as the application’s does.
top_ops counts such rows in operation_link.unlinked_by_reason — signpost,
host_op and past_profiler_capture — totals them in unlinked_row_count, and names
the first row past the profiler capture in first_past_profiler_capture_id. Read the
counts there rather than counting null ids among the ranked rows: the ranking leaves out
any row with no value for its metric, which is every signpost, so matched_rows and the
rows returned do not add up to the report.
The link is resolved once per handle, on the first top_ops or operation_detail that
needs it, and reads every captured graph at the linked rank. On a handle with both
reports, the first operation_detail therefore also generates the performance report,
which can take seconds on a large capture; every later call reuses both. A link that
failed on a read that might pass next time — a locked database, a temp-file error in
tt-perf-report — is not kept, so the next call tries again.
An operation id is not the end of the answer. memory_profile names the operations
holding its largest footprint, and operation_provenance turns such an id into the two things you need
to act on it: the arguments it was called with, and the innermost stack frame’s file,
line, function and source line.
top_ops names an operation too, but its id is a row of the performance CSV and does
not belong here (diff_reports returns no id at all: it groups by op code): the two ids are small integers in overlapping
ranges, so passing one where the other belongs answers about a different operation
rather than refusing. Pass a top_ops row’s operation_id instead, or search for the
name with find_operations. Frames are recorded innermost first, so the call site is the ttnn.<op>
call in the model code with no guessing about which frame is yours, and it is the same
frame the operation details panel shows for that operation.
Argument values are capped, with truncated and full_length set when one was cut, and
truncated_count saying how many of them were: truncation is the normal case rather than
the exception — 28% of the argument values across the local captures run past the cap,
and the longest is 177,717 characters, because a tensor repr or a Conv2dConfig lands in
one.
Two different absences are reported apart. A capture that records no trace for an operation says so. A capture that records a trace this parser cannot read — a native C++ backtrace, which 8 of 86 local captures write for every operation — says the call site is unknown rather than absent, because those are different facts and only one of them is about the operation.
operation_provenance takes a limit for the outward chain. The call site is reported
separately and never counted in it, so a small number is usually right: a real trace runs
to tens of frames, most of them the test harness that invoked the model.
find_operations takes called_from for the inverse question, which is the one that
follows: not “what about this operation” but “what about every operation from that
helper”. It matches anywhere in the stack, so a layer module finds the operations it
called through helpers rather than only those it called directly.
Limitations
Accuracy questions are not answerable. Both tensor-comparison tables exist in every local capture and are empty in every one, so there is nothing to build a PCC tool on.
Page-level memory questions — fragmentation, or the per-bank detail behind the web application’s memory plot — are not exposed. That data runs to millions of rows on an ordinary capture and needs a different shape than a tool response.
Device profiler logs carry named zones only where a kernel was instrumented to emit them.
Most captures contain only the default firmware and kernel zones — BRISC-FW,
BRISC-KERNEL and the same pair for NCRISC and TRISC, plus ERISC where a capture
used ethernet cores — so on those reports zone_timings describes RISC-level phases rather
than named model operations. A capture whose kernels do emit named zones reports them
alongside the defaults.
The transport is a minimal JSON-RPC implementation rather than the official MCP SDK, which keeps the server free of additional dependencies while the tool set settles.