n150 n300 T3000 p100 p150 p300c Galaxy 30 min Validated

Training from Scratch

Every other lesson in this track hands train_nanogpt.py a checkpoint to load. This one doesn't hand it anything — you launch a job that starts from random weights and watch it become a model, entirely from the numbers a loss curve prints to your terminal.

That's the whole job: pick a config, launch it, watch the loss, checkpoint along the way, generate from a checkpoint, and know how to scale the next run up. This lesson does not ask you to hand-write a transformer's internals — Model Architecture Basics already toured what num_heads, embedding_dim, num_blocks, and theta mean conceptually. If you want to write RoPE, grouped-query attention, and a SwiGLU block yourself instead of pointing ttml at a YAML file, that's a different track entirely — see Build It Yourself near the end.

Every number below — the loss curve, the wall-clock time, the generated text — is copied verbatim from a real training run against this extension's verified ttml build, on a Blackhole® p300c. Nothing here is projected, and nothing here overclaims what the output actually reads like.

What You'll Learn

Time: 25-30 minutes (5-10 min hands-on, ~3.5 min hardware run) | Prerequisites: Model Architecture Basics and Configuration Patterns


Where This Fits in the Track

graph LR
    A[Understand] --> B[Datasets]
    B --> C[Configuration]
    C --> D[Fine-tuning]
    D --> E[Multi-Device]
    E --> F[Experiment Tracking]
    F -.-> G[Architecture Basics]
    G -.-> H[From Scratch]

    style H fill:#1B8EB1,stroke:#092221,stroke-width:3px

Set Up the Job

Same ttml build every other lesson in this track uses. If you haven't built it, Fine-tuning Basics covers the Install tt-train command and the std::bad_cast fix in full — this lesson assumes that's done.

Set your environment honoring any value you've already exported, rather than overwriting it:

export TT_METAL_HOME="${TT_METAL_HOME:-$HOME/tt-metal}"
export TT_METAL_RUNTIME_ROOT="$TT_METAL_HOME"
: "${TT_METAL_ARCH_NAME:=wormhole_b0}"   # set to blackhole for p100 / p150 / p300c
export TT_METAL_ARCH_NAME
export TT_LOGGER_LEVEL=FATAL
cd ~/tt-metal/tt-train/sources/examples/nano_gpt

You'll need a Shakespeare corpus — either the one you built in Dataset Fundamentals, or the copy tt-metal ships at tt-train/data/shakespeare.txt, which is what the run below actually used.


Launch: the nanollama3_char Config

tt-train/configs/model_configs/ ships two architecture families for this size of job: nanogpt* (GPT-2-style — LayerNorm, learned position embeddings, plain multi-head attention) and nanollama3* (Llama-3-style — RoPE, RMSNorm, SwiGLU, grouped-query attention). Fine-tuning Basics ran the GPT-2-style config. This lesson features the modern one — nanollama3_char — because it's the exact architecture the from-scratch arc builds by hand, component by component (see Build It Yourself below).

The shapes that matter, quoted from the real files in tt-train/configs/:

Setting Value Source
Architecture model_type: llama — RoPE (theta=500000), RMSNorm, SwiGLU, grouped-query attention model_configs/nanollama3_char.yaml
Heads / KV groups 6 heads, 3 groups (2 query heads share each KV head) same
Embedding dim / blocks 384 / 6 same
Context length 256 characters same
Parameters 9,810,816 (~9.8M) printed at model creation
Tokenizer Character-level, auto-detected — 68 unique characters, rounded up to a tile-friendly 96 printed at data load
Batch size 64 training_configs/training_shakespeare_nanollama3_char.yaml
Optimizer AdamW, lr: 0.0003, weight_decay: 0.01 same
Checkpoint interval every 500 steps (model_save_interval: 500) same
Device mesh [1, 1] (default, single chip) — p300c and p100 count as one chip here, exactly like n150 no device_config: block needed

Launch it. The config's own max_steps: 5000 is overridden on the command line to run 3000:

python train_nanogpt.py \
  --config training_shakespeare_nanollama3_char.yaml \
  --data_path ~/tt-metal/tt-train/data/shakespeare.txt \
  --max_steps 3000 \
  --fresh \
  --model_save_path ~/tt-metal/tt-train/checkpoints/ct8_nanollama3

Before committing to the full 3,000-step run: sanity-check your setup with a much shorter smoke test — just swap --max_steps 3000 for --max_steps 20. On this hardware that finishes in about 14 seconds, plenty to confirm the config loads, the data path resolves, and the device initializes. It still pays the one-time kernel-compile cost on step 1, same as the full run — a smoke test skips steps, not the compile tax.

--fresh matters: it says "ignore any existing checkpoint at this path, start from random initialization." That's the entire meaning of "from scratch" — everything else in this command is the same job-launching mechanic Fine-tuning Basics already used.


Watch It Converge — The Real Curve

This ran on this extension's Blackhole p300c, against tt-metal v0.73. Total wall clock for 3000 steps: 200.74 seconds (~3.3 minutes). Steady state: ~65 ms/step, roughly 16.5 TFLOPS, ~11% model FLOPS utilization (MFU) — against a mesh peak of 148.5 TFLOPS (bf16, 1 device).

Step Loss Checkpoint written
1 4.6875
500 1.3516 ct8_nanollama3_step_500.pkl
1000 1.0938 ct8_nanollama3_step_1000.pkl
1500 0.8164 ct8_nanollama3_step_1500.pkl
2000 0.5156 ct8_nanollama3_step_2000.pkl
2500 0.2891 ct8_nanollama3_step_2500.pkl
3000 (final) 0.1836 ct8_nanollama3_final.pkl

Loss 4.6875 at step 1 sits close to ln(96) ≈ 4.56 — the entropy of guessing uniformly among 96 possible next characters. That's the honest random baseline. By step 3000 that error is down to 0.18, a much steeper drop than Fine-tuning Basics's GPT-2-style run saw over the same 3000 steps (loss 1.406, on the same corpus and step budget). model_save_interval: 500 is why a checkpoint lands every 500 steps automatically — the table above is what actually appeared on disk, no extra flag required.


Generate — and Read It Honestly

Load the final checkpoint and generate, using the same --prompt / --model_path flags every config accepts:

python train_nanogpt.py \
  --config training_shakespeare_nanollama3_char.yaml \
  --prompt "ROMEO:" \
  --model_path ~/tt-metal/tt-train/checkpoints/ct8_nanollama3_final.pkl \
  --max_new_tokens 300 --temperature 0.7 --top_k 50

Actual output, verbatim, from the checkpoint at step 3000 (loss 0.18):

etwaiynwiyounismanot ather bucoution.

LAGENIAYO:
Ahe imabaplart wellong there thou in priscian the racom to the stiffot will and years son,
There is not there in the mother we should sun
yet thou must be that duke of him so submiss'd
From the cause of thy bestray'd the death,
Must I that had body t

Read this for what it is. There's real structure: an ALL-CAPS speaker name (LAGENIAYO:) followed by a colon, line breaks, dialogue layout, an apostrophe used correctly (submiss'd, bestray'd). There's a genuine mix of real and invented words — "the," "there," "in," "we," "should," "sun," "thou," "must," "that," "from," "cause," "death" are real; "etwaiynwiyounismanot," "priscian," "racom," "stiffot," "bestray'd" are not. This is not coherent Shakespeare, and it is not correct grammar. It's the same class of output Fine-tuning Basics's GPT-2-style run produced at loss 1.406 — structure and a scattering of real words, no more.


The Overfitting Lesson

Here's the part worth sitting with: this run drove loss to 0.18 — nearly eight times lower than Fine-tuning Basics's 1.406. If loss were the whole story, this output should read dramatically more coherent. It doesn't. Both runs land in the same tier: recognizable structure, a handful of real words, mostly invented syllables.

That gap between "loss went way down" and "text didn't get more readable" is the lesson. tt-train/data/shakespeare.txt is about one megabyte of text. A 9.8M-parameter model has more than enough capacity to start memorizing that corpus's exact character sequences well before it has enough exposure to learn general English structure from them. Driving train loss to 0.18 on a dataset this small is overfitting, not mastery — the model is increasingly good at predicting this specific text, not increasingly good at language. Low loss on a tiny corpus is not a proxy for coherent output, and this run is the concrete evidence: a much lower loss bought no visible improvement in readability.

Real coherence needs scale, not just more steps against the same small file. Train It & Run for Real — the from-scratch arc lab that builds this exact nanollama3 architecture by hand — makes the same comparison against Mini-LLM, the project this whole from-scratch design follows: ~80M parameters, 361M training tokens, ~5 hours on a single A100, to get language that actually reads as language. Nine million parameters and one megabyte of characters, however low you push the loss, isn't that project — it's a controlled demonstration that the training mechanism works.


Scaling the Job

Three independent knobs, each with a real config to point at:

More steps. The featured config's own default is max_steps: 5000, not the 3000 this lesson ran — try it, but expect the same overfitting ceiling above, not qualitatively better prose, on this same corpus.

A bigger model. tt-train/configs/model_configs/ ships larger llama-family configs on the same architecture family — nanollama3.yaml (same 6-head/6-block shape, but a real 32,000-token BPE vocabulary instead of characters) and llama3_gpt2s_size.yaml (12 heads, 12 blocks, embedding_dim: 768, GPT-2-small-sized). Point --config at a training config referencing one of these via its model_config: field — every transformer_config: field maps to the concepts Model Architecture Basics covers, and to the DRAM math The Transformer Block & the Model works through for scaling toward Mini-LLM's ~80M-parameter target.

More data. Swap --data_path for a larger plain-text corpus — train_nanogpt.py takes any text file, not just Shakespeare. Dataset Fundamentals covers building one.

The mesh_shape boundary. Every config above still runs on a single chip (mesh_shape: [1, 1], the default) — p300c, p100, or n150, and everything in the nano-to-~80M range fits comfortably in one chip's DRAM. But you don't have to stay on one chip: a TT-QuietBox® 2 is a mesh — four Blackhole chips wired in a P300_X2 ring (a 2×2 mesh), not four independent chips — and data-parallel training across them scales near-linearly (measured ~1.95× on 2 chips, ~3.98× on 4). tt-train ships multi-chip examples (training_shakespeare_nanogpt_ddp_n300.yaml sets enable_ddp: true, mesh_shape: [1, 2]), and Multi-Device Training is now hardware-verified on a TT-QuietBox 2 — including the one non-obvious catch (you must set TT_MESH_GRAPH_DESC_PATH for 2/4-chip Blackhole or the fabric router times out). Concretely: the ~80M nanollama3 architecture trains from scratch across all four chips in ~2.4 hours to the structure-and-vocabulary tier; reaching coherent prose is data-bound (the Mini-LLM reference used 361M tokens), not a matter of more chips.


Build It Yourself: The From-Scratch Arc

Everything above configures and launches ttml — you never touch a matrix multiply. If you want to write this architecture instead of configuring it, that's a different track, the "Build an LLM from Scratch" arc, starting at Pick Your Altitude:

Same architecture, same config shape, two altitudes: point ttml at a YAML file (this lesson), or write every gradient step yourself (that arc).


Troubleshooting

ImportError: No module named 'ttml' or std::bad_cast on import ttml

Covered in Fine-tuning Basics — rebuild _ttnn.so after enabling tt-train, don't do a partial --target _ttml build.

RuntimeError: Device out of memory

Reduce --batch_size (default 64 for this config) — override on the command line, e.g. --batch_size 32.

Loss stays near ln(vocab_size) and never drops

Check the data actually loaded — train_nanogpt.py prints dataset size and vocabulary size at startup; if either is missing or zero, --data_path is pointing at the wrong file.

No checkpoint file after training

Three real causes, in order of how often they bite:

Multi-chip DDP dies at Fabric Router Sync: Timeout

On a TT-QuietBox 2 a 2- or 4-chip run can time out at fabric-router sync during mesh open — and it survives both a full reboot and tt-smi -r, so it looks like dead hardware. It isn't: ttml ships default mesh-graph descriptors only for T3000/Galaxy, so on a 2/4-device Blackhole mesh you must set TT_MESH_GRAPH_DESC_PATH yourself. The full fix (with the [1,4] ring descriptor) is in Multi-Device Training.

argument --resume: expected one argument (auto-resume is broken)

Any run without --fresh triggers auto-resume, which currently injects an empty --resume and dies in argparse. Run with --fresh and checkpoint often instead of relying on resume.

Device open hangs or times out on p300c / TT-QuietBox 2

First check: is another job holding the device? Only one process can own the mesh, so a live training run makes any second job (like inference) hang at device open — that's contention, not a fault. If nothing else is running, tt-smi -r to reset the board and retry; a hard-killed ttml process can wedge the device, so always let it close cleanly.

Generated text loops or turns to word-salad

Same root cause both ways: an undertrained model plus a bare decoder. The stock sample_greedy has no repetition penalty, so a low-maturity model loops ("and Ben and Ben"); bolting on a strong penalty just turns the loop into word-salad. The real fix is training maturity (loss well under 1.0 for TinyStories-scale text) plus a gentle repetition penalty — not the penalty alone.


Key Takeaways


What's Next

Next: Experiment Tracking — capture runs like the one above to a file (or Weights & Biases) instead of watching numbers scroll past in a terminal, and compare hyperparameter variations properly.

Or build every component by hand: start the from-scratch arc at Pick Your Altitude and work through the tokenizer, embeddings, attention kernel, and training loop yourself — culminating in Train It & Run for Real, which trains this exact architecture with code you wrote.


Additional Resources