n150 n300 T3000 p100 p150 p300c Galaxy 20 min Draft

Model Architecture Basics

Every field you've been editing under transformer_config: in Configuration Patternsnum_heads, embedding_dim, num_blocks, num_groups, theta — names a real piece of the model. This lesson is the map: what each piece does, and why it exists, so those fields stop being magic numbers before you set your own for a from-scratch run in Training from Scratch.

This is a conceptual tour, not a build. If you want to write every one of these components by hand — RoPE, grouped-query attention, SwiGLU, a TT-Lang kernel — that's the from-scratch arc, linked throughout and again at the end.

What You'll Learn

Time: 20 minutes | Prerequisites: Understanding Custom Training and Configuration Patterns


Where This Fits in the Track

graph LR
    A[Understand] --> B[Datasets]
    B --> C[Configuration]
    C --> D[Fine-tuning]
    D --> E[Multi-Device]
    E --> F[Experiment Tracking]
    F -.-> G[Architecture Basics]
    G -.-> H[From Scratch]

    style G fill:#1B8EB1,stroke:#092221,stroke-width:3px

The Shape of a Transformer

Text goes in, a token comes out, and in between the same block repeats num_blocks times:

graph LR
    A[Text] --> B[Tokenize]
    B --> C[Embed + position]
    C --> D["Block × num_blocks"]
    D --> E[Output projection]
    E --> F[Next-token probabilities]

Each block does the same two jobs, in order — gather context, then think about what was gathered:

graph TD
    A[Input] --> B[RMSNorm]
    B --> C["Attention (GQA)"]
    C --> D["+ residual"]
    D --> E[RMSNorm]
    E --> F["MLP (SwiGLU)"]
    F --> G["+ residual"]
    G --> H[Output]

The + residual arrows matter as much as the boxes: each sub-layer's output is added back onto its input, not used to replace it. That's what lets you stack num_blocks of these without gradients vanishing on the way back down. Five components make up everything above: tokenization, embeddings, attention, the MLP, and normalization.


1. Tokenization

A model can't consume raw text — it consumes integer IDs from a fixed vocabulary. Two ends of a spectrum:

The trade-off is direct: vocabulary size sets the size of the embedding table (below) and the final output layer — both scale with vocab_size. Want to see a BPE tokenizer built from raw bytes, merges and all? That's Tokenizer & Data in the from-scratch arc.

2. Embeddings and Position

Two lookups feed every model, combined into one vector per token:

Older models (GPT-2, BERT) learned a fixed table of position vectors, added to the token embedding. Modern models — Llama and its descendants — use RoPE (Rotary Position Embeddings) instead: position is encoded as a rotation applied directly to the query and key vectors inside attention, not as a separate table added up front. RoPE generalizes better to sequence lengths longer than anything seen in training, and it's why theta (RoPE's base frequency) is a transformer_config: field instead of a learned parameter.

Build the embedding table and RoPE's rotation math by hand in Embeddings & the Residual Stream.

3. Attention

Attention is how a token gathers context from every other token in the sequence before deciding what it means. Each token projects itself into a query (what am I looking for), a key (what can I offer), and a value (what information do I actually carry). Every query is scored against every key; the scores become weights over the values.

Multi-head attention runs several of these in parallel — splitting the embedding dimension across num_heads — so different heads can specialize in different kinds of relationships (syntax, coreference, long-range dependency) instead of averaging them all into one.

Modern models add one more move: grouped-query attention (GQA). Instead of giving every query head its own key/value heads, several query heads share one KV head — num_groups of them instead of num_heads. Fewer KV heads means a smaller KV cache at inference time, for a small accuracy cost. It's why Llama-3 and the models in this track use GQA, and why num_groups is a field you set alongside num_heads rather than a fixed multiple of it.

Full derivation — Q·Kᵀ, scaling, causal masking, softmax, the GQA head-sharing, and a hand-authored TT-Lang kernel for the whole thing — lives in Attention from Scratch.

4. The MLP (Feed-Forward Network)

Where attention mixes information across tokens, the MLP processes each token's vector individually — a small two-layer network applied at every position. It's unglamorous, but it's where most of a model's parameters actually live: in models this size, the MLP typically accounts for well over half the total parameter count, because its inner dimension is usually several times embedding_dim.

Older models used a plain ReLU or GELU activation between the two linear layers. Modern models use SwiGLU — a gated variant that's more expressive at the same parameter count, at the cost of a third weight matrix. It's the default in Llama-family models, and in the block you'll assemble in the from-scratch arc.

5. Normalization

Stacking num_blocks layers means activations can drift or explode as they pass through. A normalization step before each sub-layer keeps values in a stable range so training doesn't diverge.

Normalization has almost no parameters (one scale value per dimension), so it doesn't move your parameter count. It moves whether your loss curve is a smooth descent or a spike.


Build It Yourself: The From-Scratch Arc

Everything above is the concept. If you want the code — every matrix multiply, every rotation, every softmax, written out and then re-expressed as a TT-Lang kernel — that's a different track, starting from Pick Your Altitude:

That arc builds the exact same nanollama3 architecture — RoPE, GQA, SwiGLU, RMSNorm — that ttml runs for you in this track. Same design, two altitudes: use the framework, or write the framework's insides yourself.


From Concepts to Config

Every concept above already has a name in transformer_config: — that mapping, plus the full field reference, safety checks, and worked examples, lives in Configuration Patterns:

Concept Field
Attention heads (multi-head) num_heads
GQA key/value groups num_groups
Embedding width embedding_dim
Depth (blocks stacked) num_blocks
RoPE base frequency theta
Vocabulary size vocab_size

Bigger embedding_dim and more num_blocks mean more parameters, more memory, and more compute — in roughly the trade-offs you'd expect: width scales every matrix in a block, depth scales linearly. In Training from Scratch, you'll pick actual values for these fields and watch a model this small learn from nothing.


Key Takeaways

Five components, in a repeating loop: tokenize once, then embed → attend → process (MLP) → normalize, num_blocks times, then project to output.

Residual connections carry the signal forward — each sub-layer adds to its input rather than replacing it, which is what makes deep stacks trainable.

Modern models made four specific upgrades: RoPE over learned position tables, GQA over plain multi-head attention, SwiGLU over ReLU, RMSNorm over LayerNorm. All four show up in ttml's config, and all four get built by hand in the from-scratch arc.

The MLP holds most of the parameters — attention decides what to look at, the MLP does most of the actual work.

Every concept here is a transformer_config: field — this lesson explains the "why," Configuration Patterns is the authoritative "how."


Next Steps

Next: Training from Scratch — set real values for num_heads, embedding_dim, and num_blocks, initialize a model from random weights, and watch it learn from nothing.

Or build every piece by hand: start the from-scratch arc at Pick Your Altitude and work through the embedding table, attention kernel, and transformer block yourself, TT-Lang and all.


Additional Resources