The Full Transformer
After this lesson, you will be able to:
- Follow data from raw text through the entire Transformer, step by step
- Understand residual connections (information highways) and why they prevent vanishing gradients in deep stacks
- Distinguish pre-norm from post-norm layer normalization and know why modern LLMs switched to pre-norm
- See the feed-forward network as an expand-then-compress thinking step that holds two-thirds of the parameters per layer
- Know the three flavors of Transformers: encoder-only (BERT, bidirectional), decoder-only (GPT, causal), and encoder-decoder (T5)
- Understand the encoder-decoder cross-attention mechanism: how the decoder queries the encoder's representations
- Reason about why decoder-only models won the architecture race, and when encoder-only models still dominate
Before You Start
#The Assembly Line
We have built up the pieces across four lessons: attention, self-attention, multi-head attention, and positional encoding. Now we assemble them into the complete architecture that changed the world.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin (2017)
The paper that introduced the Transformer. Section 3 describes the full architecture. Pay special attention to Figure 1 — one of the most reproduced diagrams in ML history.
#Step Zero: Tokenization
Before anything enters the Transformer, raw text must be converted to tokens. This is often overlooked but is a critical first step.
Type a sentence above and watch BPE (Byte-Pair Encoding) merge characters step by step. Modern LLMs use BPE variants (SentencePiece, tiktoken) to split text into subword tokens. The vocabulary typically contains 32K-128K tokens. Common words become single tokens ("the", "and"), while rare words get split into subword pieces ("un" + "familiar" + "ity").
#The Architecture: End to End
#Step 1: Input Embedding
Each input token (a word or subword) is converted to a dense vector using a learned embedding table. A vocabulary of 50,000 tokens with d_model = 512 means a 50,000 x 512 embedding matrix — about 25 million parameters just for the embedding layer. The token "cat" maps to token ID 8765, which indexes into row 8765 of the embedding matrix to retrieve a 512-dimensional vector.
#Step 2: Positional Encoding
#Step 3: Multi-Head Self-Attention
Every token attends to every other token (and itself). Each of h heads computes its own attention pattern in parallel. The outputs are concatenated and projected through W_O. Each token's representation is now enriched with context from the entire sequence. This is where "it" learns to look at "animal" and "tired" in our earlier example.
#Step 4: Add and LayerNorm (First)
#Step 5: Feed-Forward Network
Each token passes independently through a two-layer neural network: expand from d_model to 4*d_model (e.g., 512 to 2048), apply an activation function (ReLU, GELU, or SwiGLU), then contract back to d_model. This is the "thinking step", where each token processes the context it gathered from attention. Critically, the FFN applies to each token independently with shared weights.
#Step 6: Add and LayerNorm (Second)
#Step 7: Repeat N Times
Stack N identical blocks (the original paper used N = 6). Each block refines the representations further. Early layers tend to capture local syntax; middle layers capture semantics and entity relationships; later layers handle abstract reasoning and task-specific patterns. Modern LLMs use 32-96+ layers.
#Step 8: Output Projection and Softmax
The final layer's output passes through a linear projection from d_model to vocabulary size (e.g., 512 to 50,000), producing logits for each vocabulary token. Softmax converts these to probabilities. The highest-probability token is the model's prediction for the next token. For generation, this token is appended to the input and the process repeats.
#Interactive: The Full Transformer
Click through each component to see what happens at every step of the pipeline. Follow a token from raw input through embedding, attention, FFN, and out to prediction.
#Residual Connections: Information Highways
Every sub-layer (attention and FFN) has a residual connection — the input is added directly to the output:
Without residual connections, training deep Transformers (32+ layers) would be nearly impossible. Gradients would vanish as they flow backward through dozens of layers. The residual connection gives the gradient a direct path back to earlier layers — an "information highway" that keeps learning alive.
#Layer Normalization
Layer norm normalizes each token's activation vector to have zero mean and unit variance, then applies a learned scale and shift:
Why does this matter? Without normalization, activations in deep networks tend to drift — growing or shrinking as they pass through layers. This makes training unstable and slow. Layer norm acts as a stabilizer, ensuring each layer receives inputs in a predictable range. Modern alternatives include RMSNorm (used in LLaMA), which skips the mean-centering step for efficiency.
#The Feed-Forward Network
The FFN expands from 512 to 2048 dimensions and then contracts back to 512. Why expand at all if we end up at the same size?
The FFN applies the same two-layer network to every token independently:
The expansion to 4x the model dimension creates a higher-dimensional "thinking space." In this expanded space, the network can represent more complex, nonlinear functions — combining and transforming the contextual information gathered by attention. The contraction back to d_model ensures the output can be fed to the next layer.
#Encoder vs. Decoder vs. Both
#Encoder (Bidirectional)
Every token can attend to every other token — past and future. The word "it" can look both backward (at "animal") and forward (at "was tired") simultaneously. This makes the encoder ideal for tasks requiring understanding the full context.
#Decoder (Causal Masking)
#Encoder-Decoder (Full)
The encoder processes the full input bidirectionally. The decoder generates the output autoregressively while cross-attending to the encoder's representations. The decoder has both self-attention (causal) and cross-attention (to encoder).
#Putting It All Together
Here is the full picture of one Transformer layer, in order:
- Multi-Head Self-Attention: Tokens exchange information across the sequence.
- Residual Add + Layer Norm: Stabilize and preserve the original input signal.
- Feed-Forward Network: Each token processes independently in a higher-dimensional space.
- Residual Add + Layer Norm: Stabilize again before the next layer.
Stack N of these layers. Precede them with embeddings + positional encoding. Follow them with a final linear projection (to vocabulary size for language models) and softmax (to get next-token probabilities). That is the entire Transformer.
The original paper: 6 layers, 8 heads, d_model = 512, d_ff = 2048, ~65 million parameters. GPT-3: 96 layers, 96 heads, d_model = 12,288, d_ff = 49,152, 175 billion parameters. LLaMA 3 (405B): 126 layers, 128 heads, d_model = 16,384. Same architecture, scaled up thousands of times.
#Try It: Count Transformer Parameters
#The Three Pillars of Scaling
The Transformer's success comes from three properties that make it uniquely scalable:
-
Parallelism: Unlike RNNs, all tokens in a sequence are processed simultaneously. This maps perfectly to GPU architecture, which excels at massive parallel computation. Training a Transformer on 1000 tokens takes the same wall-clock time as 10 tokens (ignoring memory).
-
Uniformity: Every layer has the same architecture. Adding more layers is trivial — just stack more identical blocks. This makes the architecture easy to scale from 6 layers to 126 layers.
-
Compositionality: Each layer builds on the previous layer's output. Early layers capture simple patterns; deeper layers compose these into complex representations. This hierarchical composition is what enables reasoning, analogy, and abstraction at scale.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei (2020)
Demonstrates that Transformer language model performance follows smooth power laws in model size, dataset size, and compute. This paper showed that scaling is predictable, enabling deliberate decisions about how to allocate compute budgets.
#Key Takeaways
- Residual connections and layer normalization enable deep stacking. Residual connections let gradients flow directly through the network, while layer normalization stabilizes training, allowing transformers to scale to hundreds of layers
- The feedforward network expands then contracts. Each transformer block contains an FFN that projects to a wider dimension (typically 4x), applies a nonlinearity (GELU), then projects back, acting as a per-token processing step
- Encoder-only, decoder-only, and encoder-decoder serve different purposes. BERT (encoder-only) excels at understanding, GPT (decoder-only) excels at generation, and T5 (encoder-decoder) handles sequence-to-sequence tasks
- Scaling laws predict performance from compute budget. Larger models trained on more data with more compute systematically improve, following power laws that let researchers predict performance before training
- Tokenization is the critical first step. BPE and SentencePiece convert raw text into token sequences that the model processes; tokenizer quality directly impacts model capability, especially for code and non-English languages
#Quick Check
What is the purpose of residual connections in the Transformer?