Every layer of every neural network on Earth is one operation: a matrix multiplied by a vector. A 175-billion-parameter model is just billions of numbers organized into matrices that warp vectors as they flow through. Once you see matrices as transformations, not grids of numbers, you understand what deep learning actually does.
Learning Objectives
After this lesson, you will be able to:
See matrices as machines that stretch, rotate, flip, and reshape space — like Instagram filters for numbers
Multiply a matrix by a vector and understand what the result means as a visual transformation
Connect matrix operations to how AI models process data, and understand why the order of operations matters
Interpret the determinant as the area-scaling factor of a transformation, and explain why a zero determinant means irreversible information loss
If you understood vectors, you can understand matrices. A vector is a single arrow. A matrix is a machine that bends, stretches, and rotates ALL the arrows at once. You already have the foundation — now you are just adding one more tool to your belt.
A matrix is a grid of numbers. The numbers sit in rows (going across) and columns (going down). Here is a small one, and we will use it all lesson:
A=[2013]
A matrix with m rows and n columns has shape (m, n), rows always first. So A has shape (2, 2). A vector from the Vectors lesson is the same idea with a single column.
What does a matrix do? It takes a vector in and gives a vector out. Take the input [1, 1]. There are two ways to compute A times [1, 1], and they must agree.
View 1: row by row
Take each row, multiply it entry by entry with the vector, and add. That is the dot product from the Vectors lesson. Row 1 gives the first output number, row 2 gives the second.
Row 1 dot vector: 2 × 1 + 1 × 1 = 2 + 1 = 3
Row 2 dot vector: 0 × 1 + 3 × 1 = 0 + 3 = 3
So A times [1, 1] is [3, 3].
View 2: column by column
Recall the basis vectors from the Vectors lesson: e₁ = [1, 0] and e₂ = [0, 1], the two unit arrows along the axes. Every vector is built from them, for example [1, 1] = 1 × e₁ + 1 × e₂. So it is enough to ask where A sends each of the two.
A times [1, 0]: 2 × 1 + 1 × 0 = 2 and 0 × 1 + 3 × 0 = 0, so [2, 0]. That is column 1 of A.
A times [0, 1]: 2 × 0 + 1 × 1 = 1 and 0 × 0 + 3 × 1 = 3, so [1, 3]. That is column 2 of A.
The columns tell you where the basis vectors land. Now rebuild the answer: 1 × [2, 0] + 1 × [1, 3] = [3, 3]. Same result as the row view. Try a different input, [2, 3]. Rows: 2 × 2 + 1 × 3 = 7 and 0 × 2 + 3 × 3 = 9. Columns: 2 × [2, 0] + 3 × [1, 3] = [4, 0] + [3, 9] = [7, 9]. Both views give [7, 9].
What this does to a whole space
A matrix does a linear transformation: it moves every vector to a new place, but the origin stays put and grid lines stay straight and evenly spaced. Nothing bends, and nothing slides off the origin.
Watch the unit square, the square with corners (0, 0), (1, 0), (1, 1) and (0, 1). Send each corner through A:
(0, 0) goes to (0, 0)
(1, 0) goes to (2, 0), which is column 1
(1, 1) goes to (3, 3)
(0, 1) goes to (1, 3), which is column 2
The square becomes a slanted parallelogram. Its base is 2 wide and its height is 3, so its area is 2 × 3 = 6. The original square had area 1, so A scaled the area by a factor of 6. That factor has a name: the determinant, the number by which a matrix scales areas. For a 2×2 matrix with rows [a, b] and [c, d], it equals a × d − b × c. For A that is 2 × 3 − 1 × 0 = 6, matching the picture. A later section explores it in depth.
The picture below draws exactly this. Type in the numbers and watch the grid move.
Try it: Set the matrix entries and watch space warpInteractive
Loading visualization...
Set the top-left entry to 2 and watch the x-axis stretch. Set the bottom-right entry to 0 and watch the y-axis collapse. Then enter A = [[2,1],[0,3]] and check that the unit square's corners land on the numbers above. Each number in the matrix controls a specific aspect of the transformation.
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type these lines yourself.
Written as one formula, the row-by-row computation you just did is:
[2013][11]=[2⋅1+1⋅10⋅1+3⋅1]=[33]
The symbol · is just multiplication. Each output entry is built by taking one row, pairing it with the vector entry by entry, and adding. The input [1, 1] landed on [3, 3]. Geometrically, the picture above shows the matrix stretching space unevenly, more along the y-axis than the x-axis, and adding a shearing tilt. To truly understand a matrix, look at what it does to the entire space, not just one vector.
You already found this in View 2 above. It is the key insight that makes matrices click: the columns of a matrix tell you where the basis vectors end up.
A=[2013]⇒e^1→[20],e^2→[13]
What Do You Think?
If a matrix has columns [1, 0] and [0, 1] (the identity matrix), what happens to any vector you multiply it by?
The identity matrix sends each basis vector to itself. Since every vector is built from basis vectors, every vector stays put. The identity matrix is the "do nothing" transformation — the mathematical equivalent of a pass-through layer.
Quick check
A 2×2 matrix has columns [3, 0] and [0, 3]. What does this transformation do?
A rotation matrix spins every vector by an angle θ (theta, the Greek letter we use for an angle) without changing its length. Two functions describe the turn. Stand at the point (1, 0), the tip of e₁, and walk along a circle of radius 1 around the origin, counterclockwise, until you have turned by θ. Where you end up is the point (cos θ, sin θ). So cos θ is how far right you are and sin θ is how far up.
Take θ = 90 degrees, a quarter turn. You end up straight above the origin at (0, 1), so cos 90° = 0 and sin 90° = 1. The tip of e₁ moves to [0, 1]. The tip of e₂ = [0, 1] turns a quarter turn too and lands at [-1, 0].
Those two landing spots are the columns, so the 90-degree rotation matrix is [[0, -1], [1, 0]]. Check it on [3, 4]: 0 × 3 + (-1) × 4 = -4 and 1 × 3 + 0 × 4 = 3, giving [-4, 3]. Both [3, 4] and [-4, 3] have length 5, so nothing stretched. For any angle, the same recipe gives:
A diagonal matrix has numbers only on its top-left to bottom-right diagonal and zeros elsewhere. It stretches each axis independently. Here s₁ and s₂ are just two numbers you choose, one per axis:
What if a column is the zero vector? Then that dimension gets collapsed. A projection matrix crushes data onto a smaller space inside the bigger one, such as a flat line inside a plane:
When you multiply two matrices, you compose two transformations. First apply B, then apply A. The result AB is a single matrix that does both at once.
First, a shape rule so you know when multiplying is allowed. A matrix of shape (m, n) times a matrix of shape (n, p) gives a matrix of shape (m, p). The two inner numbers must match, and they vanish. A (2, 2) matrix times a vector, which counts as shape (2, 1), gives shape (2, 1).
(AB)v=A(Bv)
Here it is with numbers. Let A be the 90-degree rotation [[0, -1], [1, 0]] and B stretch x by 2, [[2, 0], [0, 1]]. Send v = [1, 1] through. Read right to left, so B touches v first:
So A(Bv) = [-1, 2]. Now build the single matrix AB. Its columns are A applied to each column of B. A times [2, 0] = [0, 2] and A times [0, 1] = [-1, 0], so AB = [[0, -1], [2, 0]]. Apply it in one shot: 0 × 1 + (-1) × 1 = -1 and 2 × 1 + 0 × 1 = 2, giving [-1, 2]. The same answer, as (AB)v = A(Bv) promises.
Angles compose the same way. Turning by 45° twice is a 90° turn, so R(45°) times R(45°) equals R(90°). Multiplying the matrices out confirms it: the product comes out as [[0, -1], [1, 0]].
Set A and B in the picture below, then flip their order and watch the final shape end up somewhere else.
Composition: apply B then A, then flip the orderInteractive
Loading visualization...
Quick check
Let R rotate by 90 degrees and S scale x by 2. Compute the product RS — what does it do to a vector, in what order?
You met the determinant when the unit square became a parallelogram of area 6: it is the factor by which a matrix scales areas. If det(A) = 2, areas double. If det(A) = 0.5, areas halve. If det(A) = 0, the transformation collapses space into a lower dimension and information is irreversibly lost. For a 2×2 matrix the rule is simple:
det[acbd]=ad−bc
Multiply the main diagonal (a and d), then subtract the product of the other two (b and c). Worked examples:
A = [[2, 1], [0, 3]]: 2 × 3 − 1 × 0 = 6, areas scale by 6
Scaling every direction by 3, [[3, 0], [0, 3]]: 3 × 3 − 0 × 0 = 9, since the square is 3 times wider and 3 times taller
The 90-degree rotation [[0, -1], [1, 0]]: 0 × 0 − (-1) × 1 = 1, so a rotation preserves area
The shear [[1, 0.5], [0, 1]]: 1 × 1 − 0.5 × 0 = 1, so a shear preserves area too, even though it tilts shapes
Now try it. Change the entries and watch the blue square's area match ad − bc.
Determinant — Watch the Unit Square Scale (or Collapse)Interactive
What a matrix does to area, and what a determinant of zero means.
If a matrix A transforms space, the inverse A^(-1) undoes that transformation:
A−1A=I
For our A = [[2, 1], [0, 3]] the inverse is [[0.5, -0.1667], [0, 0.3333]] (rounded). Apply it to the output [3, 3] and you get back [1, 1], the vector we started from. A matrix has an inverse only when its determinant is nonzero — when the transformation does not collapse any dimension. If information is lost (determinant = 0), there is no way to recover it. This is why singular matrices are problematic in ML: they indicate redundancy or degeneracy in the data or model.
Let's see all of this come alive with real numpy. The playground below applies a 2×2 matrix to a square of points and prints the transformed coordinates. Change the matrix entries and watch how the shape morphs.
You have computed A times a vector two ways, built a determinant by hand, and multiplied two matrices in both orders. Now make Python do each step, so you can see the views agree. You will use A = [[2, 1], [0, 3]], the vector [2, 3], and the rotation and stretch from the composition section.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
Tests · Verify mat_vec_rows and mat_vec_cols both return [7, 9] and equal A @ v, det2 gives 6, 1 and 0 for A, the shear and the projection and matches np.linalg.det, rot @ stretch is [[0, -1], [2, 0]] sending [1, 1] to [-1, 2], stretch @ rot is [[0, -2], [1, 0]] sending [1, 1] to [-2, 1], and the reflection sends [3, 4] to [3, -4] with determinant -1.
Running the solution prints these values. The row view gives [7 9] and the column view gives [7 9], and both equal A @ v. The determinants are 6.0 for A, 1.0 for the shear and 0.0 for the projection, each matching numpy. The product rot @ stretch is [[0, -1], [2, 0]] and sends [1, 1] to [-1 2], while stretch @ rot is [[0, -2], [1, 0]] and sends it to [-2 1], so order matters. The reflection sends [3, 4] to [3 -4] and has determinant -1.
The sign is the lesson. A determinant of -1 says the reflection kept areas the same (the size is 1) but flipped orientation (the minus sign). A determinant of 0, as for the projection, is the warning that no inverse exists.
Visualizing and Understanding Neural Networks
Jason Yosinski, Jeff Clune, Yoshua Bengio, Hod Lipson (2014)
Pioneering work on visualizing what neural network layers actually learn. Shows how each layer progressively transforms representations from raw pixels to abstract concepts.
⚡ Playground:Vectors & Matrices → — multiply matrices and watch vectors rotate, scale, and shear in real time.
Matrices are space-warping machines. A matrix transforms every vector in a space simultaneously through rotation, scaling, shearing, or projection, and every neural network layer is fundamentally a matrix multiplication
Columns reveal the transformation. The columns of a matrix tell you exactly where the basis vectors land after the transformation, which completely determines what happens to every other vector
Determinant measures area scaling. A determinant of zero means the matrix collapses space into a lower dimension, permanently destroying information; a negative determinant means the transformation flips orientation
Order of multiplication matters. Matrix multiplication is not commutative (AB does not equal BA), which is why the order of layers in a neural network profoundly affects its behavior
A chain of matrix multiplications is itself a single matrix, so stacking them alone adds no new power. A later lesson on neural networks shows what is added between layers to fix that
What does it mean if a matrix has a determinant of zero?
Next up: Eigenvalues & SVD. You will find the special directions that a matrix stretches without turning, and learn to break any transformation into simple pieces, the key to PCA and recommendation systems.
Your Reflection
Saves automatically
What’s one thing you learned? What’s still confusing?