A pin on a map is two numbers: how far east and how far north. A song can be described by three numbers: how energetic, how danceable, how fast. A list of numbers that describes one thing is called a vector, and machine learning stores almost everything this way. In this lesson you will add vectors, measure them, and compare them with one small multiply-and-add recipe. If you can add and multiply, you can do all of it.
Before you start: the Math Toolkit lesson covers squares, square roots and the sigma sum (adding up a list of terms). That is all the background you need. Every code cell here is plain Python with lists and loops, with no extra libraries.
Learning Objectives
After this lesson, you will be able to:
Read a vector two ways: as a list of numbers and as an arrow with a direction and a length
Compute a dot product by hand and explain what its sign tells you
Turn the dot product into cosine similarity, and explain why it compares direction while ignoring length
Measure length and distance with the L1, L2 and L-infinity norms
Explain basis vectors and span, and why data often uses fewer directions than it has numbers
Math is not a talent — it is a skill. If you can count, add, and follow directions on a map, you already have everything you need to understand vectors. This lesson starts from zero and builds up one piece at a time. Take it slow and you will surprise yourself.
This lesson begins with numbers you can draw and adds one idea at a time. By the end you will see how data in machine learning is stored, and why.
In mathematics, we write a 2D vector as a column of numbers:
v=[34]
Try it! Open the Python REPL (bottom-right of the screen: click Quick Actions, then Python) and type v = [3, 4] yourself.
The top number (3) says "move 3 units along the x-axis." The bottom number (4) says "move 4 units along the y-axis." Together, they define a direction and a distance.
The magnitude (or length, or norm) of this vector is the total straight-line distance traveled:
∥v∥=32+42=9+16=25=5
The claim is that this is the Pythagorean theorem in disguise. Here is the check, on a grid. Walk 3 east, then 4 north. The straight line from where you started to where you ended is the vector itself, and it closes a right triangle (a sketch, not to scale):
y
4 | B (3, 4)
| / |
3 | / |
| / | 4 north
2 | / |
| / |
1 | / |
|/ |
0 O-----------------+---- x
0 1 2 3
3 east
Now build a square on each side of that triangle. The square on the 3-side holds 3 × 3 = 9 unit squares. The square on the 4-side holds 4 × 4 = 16. The Pythagorean theorem says the square on the slanted side holds exactly 9 + 16 = 25 unit squares, so each of its sides is √25 = 5. That is the magnitude formula: square the components, add them, take the square root.
Nothing changes when we go from 2 components to 3, or to 100. For [3, 4, 12] you walk 3 along the first axis, 4 along the second and 12 along the third. The length just gets one more term under the root: 9 + 16 + 144 = 169, and √169 = 13.
Models describe words and images with lists of numbers like this too. A list a model has learned to describe a word or an image is called an embedding vector, and similar things end up with similar lists. Modern language models use embedding vectors with hundreds to thousands of numbers. You cannot draw 768 dimensions, and you do not need to. The recipe is the same: one more component, one more term under the root.
Multiplying a vector by a number (a scalar) stretches or shrinks it without changing the line it points along. Multiply by -1 and it flips to point the opposite way.
2⋅[34]=[68]
What Do You Think?
If you multiply the vector [3, 4] by -0.5, what happens geometrically?
Multiplying by -0.5 does two things: the 0.5 halves the length, and the negative sign flips the direction. The result is [-1.5, -2] — a vector half as long, pointing in the opposite direction.
To make the connection between vectors and operations concrete, here is an interactive sandbox. Drag the head of vector a or b and watch the addition and scalar multiplication update in real time.
Loading visualization...
Quick check
A vector v has magnitude 10. You multiply it by the scalar -2. What are the magnitude and direction of the result?
How alike are two lists of numbers? One small recipe answers that, and you can do it by hand. Take a = [1, 2] and b = [3, 4]. Multiply the first components, multiply the second components, then add the results:
1 × 3 + 2 × 4 = 3 + 8 = 11
That 11 is the dot product of a and b. Try two more.
a = [1, 0] and b = [0, 1]: 1 × 0 + 0 × 1 = 0 + 0 = 0. The first arrow runs along the x-axis and the second straight up the y-axis, so they are perpendicular (at a right angle). The dot product is 0.
a = [3, 4] and b = [1, -2]: 3 × 1 + 4 × (-2) = 3 - 8 = -5. The result is negative.
Three results, three stories. A positive result (11) means the arrows lean broadly the same way. Zero means they are perpendicular. A negative result (-5) means they lean away from each other.
a⋅b=a1b1+a2b2+⋯+anbn
Read the formula as the recipe you just used: a₁ is the first component of a, b₁ the first component of b, and so on up to the n-th. The three dots mean "keep going". The dot product is the single most important operation in all of ML, because "how much do these two lists agree?" is the question it answers.
The dot product also hides an angle. First, the one piece you need: what the cosine measures.
Now the geometric identity:
a⋅b=∥a∥∥b∥cosθ
Check it against the first example. For a = [1, 2] and b = [3, 4] the lengths are √5 ≈ 2.236 and 5, and the angle between them is about 10.3 degrees, where cos is about 0.984. Then 2.236 × 5 × 0.984 = 11, exactly the dot product you computed by hand. The perpendicular example agrees too: [1, 0] and [0, 1] both have length 1 and a 90 degree angle, so 1 × 1 × 0 = 0.
Here is the same idea as a picture: drag the tips of the two arrows, and watch the amber shadow of b on a (the dot product is the length of a times that shadow) shrink to nothing at a right angle and turn negative once b leans away from a.
Now that the dot product means something, one more word. To normalise a vector, divide every component by the vector's own length. The result points the same way but has length exactly 1.
Take a = [3, 4], whose length is 5. Normalised, it is [3/5, 4/5] = [0.6, 0.8]. Check its length: 0.36 + 0.64 = 1. Take b = [4, 3], also of length 5, which normalises to [0.8, 0.6].
The raw dot product is 3 × 4 + 4 × 3 = 24.
The dot product of the two normalised vectors is 0.6 × 0.8 + 0.8 × 0.6 = 0.48 + 0.48 = 0.96.
The formula above gives cos θ = 24 / (5 × 5) = 0.96.
Same number. Once both lengths are 1, the product of lengths in |a||b| cos θ is just 1, so the dot product of normalised vectors IS the cosine of the angle (here about 16.3 degrees). That number is called cosine similarity, and it is used in everything from search engines to word embeddings:
cosine_similarity(a,b)=∥a∥∥b∥a⋅b
In a transformer, the attention step starts by taking a dot product between a query vector and each key vector. The Transformers track explains why. The arithmetic is exactly the multiply-and-add you just did.
Let us try it with numbers. Here are made-up 5-number vectors for "king", "queen" and "car". This is a toy: the numbers are not from a real model, they only show the mechanics.
king = [0.8, 0.2, -0.1, 0.9, 0.3]
queen = [0.7, 0.3, -0.2, 0.85, 0.35]
car = [-0.1, 0.8, 0.7, -0.3, 0.1]
The dot product of king and queen, one component at a time: 0.8 × 0.7 = 0.56, 0.2 × 0.3 = 0.06, (-0.1) × (-0.2) = 0.02, 0.9 × 0.85 = 0.765, 0.3 × 0.35 = 0.105. Add them: 0.56 + 0.06 + 0.02 + 0.765 + 0.105 = 1.51. The lengths come out as about 1.261 for king and 1.210 for queen, so the cosine similarity is 1.51 / (1.261 × 1.210) ≈ 0.989. Doing the same for king and car gives a dot product of -0.23 and a cosine of about -0.164. In this toy, king and queen point almost the same way, and car points slightly away.
Now compute it for real. The playground below runs plain Python in your browser, with no setup and no install. Edit the vectors, click Run, and watch the dot product, lengths and cosine similarity update.
The length we have been using, 5 for [3, 4], is one way to say how big a vector is. A norm is any rule that turns a vector into one non-negative number, its size. Three are common, and you can work all of them by hand for [3, 4]:
L1 norm: add the absolute values of the components (the absolute value drops a minus sign). 3 + 4 = 7. Think of city blocks, where you can only walk along streets.
L2 norm: the straight-line length we already used. √(3² + 4²) = √25 = 5.
L-infinity norm: the largest absolute component. The larger of 3 and 4 is 4.
So the same arrow has size 7, 5 or 4, depending on the question you are asking. L2 is the default when people say "length".
A norm also gives you a distance. The distance between two points is the size of the vector that gets you from one to the other. From a = [1, 2] to b = [4, 6], that vector is b - a = [3, 4]. Its L2 size is 5, so the straight-line distance is 5. Its L1 size is 7 (walking the blocks) and its L-infinity size is 4.
Drag the point in the next picture to see all three sizes of the same arrow with the arithmetic written out, and watch the shape formed by every point whose size is exactly 1: a diamond, a circle or a square.
Loading visualization...
A later lesson meets L1 and L2 again as penalties that keep a model's numbers small, and L2 distance is one way nearest-neighbour search measures "closest".
Everything starts as raw data: image pixels, text, audio samples, or spreadsheet rows. At this stage, the data is not yet in a form that any ML model can work with. A 28x28 grayscale image is just a grid of brightness values. A sentence is just a sequence of characters.
The raw data is converted into vectors, lists of numbers. A 28x28 image becomes a vector with 28 × 28 = 784 components. A word becomes an embedding vector of hundreds of numbers. This is the critical step: vectorization maps messy real-world data into the precise mathematical world where ML algorithms operate.
Once data is vectorized, ML algorithms use vector operations to extract meaning. The dot product measures how similar two vectors are. Cosine similarity removes the effect of length. A matrix is a grid of numbers that turns one vector into another, and the next lesson, Matrices as Transformations, covers it properly. The picture on the right is a preview: it shows how one matrix warps the whole plane.
A vector space is the arena where vectors live: a collection of vectors where adding two of them, or stretching one by a number, always lands you back inside the collection. The two operations must also follow the ordinary rules of arithmetic.
The most common vector space in ML is R^n (read "R to the n"). It is the name for all lists of n real numbers, where an n-tuple just means a list with exactly n entries. R^2 is the flat plane you have been drawing on, and R^3 is ordinary 3D space. When we say "a 768-dimensional vector", we mean one point in R^768.
Any vector in a space can be built by combining a set of basis vectors. In 2D, the standard basis is:
e^1=[10],e^2=[01]
The vector [3, 4] is 3 times e1 plus 4 times e2: [3, 0] + [0, 4] = [3, 4]. Stretching each vector by a number and adding the results is called a linear combination. The span of a set of vectors is every vector you can create through linear combinations. If two vectors are not parallel, they span the entire 2D plane. If they are parallel, they only span a line.
The standard basis is not the only one. With [1, 1] and [1, -1] you can also reach [3, 4]: 3.5 × [1, 1] + (-0.5) × [1, -1] = [3.5 - 0.5, 3.5 + 0.5] = [3, 4]. Different building blocks, same destination, different stretch amounts.
Slide the two stretch amounts in the next picture to walk to different points, then make the second arrow parallel to the first and watch the reachable plane collapse onto a line.
Loading visualization...
Quick check
You have three vectors in 3D: v1 = [1, 0, 0], v2 = [0, 1, 0], v3 = [2, 3, 0]. What is the dimension of their span?
Time to build the lesson's tools in code. You will write the dot product, the length, cosine similarity and normalising from scratch in plain Python, then check them against the numbers worked by hand above. The toy king, queen and car vectors are provided.
pythonplayground.py · Pyodide
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
Tests · Verify dot(king, queen) is 1.51 and dot(king, car) is -0.23, cosine(king, queen) is 0.9894 and cosine(king, car) is -0.1638, the dot product of the two normalised vectors equals the cosine (0.9894), length(normalise(king)) is 1.0, tripling king triples the dot product (4.77 against 1.59) while the cosine stays at 1.0.
When you run the solution you should see dot(king, queen) = 1.51 and dot(king, car) = -0.23, and cosines of 0.9894 and -0.1638. The normalised dot product matches the cosine at 0.9894, and the normalised king has length 1.0. In the stretch step, tripling king triples the dot product (4.77 against 1.59 for king with itself) but the cosine stays at 1.0. The dot product feels the length, and the cosine does not.
Real embedding vectors have hundreds to thousands of numbers instead of 5. The code is identical, just with longer lists.
A vector is a list of numbers, and you can also picture it as an arrow with direction and magnitude. Every data point in ML is represented as a vector
Dot product measures alignment. Multiply matching components and add: it is positive when vectors lean the same way, zero when perpendicular, negative when they lean apart
Cosine similarity normalizes for magnitude. By dividing the dot product by both vector lengths, it measures pure directional alignment regardless of scale, which is why it powers embedding-based search
Norms measure size. For [3, 4] the L1, L2 and L-infinity norms are 7, 5 and 4, and the norm of a difference is a distance
Basis vectors and span define the space. Any vector can be built from linear combinations of basis vectors, and data often uses fewer independent directions than it has numbers
High dimensions work the same way. The math for 2D vectors extends identically to 768D or higher; you cannot visualize it, but the operations are unchanged
u = [2, -1] and v = [1, k]. For which value of k are u and v perpendicular?
Key Terms12 terms
An ordered list of numbers that describes one thing, such as a point or a displacement; you can also picture it as an arrow with a direction and a magnitude.
One entry of a vector. [3, 4] has two components, 3 and 4.
The number of components in a vector. [3, 4] is 2-dimensional; a 28x28 image flattened into a list is 784-dimensional.
A single real number with no direction. Multiplying a vector by a scalar stretches or shrinks it without changing the line it points along (a negative scalar also flips it).
The size of a vector. The usual L2 norm is the square root of the sum of its squared components, written ||v||. L1 adds absolute values and L-infinity takes the largest absolute component.
A single number computed by multiplying corresponding components of two vectors and summing them. Measures how aligned two vectors are.
The dot product divided by both vector magnitudes, giving a value in [-1, 1] that measures only angular alignment, ignoring length.
Divide every component of a vector by its own length, so the result has length 1 and points the same way.
A list of numbers a model has learned to describe a word, image or other item, so that similar items get similar lists.
Two vectors are orthogonal (perpendicular) when their dot product is zero. In ML, orthogonal (mean-centred) features are uncorrelated: neither linearly predicts the other. Uncorrelated is weaker than independent.
A smallest set of vectors that can build every vector in a space by linear combinations (none of them can be built from the others).
The set of all vectors you can produce by linear combinations of a given set of vectors. In 2D, two non-parallel vectors span the entire plane.
Where This Matters
Search and retrieval systems
Semantic Search
An embedding model turns documents and queries into vectors with hundreds to thousands of numbers, so a retrieval system can rank results by cosine similarity instead of keyword overlap. (Semantic means about meaning.)
↑
Finds results by meaning, not only by shared words
Streaming services
Music Recommendations
Recommendation systems commonly describe tracks and listeners with vectors, so 'songs like this' becomes 'find the vectors closest to this one'.
↑
Turns taste into a nearest-neighbour search
Image search products
Visual Search
A vision model encodes each image as a vector, and 'visually similar' lookups are nearest-neighbour searches in that vector space.
↑
Lets you search with a picture instead of a keyword
Interview Practice
Next up: Matrices as Transformations. You will see how a grid of numbers turns one vector into another, stretching, rotating or flipping the whole plane, and why every neural-network layer does exactly this.
The model processes these vectors through layers of transformations. In a transformer, query and key vectors are compared via dot products to compute attention. In a recommendation system, user and item vectors are compared via cosine similarity. In RAG (retrieval-augmented generation, where a system looks up relevant documents before answering), document and query vectors are compared to find relevant context. Every one of these steps is built from the vector math in this lesson.