Math Toolkit: Functions, Exponents, Logs & Sums
After this lesson, you will be able to:
- Read f(x) notation, draw a function as a set of (x, f(x)) pairs, and evaluate a composition f(g(x)) from the inside out
- Evaluate powers with positive, negative and fractional exponents, and say what the number e is and why growth and decay use it
- Use a logarithm as the undo button for an exponent, and apply the three log rules: product to sum, power to multiplier, change of base
- Explain why ML libraries work with log-probabilities instead of raw products, and show it with a number you compute yourself
- Expand sigma and pi expressions term by term, write them as a Python sum() or loop, and read a real formula such as softmax or the Gaussian density word by word
- Describe a slope as rise over run, which is the doorway to the next lessons
- Check your own algebra with six quick questions, and read the set symbols (in, union, intersection, for all, there exists, such that) that the probability lessons use
#Why a Maths Toolkit First
Open any ML paper or tutorial and you will see things like exp, log, a tall sigma, a product symbol, and little letters attached above and below. Most courses use these symbols on page one and never explain them.
This lesson explains them on purpose. It is not about getting good at algebra. It is about being able to look at a formula and know what to do with each piece.
Before the main ideas there are two short warm-ups. One is a six-question algebra self-check, so you can find out whether you are ready. The other is a table of set symbols, because the probability lessons use them.
Then you will meet seven small ideas, in an order where each one uses the one before:
- Functions, the machines that turn inputs into outputs.
- Exponents and the special number e.
- Logarithms, the undo button for exponents.
- Why ML lives on logs.
- Sigma and pi, the notation for sums and products.
- A symbol table and two formulas read word by word.
- Sigmoid and softmax, and then a first look at slopes.
Every Python cell on this page runs in your browser. Run them. Reading about a number and seeing your own machine print it are different experiences.
#Algebra Self-Check
Try these six questions on paper before you read the answers. They use only school arithmetic, and they are the skills the rest of this lesson leans on.
- Solve 4x = 20. What is x?
- What is 1/2 + 1/3?
- What are 3 squared and 10 cubed?
- What are -5 + 8 and (-3) x (-2)?
- What is 2 + 3 x 4? What is (2 + 3) x 4?
- Solve 3x - 4 = 11. What is x?
| Question | Answer | Working |
|---|---|---|
| 1 | x = 5 | Divide both sides by 4: 20 / 4 = 5. |
| 2 | 5/6 | Common bottom of 6: 3/6 + 2/6 = 5/6. |
| 3 | 9 and 1000 | 3 x 3 = 9, and 10 x 10 x 10 = 1000. |
| 4 | 3 and 6 | Start at -5 and move 8 steps up to land on 3. Two negatives multiplied give a positive: 3 x 2 = 6. |
| 5 | 14 and 20 | Multiplication goes before addition: 3 x 4 = 12, then 2 + 12 = 14. Brackets go first: 5 x 4 = 20. |
| 6 | x = 5 | Add 4 to both sides to get 3x = 15, then divide by 3. |
If you got at least five of the six, you are ready. If you missed several, do not stop. Read this lesson slowly, run every Python cell, and redo any question you missed when its idea appears below: powers return in the exponents section, and negatives, brackets and fractions return all through the lesson.
#Sets and Symbols
{1, 2, 3} and B = {3, 4}.| Symbol | Read as | Example with the answer |
|---|---|---|
| ∈ | "is in" (and ∉ means "is not in") | 2 ∈ A is true. 5 ∈ A is false, so 5 ∉ A. |
| ∪ | union: everything in either set | A ∪ B = {1, 2, 3, 4} |
| ∩ | intersection: only what is in both | A ∩ B = {3} |
| ∀ | "for all" | "For all x in A, x > 0" is true, because 1, 2 and 3 are all above 0. |
| ∃ | "there exists" | "There exists x in A with x > 2" is true, because x = 3 works. |
| such that | written as a colon or a bar | {x ∈ A : x > 1} reads "the x in A such that x is above 1", which is {2, 3}. |
When a later lesson says an event is a set of outcomes, "A or B" will be the union and "A and B" will be the intersection. Python has the same ideas built in, so you can check each row.
#Functions: Machines That Turn Inputs into Outputs
Take the rule "double it and add one". In symbols:
Feed in a few inputs and write down what comes out:
| Input x | Working | Output f(x) |
|---|---|---|
| -2 | 2 x (-2) + 1 | -3 |
| 0 | 2 x 0 + 1 | 1 |
| 3 | 2 x 3 + 1 | 7 |
| 10 | 2 x 10 + 1 | 21 |
#A graph is every (x, f(x)) pair
A neural network is also a function. It takes an input (the pixels of a photo) and returns an output (the probability of "cat"). The only difference is that its rule has millions of numbers inside it, and training is the process of choosing those numbers.
#Composition: one machine feeding another
Work it for x = 3, from the inside out:
- g(3) = 3 squared = 9.
- f(9) = 2 x 9 + 1 = 19.
So f(g(3)) = 19. Order matters. Reverse it: f(3) = 7, then g(7) = 49. So g(f(3)) = 49, a different answer.
A neural network is a long composition. Layer 1 feeds layer 2, which feeds layer 3, so the whole model is f3(f2(f1(x))). Later, when you learn how a small change in the input travels through a stack of layers, you will use exactly this idea of nesting. That rule is called the chain rule, and composition is its raw material.
#Exponents and the Number e
Doubling is the easiest way to feel how fast this grows. After 10 doublings you have 2 to the 10 = 1024. After 20 doublings you have 1,048,576. A power is shorthand for growth that compounds.
#Zero and negative exponents
Two extensions look odd at first, but each follows from one idea: an exponent counts how many times you multiply, and subtracting one from the exponent means dividing by the base once.
| Expression | Meaning | Value |
|---|---|---|
| 2 to the 3 | 2 x 2 x 2 | 8 |
| 2 to the 2 | 2 x 2 | 4 |
| 2 to the 1 | 2 | 2 |
| 2 to the 0 | divide 2 by 2 once more | 1 |
| 2 to the -1 | divide again, giving 1 over 2 | 0.5 |
| 2 to the -3 | 1 over (2 x 2 x 2) | 0.125 |
Read the pattern going down: every step down in the exponent halves the value. A negative exponent means "one over". Any number to the power 0 is 1.
A note on negatives, so that 2 to the -3 is not confusing. The minus sign in an exponent is not a negative number: 2 to the -3 is 0.125, which is positive. Separately, multiplying two negative numbers gives a positive one. You can see why from a pattern: 3 x (-3) = -9, 2 x (-3) = -6, 1 x (-3) = -3, 0 x (-3) = 0. Each step down in the first number adds 3, so the next row must be (-1) x (-3) = 3. That is why (-2) x (-3) = 6.
#Two rules for powers, and fractional exponents
Two handy rules follow from "an exponent counts copies". Multiplying powers of the same base adds the exponents. Count the copies: 2 to the 3 is three 2s, 2 to the 2 is two 2s, and together there are five 2s. In numbers, 8 x 4 = 32, which is 2 to the 5.
Raising a power to a power multiplies the exponents. Take (2 to the 3) to the 2. That means two copies of 2 to the 3 multiplied together, which is (2 x 2 x 2) x (2 x 2 x 2), six 2s in all. In numbers, 8 x 8 = 64, which is 2 to the 6, and 3 x 2 = 6.
| Expression | Meaning | Value |
|---|---|---|
| 9 to the 1/2 | the number that squares to 9 | 3 |
| 8 to the 1/3 | the number that cubes to 8 | 2 |
#The number e
Suppose a bank pays 100% interest per year on 1 unit of money. Paid once at year end, you hold 2. Paid twice, as 50% every half-year, you hold (1 + 1/2) squared = 2.25. Split into more and more slices, the total creeps up but never explodes:
| Slices n | (1 + 1/n) to the n |
|---|---|
| 1 | 2.0000 |
| 10 | 2.5937 |
| 100 | 2.7048 |
| 1,000 | 2.7169 |
| 1,000,000 | 2.7183 |
For exp, the steepness at any spot equals the height of the curve there. At x = 2 the height is 7.389 and the steepness is also about 7.39 (between x = 2 and x = 2.001 the rise over run is 7.3928). At x = 3 both are about 20.09. No other base does this, and that is why e is the natural base. The last section of this lesson takes slopes further, and the derivatives lesson explains why exp behaves this way.
Why ML cares: exp turns any real number, positive or negative, into a positive one. Probabilities must be positive, so exp is how a model turns raw scores into valid ones. You will see that in the softmax section below.
#Logarithms: The Undo Button for Exponents
Here is the picture. Start at 1 and keep doubling. Each doubling is one step up a staircase, and the log base 2 of a number is which step you are standing on.
| Step | Value after that many doublings | log2 of the value |
|---|---|---|
| 0 | 1 | 0 |
| 1 | 2 | 1 |
| 2 | 4 | 2 |
| 3 | 8 | 3 |
| 4 | 16 | 4 |
| 5 | 32 | 5 |
Reading the table sideways: to get from the value 32 back to the step number, take log2, and you get 5. The exponent walks up the stairs and the log walks back down.
| Name | Written | Meaning | Example |
|---|---|---|---|
| Log base 2 | log2 | how many doublings | log2(8) = 3 |
| Common log | log10 | how many factors of 10, roughly "number of digits minus 1" | log10(1000) = 3 |
| Natural log | ln | log with base e | ln(e squared) = 2 |
log almost always means the natural log ln. np.log and math.log both do. When a paper says log without a base, assume ln unless told otherwise.Because a log is the undo of an exponent, they cancel in either order: ln(exp(x)) = x and exp(ln(x)) = x. Feed a number in and then undo it, and you are back where you started. Also note that logs only exist for positive numbers. There is no power of 2 that equals 0 or a negative number.
The log of a number below 1 is negative. Check it with base 2: log2(0.5) = -1, because 0.5 is 1 over 2, and "one over" is what a negative exponent means, so 2 to the -1 = 0.5. The natural log works the same way: ln(0.1) = -2.3026, because 0.1 is 1 over 10, so the power of e that gives 0.1 must be negative (e to the -2.3026 is 0.1). Remember this, because every probability sits between 0 and 1, so the log of a probability is always negative or zero.
#The three rules
These three rules carry almost everything. Each one is just an exponent rule turned inside out.
| Rule | Statement | Worked number |
|---|---|---|
| Product becomes a sum | log(a x b) = log(a) + log(b) | log2(2 x 8) = log2(16) = 4, and log2(2) + log2(8) = 1 + 3 = 4 |
| Power becomes a multiplier | log(a to the b) = b x log(a) | log2(8 squared) = log2(64) = 6, and 2 x log2(8) = 2 x 3 = 6 |
| Change of base | log_b(x) = ln(x) / ln(b) | log2(10) = ln(10) / ln(2) = 2.3026 / 0.6931 = 3.3219 |
The first rule is the one to remember. It comes straight from the exponent rule you already saw: multiplying powers adds their exponents. A log reads off exponents, so multiplying the numbers adds their logs.
The third rule has a reason too. Let x be log2(10), which means 2 to the x is 10. Take ln of both sides. The power rule moves the exponent to the front, so x times ln(2) equals ln(10). Divide by ln(2) and you get x = ln(10) / ln(2) = 2.3026 / 0.6931 = 3.3219. As a check, 2 to the 3.3219 is 9.9998, which is 10 up to rounding.
What is log2(32)?
#Why Machine Learning Lives on Logs
Here is the practical reason logs matter so much, as a short story.
Three heads in a row: 0.5 x 0.5 x 0.5 = 0.125. Ten heads in a row: 0.5 multiplied ten times is 0.000977. A hundred heads in a row is 7.9e-31, where "e-31" means the decimal point moves 31 places to the left. Four hundred heads in a row is 3.9e-121. Every extra event multiplies in another number below 1, so the total keeps shrinking toward zero.
Make it concrete with a smaller probability. Say 400 independent events each have probability 0.1. The likelihood of all 400 is 0.1 multiplied by itself 400 times, which mathematically is 10 to the power -400. Computers write that as 1e-400, meaning 1 divided by a 1 followed by 400 zeros.
In Python, you start with total = 1.0 and multiply it by 0.1 four hundred times in a loop. What does the program print for total at the end?
A 64-bit float holds numbers down to about 1e-308 at full precision and, with reduced precision, down to about 5e-324. Below that it silently rounds to zero. 0.1 multiplied by itself 100 times is 1e-100, which still fits. At 300 copies you reach 1e-300, still fine. At 320 copies the value is about 1e-320 and has lost most of its precision. At 400 copies it is gone, and the program says 0.0.
Real models are far worse than this. A language model scoring a long sentence, or a classifier scoring thousands of training rows, multiplies far more than 400 factors.
Even with probability 0.5, which is far larger than 0.1, the plain product underflows to 0.0 after 1075 events. Python gives no warning at all.
Now use the product rule. The log of a product is the sum of the logs, and adding numbers does not shrink anything:
The answer, -921.03, loses nothing. You can still compare two models by it: the one with the larger (less negative) total log-likelihood explained the data better. Since log is increasing, "larger likelihood" and "larger log-likelihood" always pick the same winner.
Run it yourself. The cell prints the naive product, the sum of logs, and where each one breaks.
Next, try the explorer below. It draws exp and log as mirror images and shows the product-to-sum rule on numbers you choose. Its third panel repeats the underflow demo from the cell above as a second look, so spend your time on the first two.
Move the sliders and watch the exp and log curves reflect across the diagonal, then enter two numbers in the product panel and compare the log of their product with the sum of their logs.
#Cross-entropy is the same idea
#Sigma and Pi: Sums and Products in One Symbol
Start with a small sum worked by hand: add the first four square numbers, 1 squared, 2 squared, 3 squared and 4 squared. Keep a running total as you go.
| i | Term (i squared) | Running total |
|---|---|---|
| 1 | 1 | 1 |
| 2 | 4 | 1 + 4 = 5 |
| 3 | 9 | 5 + 9 = 14 |
| 4 | 16 | 14 + 16 = 30 |
The answer is 30. Sigma is the short way to write that whole table. Here is the plainest case, adding four numbers x_1, x_2, x_3 and x_4:
Here is the same sum of i squared as a plain Python loop. The total starts at 0, and each pass through the loop adds the next term:
total = 0
for i in range(1, 5): # range(1, 5) counts 1, 2, 3, 4 and stops before 5
total = total + i ** 2
print(total) # 30sum(...) adds up everything it is given, so sum(i ** 2 for i in range(1, 5)) is the one-line version of the loop. Later you will also see zip(a, b), which walks along two lists at once and hands you the pair in each position: zip([1, 2, 3], [4, 5, 6]) gives (1, 4), (2, 5), (3, 6).The stepper below does the same table for several expressions.
Pick an expression, press the step button, and watch each term join the running total, then switch to a pi expression and see the running product.
#Four sigma expressions you will meet constantly
Each of these is a sum in disguise. Work through them once and the later lessons will read easily.
#Pi: the same idea for products
The capital Greek letter pi, written with its big product symbol, means "multiply these" in exactly the same way sigma means "add these". The index, bounds and subscripts work identically.
Now the bridge from the previous section. Take the log of a pi expression and the product rule turns it into a sigma expression:
That one line is the whole reason maximum likelihood methods are written as sums of log-probabilities.
#Reading a Formula Word by Word
A formula is a sentence. It has a handful of symbols, and once you know the alphabet you can read it left to right. Here is the alphabet that ML uses most.
| Symbol | Name | Typical meaning |
|---|---|---|
| α (alpha) | alpha | a learning rate or a small tuning constant |
| β (beta) | beta | a coefficient, or a momentum setting |
| θ (theta) | theta | the parameters (weights) of a model |
| μ (mu) | mu | the mean of a distribution |
| σ (sigma, lowercase) | sigma | a standard deviation. Do not confuse with capital Σ, which means "sum" |
| λ (lambda) | lambda | a regularisation strength, or a rate |
| ε (epsilon) | epsilon | a tiny number, or noise |
| x with a hat, x̂ | x-hat | an estimate or prediction of x |
| x with a bar, x̄ | x-bar | the mean of x |
| x_i | subscript | the i-th item, or a label such as x_train |
| x squared, x to the n | superscript | a power, or sometimes an index in brackets such as x(i) |
| a over b | fraction bar | divide the top by the bottom |
| arg max | argument of the maximum | the input that gives the largest output, not the largest output itself |
Two reading habits help. First, the same letter can mean different things in different books, so check the line near the formula that says what each symbol stands for. Second, brackets and bars group things. Work the inside first.
#Reading softmax
Softmax is the function that turns a list of raw scores z into probabilities. Before the formula, build it by hand for just two scores, z = [1, 0].
- Make each score positive with exp: exp(1) = 2.718 and exp(0) = 1.
- Add them to get the total: 2.718 + 1 = 3.718.
- Divide each by the total: 2.718 / 3.718 = 0.731 and 1 / 3.718 = 0.269.
The two answers add to 1, and the larger score got the larger share. The formula is that recipe written for any number K of scores (K is just how many classes there are).
Read it word by word:
- softmax(z) with subscript i means "the answer for class number i".
- The top, e to the z_i, takes class i's score and makes it positive.
- The fraction bar says: divide.
- The bottom is a sigma. Its index j runs from 1 to K, meaning "all K classes". It adds up e to the z_j for every class.
- Together: class i's share of the total of all the positive scores.
#Reading the Gaussian density
| Piece | What it computes | Value at x = 1 |
|---|---|---|
| x minus mu | distance from the centre | 1 |
| squared | so left and right count the same | 1 |
| divided by two sigma squared | scale by the width | 1 / 2 = 0.5 |
| minus, then exp | 1 at the centre, shrinking outward | exp(-0.5) = 0.6065 |
| times one over sigma root two pi | a constant (pi is the circle number, 3.14159) | 0.3989 x 0.6065 = 0.2420 |
Stack those five pieces together and you have the formula.
Word by word, working from the middle outward:
- x minus mu is how far the point is from the centre.
- Squared, so left and right count the same.
- Divided by two sigma squared, so a wide curve (big sigma) treats the same distance as less surprising.
- A minus sign, then exp. At the centre the exponent is 0 and exp(0) = 1. Away from the centre it goes negative, so exp gives a number between 0 and 1.
- The front factor rescales everything so the total area under the curve is exactly 1. The area under a curve is the amount of space between the curve and the flat axis beneath it. For a probability curve that area is the total probability of every possible outcome, so it must come to exactly 1.
#Sigmoid and Softmax: Exp Plus Normalise
| Class | Score z | exp(z) | Probability |
|---|---|---|---|
| A | 2.0 | 7.389 | 7.389 / 11.213 = 0.659 |
| B | 1.0 | 2.718 | 2.718 / 11.213 = 0.242 |
| C | 0.1 | 1.105 | 1.105 / 11.213 = 0.099 |
| Total | 11.213 | 1.000 |
The biggest score gets the biggest share, but every class keeps some probability. Because exp grows quickly, a score gap of one point already gives A nearly three times B's share.
Change one score in the widget above and watch all the probabilities move together, then add 100 to every score and see that nothing changes.
Notice the same σ letter again. Here, sigma(t) is the sigmoid function. Context tells you which sigma you are looking at: capital Σ is a sum, lowercase σ is usually a standard deviation, and σ followed by a bracket is often the sigmoid.
Why ML cares: softmax and sigmoid sit at the end of nearly every classifier, and cross-entropy then takes the log of the probability they produce. Exp and log are two halves of one pipeline.
#A Slope Is Rise Over Run
For a straight line like f(x) = 2x + 1 the slope is the same everywhere: 2. For a curve it changes. Take f(x) = x squared and look at the point x = 3 with a shrinking run:
| Between | Rise | Run | Slope |
|---|---|---|---|
| x = 3 and x = 4 | 16 - 9 = 7 | 1 | 7.0 |
| x = 3 and x = 3.1 | 9.61 - 9 = 0.61 | 0.1 | 6.1 |
| x = 3 and x = 3.01 | 9.0601 - 9 = 0.0601 | 0.01 | 6.01 |
As the run shrinks, the slope settles toward 6. That settled value is the slope of the curve exactly at x = 3. Chasing it as the run shrinks to nothing is the idea behind derivatives, and it is where the next lessons pick up. In ML, the slope of the loss tells a model which way to nudge its weights to improve.
#Try It Yourself
Tests · Verify the naive product equals 0.0 (underflow), the log-likelihood is a normal finite negative number, the average per event is the log-likelihood divided by 400, and the cross-entropy for p = 0.7 on the true class is about 0.3567.
When you run the solution you should see the naive product print as 0.0, so the "product is zero" check prints True. The log-likelihood is -794.38, the average per event is -1.9859, and the cross-entropy is 0.3567.
#Key Takeaways
- A function is a rule with one output per input, and f(g(x)) means do g first, then f. A neural network is a long composition of simple functions.
- Exponents count repeated multiplication. A negative exponent means one over, a fractional exponent means a root, and e = 2.71828 is the base where exp(x) has a slope equal to its own value.
- A logarithm is the undo of an exponent. Its three rules are product to sum, power to multiplier, and change of base.
- Multiplying 400 probabilities of 0.1 gives exactly 0.0 in 64-bit floats, while the sum of their logs is a normal number, -921.03. This is why ML works with log-probabilities.
- Sigma means sum and pi means product. The log of a pi is a sigma, so a likelihood becomes a log-likelihood.
- Mean, dot product, variance and cross-entropy are all sigma expressions, and each one has a one-line Python sum() version.
- Softmax and sigmoid are exp followed by a normalising division. A slope is rise over run, and a derivative is what that slope settles to when the run shrinks.
#Quick Check
With f(x) = x + 3 and g(x) = x squared, what is f(g(2))?