1. Many rules, one table
Put mathematics aside for a moment. Imagine you have taken over a bubble tea shop. Your menu has three signature drinks, each with its own recipe using the same four ingredients: tea, milk, sugar and pearls.
Here are the recipes, with numbers representing portions:
- Pearl milk tea: 2 portions of tea + 1 of milk + 0 of sugar + 3 of pearls.
- Milk pudding: 0 portions of tea + 4 of milk + 1 of sugar + 2 of pearls.
- Traditional sweet tea: 1 portion of tea + 0 of milk + 3 of sugar + 1 of pearls.
Today you are buying ingredients. Let the price per portion of each ingredient be , , and . You want to know immediately: how much does each drink cost to make?
The calculation is simple. Use the portions in each recipe as weights and take a weighted sum of the ingredient prices:
Three drinks give us three weighted sums. The weights are simply the quantities of each ingredient in the recipes. In practice, there might be hundreds of products and hundreds of ingredients. A neural network can have millions of weights.
Rather than writing each formula separately, mathematicians arrange all the recipe quantities into a rectangular table, called a matrix.
A recipe table is a matrix.
| Recipe | Tea | Milk | Sugar | Pearls |
|---|---|---|---|---|
| Pearl milk tea | 2 | 1 | 0 | 3 |
| Milk pudding | 0 | 4 | 1 | 2 |
| Traditional sweet tea | 1 | 0 | 3 | 1 |
We describe a matrix's size as : rows — the number of drinks here, or the number of outputs — and columns — the number of ingredients, or the number of inputs. In this example, and .
Keep this picture in mind: one recipe per row, one ingredient per column. We can return to this table to understand the matrix operations that follow.
2. Matrix times vector: one dot product per row
Continue with the bubble tea shop. Suppose today's prices per portion are: tea 1 yuan, milk 2 yuan, sugar 1 yuan, and pearls 0 yuan — the pearls are free as part of a promotion. Write these prices as the column vector , with components .
To calculate the costs of all three drinks at once, multiply . This means: take the dot product of each row of — each recipe — with — the ingredient prices — and collect the results in a cost vector. We learned the dot product in the previous lesson: multiply corresponding components and add them.
Row 1 · Pearl milk tea
Row 2 · Milk pudding
Row 3 · Traditional sweet tea
The results are 4 yuan for pearl milk tea, 9 yuan for milk pudding, and 4 yuan for traditional sweet tea. One matrix multiplication calculates the costs of the entire menu together.
This is what makes matrices useful: they turn “calculate many weighted sums at once” into a single operation. Each neural network layer uses this operation, with a much larger table of weights. A complete layer can also add a bias and apply an activation function; here we are focusing on the matrix–vector multiplication.
3. From n dimensions to m dimensions
There is one central rule for matrix–vector multiplication. Remember the shapes:
Two points matter:
- The columns must match the input: if has columns, must have components. A mismatch means the product cannot be calculated.
- The rows determine the output: if has rows, has components. A matrix can map an -dimensional input space to an -dimensional output space.
This ability to move “from one number of dimensions to another” is central to the next lesson, What is a linear transformation? A matrix is more than a calculation tool. Geometrically, it represents a transformation of space, such as rotation, scaling, compression or stretching.
4. What you have discovered
The takeaway. A matrix neatly arranges many sets of weights: each row is a set of coefficients, and each column corresponds to one input. Multiplying a matrix by a vector, , means taking the dot product of each row with . The rows produce results, forming an -dimensional output vector. The dimension rule is : the inner dimensions must match, and the outer dimensions determine the result.
What you have discovered in this lesson:
- Matrix: an table of numbers with rows and columns, organising multiple sets of coefficients into one object.
- Matrix–vector multiplication: takes the dot product of each row with , calculating several weighted sums at once.
- Dimension rule: . turns an -dimensional input into an -dimensional output.
- Matching inner dimensions: the number of columns of must equal the number of components of , so that the dot products can line up.
Check your understanding
Choose an answer and select “Submit answer” to see the correct option and an explanation.