← All notes
Course chapters
On this page

Lesson 03 · Foundational mathematics

What is a matrix?

Each layer of a neural network turns one collection of numbers into another. How can we represent that transformation?

1. Many rules, one table

Put mathematics aside for a moment. Imagine you have taken over a bubble tea shop. Your menu has three signature drinks, each with its own recipe using the same four ingredients: tea, milk, sugar and pearls.

Here are the recipes, with numbers representing portions:

  • Pearl milk tea: 2 portions of tea + 1 of milk + 0 of sugar + 3 of pearls.
  • Milk pudding: 0 portions of tea + 4 of milk + 1 of sugar + 2 of pearls.
  • Traditional sweet tea: 1 portion of tea + 0 of milk + 3 of sugar + 1 of pearls.

Today you are buying ingredients. Let the price per portion of each ingredient be x1x_1, x2x_2, x3x_3 and x4x_4. You want to know immediately: how much does each drink cost to make?

The calculation is simple. Use the portions in each recipe as weights and take a weighted sum of the ingredient prices:

y1=2x1+1x2+0x3+3x4(pearl milk tea)y2=0x1+4x2+1x3+2x4(milk pudding)y3=1x1+0x2+3x3+1x4(traditional sweet tea)\begin{aligned} y_1 &= 2x_1+1x_2+0x_3+3x_4 &&\text{(pearl milk tea)} \\ y_2 &= 0x_1+4x_2+1x_3+2x_4 &&\text{(milk pudding)} \\ y_3 &= 1x_1+0x_2+3x_3+1x_4 &&\text{(traditional sweet tea)} \end{aligned}

Three drinks give us three weighted sums. The weights are simply the quantities of each ingredient in the recipes. In practice, there might be hundreds of products and hundreds of ingredients. A neural network can have millions of weights.

Rather than writing each formula separately, mathematicians arrange all the recipe quantities into a rectangular table, called a matrix.

A recipe table is a matrix.

WW — portions per recipe
RecipeTeaMilkSugarPearls
Pearl milk tea2103
Milk pudding0412
Traditional sweet tea1031
Figure 3-1. WW is a 3×43\times4 matrix: each row is a recipe, and each column is an ingredient. Read “3 × 4” as “three rows and four columns”.

We describe a matrix's size as m×nm \times n: mm rows — the number of drinks here, or the number of outputs — and nn columns — the number of ingredients, or the number of inputs. In this example, m=3m = 3 and n=4n = 4.

Keep this picture in mind: one recipe per row, one ingredient per column. We can return to this table to understand the matrix operations that follow.

2. Matrix times vector: one dot product per row

Continue with the bubble tea shop. Suppose today's prices per portion are: tea 1 yuan, milk 2 yuan, sugar 1 yuan, and pearls 0 yuan — the pearls are free as part of a promotion. Write these prices as the column vector x\mathbf{x}, with components [1,2,1,0][1, 2, 1, 0].

To calculate the costs of all three drinks at once, multiply WxW\mathbf{x}. This means: take the dot product of each row of WW — each recipe — with x\mathbf{x} — the ingredient prices — and collect the results in a cost vector. We learned the dot product in the previous lesson: multiply corresponding components and add them.

WW

[210304121031]\begin{bmatrix}2 & 1 & 0 & 3 \\ 0 & 4 & 1 & 2 \\ 1 & 0 & 3 & 1\end{bmatrix}

x\mathbf{x}

[1210]\begin{bmatrix}1 \\ 2 \\ 1 \\ 0\end{bmatrix}

y\mathbf{y}

[494]\begin{bmatrix}4 \\ 9 \\ 4\end{bmatrix}

Row 1 · Pearl milk tea

2×1+1×2+0×1+3×0=42 \times 1 + 1 \times 2 + 0 \times 1 + 3 \times 0 = 4

Row 2 · Milk pudding

0×1+4×2+1×1+2×0=90 \times 1 + 4 \times 2 + 1 \times 1 + 2 \times 0 = 9

Row 3 · Traditional sweet tea

1×1+0×2+3×1+1×0=41 \times 1 + 0 \times 2 + 3 \times 1 + 1 \times 0 = 4

Figure 3-2. Take the dot product of each row with x. The three results form the output column vector y=[4,9,4]\mathbf{y}=[4,9,4].

The results are 4 yuan for pearl milk tea, 9 yuan for milk pudding, and 4 yuan for traditional sweet tea. One matrix multiplication calculates the costs of the entire menu together.

This is what makes matrices useful: they turn “calculate many weighted sums at once” into a single operation. Each neural network layer uses this operation, with a much larger table of weights. A complete layer can also add a bias and apply an activation function; here we are focusing on the matrix–vector multiplication.

Pause and think

WW is a 3×43\times4 matrix. Why must x\mathbf{x} have four components, rather than three or five?

I've thought about it — show the answerHide answer

Each column of WW corresponds to one component of x\mathbf{x}. They must line up to calculate a dot product. If WW has nn columns, x\mathbf{x} must have nn components; otherwise, we cannot pair the positions one by one, and the multiplication is not defined. The general rule is: an m×nm\times n matrix multiplies an nn-dimensional vector, producing an mm-dimensional vector.

3. From n dimensions to m dimensions

There is one central rule for matrix–vector multiplication. Remember the shapes:

W⏟m×nx⏟n×1=y⏟m×1\underbrace{W}_{m\times n}\underbrace{\mathbf{x}}_{n\times1}=\underbrace{\mathbf{y}}_{m\times1}
WW
m×nm \times n
mm
nn
×\times
x\mathbf{x}
n×1n \times 1
==
y\mathbf{y}
m×1m \times 1
Inner dimensions must match: n\text{Inner dimensions must match: } n
m outputsm\text{ outputs}
Figure 3-3. The number of columns of W must match the input dimension. The number of rows determines the output dimension: (m×n)×(n×1)=(m×1)(m\times n)\times(n\times1)=(m\times1).

Two points matter:

  • The columns must match the input: if WW has nn columns, x\mathbf{x} must have nn components. A mismatch means the product cannot be calculated.
  • The rows determine the output: if WW has mm rows, y\mathbf{y} has mm components. A matrix can map an nn-dimensional input space to an mm-dimensional output space.

This ability to move “from one number of dimensions to another” is central to the next lesson, What is a linear transformation? A matrix is more than a calculation tool. Geometrically, it represents a transformation of space, such as rotation, scaling, compression or stretching.

4. What you have discovered

The takeaway. A matrix neatly arranges many sets of weights: each row is a set of coefficients, and each column corresponds to one input. Multiplying a matrix by a vector, WxW\mathbf{x}, means taking the dot product of each row with x\mathbf{x}. The mm rows produce mm results, forming an mm-dimensional output vector. The dimension rule is Wm×nxn×1=ym×1W_{m\times n}\mathbf{x}_{n\times1}=\mathbf{y}_{m\times1}: the inner dimensions must match, and the outer dimensions determine the result.

What you have discovered in this lesson:

  • Matrix: an m×nm \times n table of numbers with mm rows and nn columns, organising multiple sets of coefficients into one object.
  • Matrix–vector multiplication: WxW\mathbf{x} takes the dot product of each row with x\mathbf{x}, calculating several weighted sums at once.
  • Dimension rule: (m×n)×(n×1)=(m×1)(m \times n) \times (n \times 1) = (m \times 1). WW turns an nn-dimensional input into an mm-dimensional output.
  • Matching inner dimensions: the number of columns of WW must equal the number of components of x\mathbf{x}, so that the dot products can line up.

Check your understanding

Choose an answer and select “Submit answer” to see the correct option and an explanation.

1. A neural network layer turns one collection of numbers into another. Which mathematical object naturally represents this linear transformation?
2. When a 3×43\times4 matrix multiplies a vector, how many components must the input have, and how many does the output have?