1. The same , a different view
In the previous lesson, we learned matrix–vector multiplication, : take the dot product of each row of with to obtain each component of the output. That was a calculation-based view — treat the matrix as a calculator, feed it one vector, and receive another.
In this lesson, we will look at the same through a different lens: what is it doing geometrically? Do not worry about big words such as “space” and “transformation” just yet. Start with one vector and work step by step.
First, recall a fact from Lesson 1. A vector is a list of numbers, but we can also draw it as an arrow starting at the origin, with its tip at a point on graph paper. For example, points one unit right and one unit up.
Second, here is the key observation: if the input is a point and the output is also a point, the geometric action of matrix multiplication is quite simple: it sends the input point to another position.
Try a concrete example:
This sends to . Drawing it on a coordinate plane makes the action clear:
One matrix multiplication assigns a new position to a point. So far, this is still the story of just one point.
Third comes the real leap: what if the same transforms every point on the graph paper? Instead of immediately imagining infinitely many points, take a small step: from one point to four.
Take a unit square, with corners , , and . Multiply each corner by the same — C is the point we just used in Figure 4-1:
Before · square, area 1
After · parallelogram, area 3
Connect the new corners in their original order, and the square becomes a parallelogram. Look at it for a few seconds. The origin A has not moved. The edges are still straight, opposite sides are still parallel, and the area has increased from 1 to 3. has reshaped the entire region.
Now extend this idea: not only the four corners, but every point inside and outside the square follows the same rule. Think of the “free transform” tool in an image editor: move a control point, and thousands of pixels move together according to one rule.
When the same acts on every point, the entire coordinate grid takes on a new shape. It might rotate, stretch, compress or shear. This is what it means to view as a transformation of the whole space. It is no longer about moving a single isolated point, but applying one rule everywhere. For the non-degenerate examples here, the grid lines stay straight and each family stays parallel and evenly spaced:
Before · original grid
After · transformed grid
2. Watch the basis vectors
The coordinate plane has infinitely many points. Must we calculate each new position separately? No. There is a shortcut: if we know where the two basis vectors go, we know where the entire plane goes.
The basis vectors are the two unit-length arrows along the coordinate axes: , pointing right, and , pointing up. Every point in the plane is a linear combination of them. For example, means “three steps along , then two steps along ”: .
A matrix transformation has a useful property: it preserves this combination rule. If moves to and moves to , then moves to . Knowing just those two new positions lets us determine every point's new position. And those positions are exactly the two columns of :
To understand a matrix geometrically in two dimensions, take its two columns, draw them as arrows, and see where they point. They tell you how the plane transforms.
One clarification matters: reading the columns works in any dimension; having exactly two columns is specific to a two-dimensional input. An -dimensional input has standard basis vectors, one along each axis, so its transformation matrix has columns. Column tells you where the basis vector on axis goes. The principle is unchanged; there are simply more arrows than we can draw. The matrices with thousands of rows and columns in neural networks are this same idea in higher dimensions.
3. Rotation, scaling, shear and reflection
Different matrices produce different transformations. Use the same unit square to explore four common types. The dashed outline shows the original square; the coloured shape shows the result. Each diagram includes its matrix. Remember that its columns give the new positions of and :
Scaling
Stretch horizontally; compress vertically.
Rotation
Rotate anticlockwise.
Shear
Push the square into a parallelogram.
Reflection
Reflect across the horizontal axis.
For these examples, rotation, scaling, shear and reflection preserve straight edges, parallel opposite sides and the origin. The square remains a parallelogram. More generally, a linear transformation can also collapse a dimension, turning a square into a line or even a point.
The defining rules of linearity are that the transformation preserves vector addition and scalar multiplication:
Geometrically, the origin stays fixed. Straight lines do not become curves: they map to straight lines or, if a dimension collapses, to points. This is the mathematical meaning behind “linear”.
What are linear transformations used for? At heart, they let us use one matrix to rearrange a whole collection of vectors according to the same rule. This simple operation supports many applications:
- Geometric transformations: rotation, scaling, reflection and projection in games, 3D rendering, image editing and CAD. Transform the coordinates of each vertex using the same matrix, just as in Figure 4-5.
- Looking at data in a new coordinate system: project data onto new axes to emphasise informative directions and reduce redundant dimensions. Dimensionality reduction, such as PCA, uses linear projections. PCA commonly centres the data first.
- Neural network layers: the core operation transforms the data space. Layers can reshape the representation to make samples easier to separate. Biases and nonlinear activation functions extend what the network can express; linear transformations alone cannot make every classification problem linearly separable.
Here is a concrete example: rotate a photo anticlockwise. Imagine its pixels in a Cartesian coordinate system, with increasing to the right and increasing upwards. Apply the same rotation matrix to every pixel coordinate:
Each pixel moves to . For example, moves to . Thousands of pixels follow the same rule, rotating the image together. A real image editor also accounts for the rotation centre, image bounds and its pixel-coordinate convention, but this matrix describes the geometric rotation.
Linear transformations are a basic tool for moving and reorganising data in a neural network. Understanding them gives you a geometric view of the matrix operation within a layer.
4. To translate the space, add
Linear transformations are powerful, but there is one thing they cannot do: translate the whole space. The reason is already familiar: they keep the origin fixed. No matter how we rotate or stretch the space, still maps to .
Neural networks often need a shift as well. The solution is straightforward: add a bias vector after , shifting the transformed space to a new position:
· origin stays fixed
· translate by
This is the core operation of a fully connected layer. 's shape determines the input and output dimensions. Each component of provides a learnable baseline offset for one output. With non-zero , this is an affine transformation, rather than a linear transformation. A layer may then apply an activation function.
5. Many layers can still be one
If each layer computes , can stacking many of them produce a much more complex transformation? Follow the mathematics for two layers, without an activation function between them:
Two layers equal one. Multiply and to form one new matrix. Combine into one new bias. Two affine layers therefore give the same result as one affine layer. The same reasoning applies no matter how many such layers we stack. Geometrically, rotating and then stretching still amounts to another linear transformation; adding translations gives another affine transformation.
A network containing only these operations has an expressive limit. With a linear classification score and a threshold, it can only draw a straight-line boundary — a hyperplane in higher dimensions. It cannot, for example, precisely enclose a circular region. To go further, place nonlinearity between the layers. That is the role of the activation functions we will study later.
The takeaway. applies one transformation to the whole space: rotation, stretching, shear or reflection. Read a matrix by looking at its columns: each is the transformed position of a standard basis vector. A linear transformation preserves addition and scalar multiplication and fixes the origin. Add to translate the result, giving the affine operation used in a fully connected layer. Stacking these operations without nonlinear activations is still equivalent to one affine operation.
What you have learned to see differently:
- Geometric view: transforms the whole coordinate grid, not just a single point.
- Basis-vector shortcut: the matrix columns specify where the basis vectors go. There are two for a two-dimensional input, and for an -dimensional input.
- Four transformations: scaling, rotation, shear and reflection. The examples preserve straight, parallel, evenly spaced grid lines.
- : reshapes the space; translates it. Together, they form the affine part of a fully connected layer.
- The expressive limit: stacking linear or affine transformations without nonlinearity does not increase the family of transformations the network can represent.
Check your understanding
Choose an answer and select “Submit answer” to see the correct option and an explanation.