← All notes
Course chapters
On this page

Lesson 02 · Foundational mathematics

Common vector operations

Can similarity be calculated? An introduction to vector addition, the dot product and cosine similarity.

1. When distance gets it wrong

Imagine you are building a program that finds similar articles. Represent each article's content with a word-frequency vector: how many times “AI” appears, how many times “data” appears, and so on. Consider these three articles:

024681002468Occurrences of ‘AI’ →Occurrences of ‘data’ ↑
A(8,6)A (8, 6)
B(2,1)B (2, 1)
C(1,7)C (1, 7)
Figure 2-1. Counts of ‘AI’ and ‘data’ form two-dimensional vectors. A is a long AI article, B is a short AI article, and C focuses on databases.
  • Article A (8 occurrences of “AI”, 6 of “data”): a detailed article discussing AI and data, around 3,000 Chinese characters long.
  • Article B (2 occurrences of “AI”, 1 of “data”): similar content to A, but a short summary, around 800 Chinese characters long.
  • Article C (1 occurrence of “AI”, 7 of “data”): mostly about databases, with occasional mentions of AI.

Intuitively, A and B are the most alike: the same topic, with different lengths. Let us measure similarity in the most natural way: how far apart the two points are on the diagram. The Pythagorean theorem is enough. Find the horizontal and vertical differences; the hypotenuse gives the distance:

d(A,B)=(8−2)2+(6−1)2=36+25≈7.8d(A,C)=(8−1)2+(6−7)2=49+1≈7.1\begin{aligned} d(A,B) &= \sqrt{(8-2)^2+(6-1)^2} = \sqrt{36+25} \approx 7.8 \\ d(A,C) &= \sqrt{(8-1)^2+(6-7)^2} = \sqrt{49+1} \approx 7.1 \end{aligned}

According to this calculation, A and C are closer. A long article about AI seems more similar to a database article that barely mentions AI than to a summary on the same topic. That is clearly the wrong result for this task.

2. Why length interferes

Why did distance fail? Draw A, B and C as arrows starting at the origin, and the answer becomes clear. A and B point in very similar directions — both towards “AI and data” — but B is shorter. C points in a different direction, closer to the “data” axis.

Straight-line distance measures how far apart the arrow tips are, so the short length of B distorts the comparison. For this task, similarity of content should depend on the arrows' directions, rather than their lengths.

024681002468Occurrences of ‘AI’ →Occurrences of ‘data’ ↑
A(8,6)A (8, 6)
B(2,1)B (2, 1)
C(1,7)C (1, 7)
Figure 2-2. A and B point in almost the same direction, about 10° apart. A and C are about 45° apart. Direction distinguishes the topic from the article's length.
Pause and think

Ignore length and use only direction. Can you think of a similarity measure that looks only at direction? Pause and think for 30 seconds.

I've thought about it — show the answerHide answer

Use the angle between two vectors to measure similarity: the smaller the angle, the more closely their directions align, and the more similar they are. The angle between A and B is about 10∘10^\circ, while the angle between A and C is about 45∘45^\circ. The comparison now matches our intuition. One question remains: how do we calculate the angle?

3. The geometry behind the dot product

To calculate the angle, first meet a new tool: the dot product. Its calculation is simple: multiply corresponding components, then add the results:

a⋅b=a1b1+a2b2+⋯+anbn=∑i=1naibi\mathbf{a}\cdot\mathbf{b}=a_1b_1+a_2b_2+\cdots+a_nb_n=\sum_{i=1}^{n}a_ib_i

This formula is easy to remember, but its geometric meaning matters more. Imagine projecting the tip of vector b\mathbf{b} perpendicularly onto the line containing vector a\mathbf{a}. The signed distance from the origin to that point is the projection of b\mathbf{b} in the direction of a\mathbf{a}:

θ\theta
a\mathbf{a}
b\mathbf{b}
∥b∥cos⁡θ=projection length\lVert\mathbf{b}\rVert\cos\theta=\text{projection length}
a⋅b=∥a∥×projection length\mathbf{a}\cdot\mathbf{b}=\lVert\mathbf{a}\rVert\times\text{projection length}
Figure 2-3. Project b onto a. The dot product equals the length of a multiplied by the signed length of that projection.

This gives us an intuitive interpretation of the dot product: how closely do the two vectors align? The more of b\mathbf{b} points along a\mathbf{a}, the larger their dot product is.

Similar directions

a\mathbf{a}
b

a⋅b>0\mathbf{a}\cdot\mathbf{b}>0

Perpendicular

a\mathbf{a}
b

a⋅b=0\mathbf{a}\cdot\mathbf{b}=0

Opposing directions

a\mathbf{a}
b

a⋅b<0\mathbf{a}\cdot\mathbf{b}<0

Figure 2-4. The sign of the dot product reflects directional alignment: positive for alignment, zero for perpendicular vectors, and negative for opposing directions.
Pause and think

A neuron calculates w⋅x\mathbf{w}\cdot\mathbf{x}. What does that mean geometrically?

I've thought about it — show the answerHide answer

w⋅x\mathbf{w}\cdot\mathbf{x} equals the projection of the input vector x\mathbf{x} in the direction of the weight vector w\mathbf{w}, multiplied by ∥w∥\lVert\mathbf{w}\rVert. A larger projection means the input aligns more closely with the direction preferred by this neuron, producing a larger dot-product response. If x\mathbf{x} and w\mathbf{w} are perpendicular, their dot product is zero. This is the geometry behind a neuron's selective response. Here we are considering the dot-product calculation alone, before adding a bias or applying an activation function.

4. Cosine similarity

From the geometric dot-product formula, we can isolate cos⁡θ\cos\theta:

a⋅b=∥a∥∥b∥cos⁡θcos⁡θ=a⋅b∥a∥∥b∥\begin{aligned} \mathbf{a}\cdot\mathbf{b} &= \lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert\cos\theta \\ \cos\theta &= \frac{\mathbf{a}\cdot\mathbf{b}}{\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert} \end{aligned}

This value is called cosine similarity. Dividing by the lengths of both vectors removes the influence of length, leaving direction. For non-zero vectors, it ranges from −1 to 1:

  • 1: exactly the same direction — fully aligned.
  • 0: perpendicular directions — no directional alignment.
  • −1: opposite directions — fully opposed.

A and B

A⋅B=8×2+6×1=22∥A∥=10,∥B∥≈2.24cos⁡θ≈2210×2.24≈0.98\begin{aligned} A\cdot B&=8\times2+6\times1=22\\ \lVert A\rVert&=10,\quad\lVert B\rVert\approx2.24\\ \cos\theta&\approx\frac{22}{10\times2.24}\approx0.98\end{aligned}

Closely aligned directions.

A and C

A⋅C=8×1+6×7=50∥A∥=10,∥C∥≈7.07cos⁡θ≈5070.7≈0.71\begin{aligned} A\cdot C&=8\times1+6\times7=50\\ \lVert A\rVert&=10,\quad\lVert C\rVert\approx7.07\\ \cos\theta&\approx\frac{50}{70.7}\approx0.71\end{aligned}

Noticeably different directions.

Figure 2-5. Cosine similarity identifies A and B as closely aligned (about 0.98), while A and C differ more (about 0.71). Article length no longer determines the comparison.

This is a widely used tool for measuring similarity in AI. Search engines, recommendation systems and semantic search can compare text vectors using cosine similarity rather than straight-line distance.

5. Vector addition: arithmetic with meaning

Vectors can also be added and subtracted. The rule is simple: add or subtract corresponding components:

[a1,a2,…]+[b1,b2,…]=[a1+b1,a2+b2,…][a_1,a_2,\ldots]+[b_1,b_2,\ldots]=[a_1+b_1,a_2+b_2,\ldots]

Geometrically, add two vectors by placing the tail of the second arrow at the tip of the first. This is the parallelogram rule. In AI, this simple operation leads to a remarkable result.

When neural networks learn word vectors from large amounts of text, researchers have found relationships such as:

vking−vman+vwoman≈vqueen\mathbf{v}_{\text{king}}-\mathbf{v}_{\text{man}}+\mathbf{v}_{\text{woman}}\approx\mathbf{v}_{\text{queen}}
‘Gender’ direction →‘Royalty’ direction ↑manwomankingqueen
+vwoman−vman+\mathbf{v}_{\text{woman}}-\mathbf{v}_{\text{man}}
+vwoman−vman+\mathbf{v}_{\text{woman}}-\mathbf{v}_{\text{man}}
+royalty+\text{royalty}
vking−vman+vwoman≈vqueen\mathbf{v}_{\text{king}}-\mathbf{v}_{\text{man}}+\mathbf{v}_{\text{woman}}\approx\mathbf{v}_{\text{queen}}
Figure 2-6. A simplified picture of two conceptual directions: gender and royalty. The same offset from man to woman takes king to queen. The axes illustrate the idea; learned embeddings do not have manually labelled dimensions.

What does this mean? Relationships between words can be encoded as directions in vector space. “Man → woman” corresponds to an offset, and “commoner → royalty” corresponds to another. These offsets can combine independently, leading approximately to the expected result.

Lesson 15, on word vectors, will explain how these representations are learned. For now, remember that vector addition and subtraction can perform arithmetic on meaning.

6. What you have discovered

The takeaway. Straight-line distance measures “how far apart the tips are”; cosine similarity measures “how closely the directions align”. When a vector's length is irrelevant — for example, when word counts grow with article length — cosine similarity removes that influence. The dot product has a geometric meaning: one vector's length multiplied by the other's signed projection onto its direction. Vector subtraction can also express analogies such as “vking−vman+vwoman≈vqueen\mathbf{v}_{\text{king}}-\mathbf{v}_{\text{man}}+\mathbf{v}_{\text{woman}}\approx\mathbf{v}_{\text{queen}}”, turning semantic relationships into offsets we can calculate.

What you have discovered in this lesson:

  • The blind spot of straight-line distance: vectors with similar directions but different lengths can be farther apart than vectors with different directions.
  • Dot product: a⋅b=∑iaibi=∥a∥∥b∥cos⁡θ\mathbf{a}\cdot\mathbf{b}=\sum_i a_ib_i=\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert\cos\theta. Geometrically, it is the length of a\mathbf{a} multiplied by the projection of b\mathbf{b} onto a\mathbf{a}'s direction.
  • Cosine similarity: cos⁡θ=a⋅b∥a∥∥b∥\cos\theta=\frac{\mathbf{a}\cdot\mathbf{b}}{\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert}. It removes the influence of length and compares direction.
  • Vector addition: relationships between meanings can support arithmetic — “vking−vman+vwoman≈vqueen\mathbf{v}_{\text{king}}-\mathbf{v}_{\text{man}}+\mathbf{v}_{\text{woman}}\approx\mathbf{v}_{\text{queen}}”.

Check your understanding

Choose an answer and select “Submit answer” to see the correct option and an explanation.

1. Which operation measures how closely two vectors align in direction, without being affected by their lengths?
2. What does the dot product of two vectors produce, and what does its value describe?