Can similarity be calculated? An introduction to vector addition, the dot product and cosine similarity.
1. When distance gets it wrong
Imagine you are building a program that finds similar articles. Represent each article's content with a word-frequency vector: how many times “AI” appears, how many times “data” appears, and so on. Consider these three articles:
Figure 2-1. Counts of ‘AI’ and ‘data’ form two-dimensional vectors. A is a long AI article, B is a short AI article, and C focuses on databases.
Article A (8 occurrences of “AI”, 6 of “data”): a detailed article discussing AI and data, around 3,000 Chinese characters long.
Article B (2 occurrences of “AI”, 1 of “data”): similar content to A, but a short summary, around 800 Chinese characters long.
Article C (1 occurrence of “AI”, 7 of “data”): mostly about databases, with occasional mentions of AI.
Intuitively, A and B are the most alike: the same topic, with different lengths. Let us measure similarity in the most natural way: how far apart the two points are on the diagram. The Pythagorean theorem is enough. Find the horizontal and vertical differences; the hypotenuse gives the distance:
According to this calculation, A and C are closer. A long article about AI seems more similar to a database article that barely mentions AI than to a summary on the same topic. That is clearly the wrong result for this task.
2. Why length interferes
Why did distance fail? Draw A, B and C as arrows starting at the origin, and the answer becomes clear. A and B point in very similar directions — both towards “AI and data” — but B is shorter. C points in a different direction, closer to the “data” axis.
Straight-line distance measures how far apart the arrow tips are, so the short length of B distorts the comparison. For this task, similarity of content should depend on the arrows' directions, rather than their lengths.
Figure 2-2. A and B point in almost the same direction, about 10° apart. A and C are about 45° apart. Direction distinguishes the topic from the article's length.
3. The geometry behind the dot product
To calculate the angle, first meet a new tool: the dot product. Its calculation is simple: multiply corresponding components, then add the results:
a⋅b=a1b1+a2b2+⋯+anbn=i=1∑naibi
This formula is easy to remember, but its geometric meaning matters more. Imagine projecting the tip of vector b perpendicularly onto the line containing vector a. The signed distance from the origin to that point is the projection of b in the direction of a:
Figure 2-3. Project b onto a. The dot product equals the length of a multiplied by the signed length of that projection.
This gives us an intuitive interpretation of the dot product: how closely do the two vectors align? The more of b points along a, the larger their dot product is.
Similar directions
a⋅b>0
Perpendicular
a⋅b=0
Opposing directions
a⋅b<0
Figure 2-4. The sign of the dot product reflects directional alignment: positive for alignment, zero for perpendicular vectors, and negative for opposing directions.
4. Cosine similarity
From the geometric dot-product formula, we can isolate cosθ:
a⋅bcosθ=∥a∥∥b∥cosθ=∥a∥∥b∥a⋅b
This value is called cosine similarity. Dividing by the lengths of both vectors removes the influence of length, leaving direction. For non-zero vectors, it ranges from −1 to 1:
1: exactly the same direction — fully aligned.
0: perpendicular directions — no directional alignment.
Figure 2-5. Cosine similarity identifies A and B as closely aligned (about 0.98), while A and C differ more (about 0.71). Article length no longer determines the comparison.
This is a widely used tool for measuring similarity in AI. Search engines, recommendation systems and semantic search can compare text vectors using cosine similarity rather than straight-line distance.
5. Vector addition: arithmetic with meaning
Vectors can also be added and subtracted. The rule is simple: add or subtract corresponding components:
[a1,a2,…]+[b1,b2,…]=[a1+b1,a2+b2,…]
Geometrically, add two vectors by placing the tail of the second arrow at the tip of the first. This is the parallelogram rule. In AI, this simple operation leads to a remarkable result.
When neural networks learn word vectors from large amounts of text, researchers have found relationships such as:
vking−vman+vwoman≈vqueenFigure 2-6. A simplified picture of two conceptual directions: gender and royalty. The same offset from man to woman takes king to queen. The axes illustrate the idea; learned embeddings do not have manually labelled dimensions.
What does this mean? Relationships between words can be encoded as directions in vector space. “Man → woman” corresponds to an offset, and “commoner → royalty” corresponds to another. These offsets can combine independently, leading approximately to the expected result.
Lesson 15, on word vectors, will explain how these representations are learned. For now, remember that vector addition and subtraction can perform arithmetic on meaning.
6. What you have discovered
The takeaway. Straight-line distance measures “how far apart the tips are”; cosine similarity measures “how closely the directions align”. When a vector's length is irrelevant — for example, when word counts grow with article length — cosine similarity removes that influence. The dot product has a geometric meaning: one vector's length multiplied by the other's signed projection onto its direction. Vector subtraction can also express analogies such as “vking−vman+vwoman≈vqueen”, turning semantic relationships into offsets we can calculate.
What you have discovered in this lesson:
The blind spot of straight-line distance: vectors with similar directions but different lengths can be farther apart than vectors with different directions.
Dot product:a⋅b=∑iaibi=∥a∥∥b∥cosθ. Geometrically, it is the length of a multiplied by the projection of b onto a's direction.
Cosine similarity:cosθ=∥a∥∥b∥a⋅b. It removes the influence of length and compares direction.
Vector addition: relationships between meanings can support arithmetic — “vking−vman+vwoman≈vqueen”.
Check your understanding
Choose an answer and select “Submit answer” to see the correct option and an explanation.