Course chapters
Notes / AI & large language models
From vectors to Transformers
A step-by-step introduction to the ideas behind large language models.
4 lessons are available. The remaining lessons are listed below and are not available yet.
Start with What is a vector? →IFoundational mathematics
- 01
- 02
- 03
- 04
- 05What is a gradient?
Not available yet
- 06Probability and information
Not available yet
IINeural networks
- 07The structure of a neuron
Not available yet
- 08Activation functions
Not available yet
- 09Neural networks and training
Not available yet
- 10Loss functions
Not available yet
- 11Softmax
Not available yet
- 12Gradient descent and optimisers
Not available yet
- 13Backpropagation
Not available yet
IIINatural language processing
- 14The probability game of language: N-grams
Not available yet
- 15Word vectors
Not available yet
- 16Feedforward neural language models
Not available yet
- 17Recurrent neural networks
Not available yet
- 18Long short-term memory
Not available yet
IVLarge language models
- 19Attention
Not available yet
- 20Multi-head attention
Not available yet
- 21The Transformer architecture
Not available yet
- 22Tokenizers
Not available yet
- 23Encoders, decoders and large language models
Not available yet
- 24Residual connections and layer normalisation
Not available yet
- 25Pretraining, supervised fine-tuning and reinforcement learning
Not available yet
- 26KV caching, sparse attention and FlashAttention
Not available yet
- 27Mixture of experts
Not available yet
- 28Model distillation
Not available yet
- 29From N-grams to Transformers: a recap
Not available yet
- 30Frontiers and the road ahead
Not available yet
VAppendices
- A1A panoramic view of the Transformer
Not available yet
- A2Attention Is All You Need: the original paper
Not available yet