In brief
- An AI learns by repeating a loop: it makes a prediction, measures its error, then slightly adjusts its parameters. After billions of repetitions, the error becomes small.
- There are three main ways to learn: from corrected examples (supervised), by finding patterns on its own (unsupervised), or through trial and reward (reinforcement).
- Assistants like ChatGPT combine several methods: predicting the next word across enormous volumes of text, then being fine-tuned with human feedback (RLHF).
On this page
The learning loop: predict, measure the error, correct
Whatever the method, learning almost always comes down to repeating the same four steps, millions or even billions of times.
- Predict
The model receives a piece of data (a photo, a sentence with a missing word) and offers an answer. At first its parameters are random, so the answer is nonsense.
- Measure the error
Its answer is compared with the right answer using a loss function, a number that tells you “how wrong the model is.”
- Correct
An algorithm works out which way, and by how much, to change each parameter to reduce the error. This correction relies on gradient descent and backpropagation, both explained below.
- Repeat
On to the next example. The error gradually shrinks.
Gradient descent: walking down a mountain in the fog
Imagine you are lost in the fog on a mountain and want to reach the valley. You can’t see the landscape, but you can feel the slope under your feet. With each step, you move a little in the direction that goes down the most. That is exactly what gradient descent does: altitude represents the error, and each step adjusts the parameters to reduce it.
Calculating the “slope” for each of the billions of parameters is done efficiently by backpropagation of the gradient. This method, popularized in 1986 by David Rumelhart, Geoffrey Hinton and Ronald Williams, works backward from the final result to the first layers to determine how much each parameter is responsible for the error.
Three main ways to learn
1. Supervised learning: learning with a teacher
The model is given labeled examples: “this image is a cat,” “this email is spam,” “this X-ray shows a tumor.” It learns to predict the label. This is the method behind image recognition (such as AlexNet in 2012 with the ImageNet dataset), spam filters and bank fraud detection. Its weak point: it needs a lot of correct examples, often annotated by humans.
2. Unsupervised and self-supervised learning: learning without labels
Here, no “right answer” is provided. The model has to discover structure on its own, for example by grouping customers with similar behavior. An essential variant, self-supervised learning, creates its own exercises from raw data: a word is hidden in a sentence, and the model has to guess it. This is how large language models are pretrained on billions of pages of text.
3. Reinforcement learning: learning through trial and reward
An “agent” acts in an environment and receives rewards or penalties. It learns the strategy that maximizes its long-term gains, like a child learning to ride a bike. This is the method that allowed AlphaGo (Google DeepMind) to beat Lee Sedol in March 2016, after playing a very large number of games against itself. It is also used to train robots and to fine-tune conversational assistants.
What about ChatGPT? Learning in several stages
A modern conversational assistant doesn’t learn in just one way. It combines several methods, one after the other:
- Self-supervised pretraining
The model reads immense volumes of text (web pages, books, code, encyclopedias) and learns to predict the next word, or more precisely the next “token.” This is the most expensive stage.
- Supervised fine-tuning
Humans write examples of good answers to prompts. The model learns to follow instructions rather than simply continue a text. OpenAI presented this approach with InstructGPT in 2022.
- Reinforcement learning from human feedback (RLHF)
Human reviewers rank several of the model’s answers from best to worst. A second model, called a “reward model,” learns to predict these preferences, then is used to fine-tune the assistant through reinforcement learning. Anthropic offers a variant, “Constitutional AI,” in which the model is guided by a list of written principles.
For more detail on these stages, see how AI is built and how ChatGPT works.
The pitfalls of learning
- Overfitting: the model memorizes its training examples instead of generalizing. It aces the exercises it already knows and fails on new ones. You can detect it by testing the model on data it has never seen.
- Data bias: a model trained on unbalanced or discriminatory data reproduces those flaws (see limits and risks).
- Label quality: if examples are poorly annotated, the model learns mistakes. “Garbage in, garbage out” is a classic saying in the field.
- Energy cost: training the largest models mobilizes thousands of processors for weeks.
Frequently asked questions
How does an AI actually learn?
It repeats a loop: it makes a prediction, measures its error with a loss function, then slightly adjusts its parameters to reduce that error. After billions of repetitions on data, the error becomes small.
What is the difference between supervised and unsupervised learning?
In supervised learning, each example comes with the right answer (for instance, an image labeled “cat”). In unsupervised learning, no right answer is provided: the model looks for patterns in the data on its own.
What is backpropagation?
It is the algorithm that calculates, for each parameter in a neural network, how much it contributed to the final error. It makes it possible to efficiently correct billions of parameters. It was popularized in 1986 by Rumelhart, Hinton and Williams.
What is RLHF?
RLHF (reinforcement learning from human feedback) is a stage in which humans rank an AI’s answers; a reward model learns those preferences, then is used to fine-tune the assistant so it responds in a more helpful and safer way.
Does an AI keep learning after it is released?
Generally not: a deployed model has frozen parameters. It can keep the context of a conversation, but it doesn’t learn from it. Providers can, however, train new versions with collected data, depending on their privacy policy.
Sources and references
- Rumelhart, Hinton & Williams, “Learning representations by back-propagating errors,” Nature, 1986
- Ouyang et al. (OpenAI), “Training language models to follow instructions with human feedback” (InstructGPT), 2022
- Wikipedia — Reinforcement learning from human feedback
- Bai et al. (Anthropic), “Constitutional AI: Harmlessness from AI Feedback,” 2022
- Wikipedia — AlphaGo versus Lee Sedol