In brief
- A modern AI is an immense network of small computing units, “artificial neurons,” connected to one another by numbers called weights or parameters.
- Each neuron adds up the signals it receives, weighs them, then decides whether or not to pass a signal on. Stacked over dozens of layers, these simple operations make it possible to recognize images or produce language.
- All of an AI’s “knowledge” is stored in its parameters: GPT-3 has 175 billion. These are not facts filed away in memory, but numerical settings.
On this page
The basic idea: rules that are worked out, not written
Imagine you had to write a program that recognizes a cat in a photo. Writing out every rule by hand (“two pointy ears, whiskers…”) is almost impossible: cats come in every shape, color, and posture. Modern AI gets around the problem: we show it millions of photos labeled “cat” or “not cat,” and it adjusts its own internal settings until it can tell them apart.
This adjustment process is called training (see how an AI learns). This article focuses instead on what you end up with: the model itself, and what happens when you use it.
The artificial neuron: a small calculator
The artificial neuron is inspired (very loosely) by the neurons of the brain. The first mathematical model dates from 1943 (Warren McCulloch and Walter Pitts). Here is what it does:
- It receives inputs
These are numbers: for example, the brightness of the pixels in an image, or the numerical representation of a word.
- It weighs each input
Each input is multiplied by a weight, which indicates its importance. A high weight means “this signal matters a lot”; a negative weight means “this signal pushes in the opposite direction.”
- It adds everything up
It computes the sum of the weighted inputs, then adds a small offset called the bias.
- It applies an activation function
A simple rule decides how strong the signal passed on will be (for example: “pass nothing on if the sum is negative”). Without this step, stacking neurons would be pointless: the result would remain a simple linear operation.
The neural network: layers that extract meaning
Neurons are organized in layers. The first receives the raw data, the last gives the result, and the intermediate layers (called “hidden” layers) gradually transform the information. This is what we call deep learning: “deep” refers to the number of layers.
In a network that recognizes images, you generally see a progression: the first layers detect simple features (edges, patches of color), the next ones shapes (eyes, ears, wheels), and the last ones whole objects (face, cat, car). Nobody programmed this hierarchy: it emerges from training.
Parameters: where the model’s “knowledge” lives
The full set of a model’s weights and biases is called its parameters. These are the “tuning knobs” that training adjusts. A model contains neither a database of facts nor sentences memorized in readable form: its knowledge is spread across billions of numbers.
The more parameters a model has, the more nuance it can capture, as long as there is enough data and computing power to train it. But bigger is not always better: smaller but better-trained models can outperform larger ones, and they cost less to run.
Training and inference: two moments not to be confused
The life of an AI model is divided into two very different phases.
| Training | Inference (use) | |
|---|---|---|
| What happens? | The model adjusts its parameters based on data | The model, now frozen, produces a response to your request |
| When? | Once (or in large phases), before it is put into service | Every time a user makes a request |
| Duration | Weeks or months for a large model | A few seconds |
| Cost | Very high: thousands of GPUs and millions of dollars for the largest models | Low per request, but enormous in total with millions of users |
| Does the model learn from my conversation? | Yes, but that is precisely what happens at this stage | No: during a conversation, it does not change its parameters (even though it can keep the context of the discussion) |
Not all neural networks are alike
The network’s architecture is adapted to the type of data:
- Convolutional networks (CNNs): specialized in images. Popularized by Yann LeCun starting in the 1990s, they had their moment of glory with AlexNet in 2012.
- Recurrent networks (RNNs, LSTMs): designed for sequences (text, sound), reading word by word. They dominated language processing until 2017.
- Transformers: introduced in 2017, they read an entire sequence in parallel thanks to an attention mechanism. They underpin ChatGPT, Claude, Gemini, and also many image and audio models (see how ChatGPT works).
- Diffusion models: specialized in generating images, video, and sound (see generative AI).
What to keep in mind
- An AI doesn’t “know”: it computes. Its answer is the result of millions of numerical operations, not of conscious reasoning.
- We don’t understand everything. Neural networks are often “black boxes”: we know how to train them, but it is hard to explain precisely why they give a particular answer. Interpretability is an active area of research.
- A model reflects its data. If it was trained on biased or incomplete data, its answers will be too (see limits and risks).
Frequently asked questions
How does a neural network work in simple terms?
A neural network is a series of layers of small computing units. Each unit adds up weighted signals and passes a result on to the next layer. By adjusting the weights based on examples, the network learns to recognize patterns: a face, a word, a sound.
What is a parameter in AI?
A parameter is a number (a weight or a bias) inside the model, adjusted during training. An AI’s knowledge is spread across its parameters: GPT-3 has 175 billion of them.
What is the difference between training and inference?
Training is the long, costly phase in which the model adjusts its parameters based on data. Inference is the use of the already-trained model to respond to a request.
Does an AI store information like a database?
No. A model does not contain searchable files: its knowledge is encoded as numbers spread across its parameters, which is why it can mix up or invent details.
Why is AI called a “black box”?
Because even when we know a neural network’s architecture and data, it is very hard to explain why it produces exactly a given answer in a given case. Model interpretability is an active field of research.