How it worksGuide 3 of 9

How Does Artificial Intelligence Work?

9 min readUpdated on By the CheblAI editorial team
Read inFrançaisEnglish中文

In brief

  • A modern AI is an immense network of small computing units, “artificial neurons,” connected to one another by numbers called weights or parameters.
  • Each neuron adds up the signals it receives, weighs them, then decides whether or not to pass a signal on. Stacked over dozens of layers, these simple operations make it possible to recognize images or produce language.
  • All of an AI’s “knowledge” is stored in its parameters: GPT-3 has 175 billion. These are not facts filed away in memory, but numerical settings.
On this page
The basic idea: rules that are worked out, not writtenThe artificial neuron: a small calculatorThe neural network: layers that extract meaningParameters: where the model’s “knowledge” livesTraining and inference: two moments not to be confusedNot all neural networks are alikeWhat to keep in mindFrequently asked questionsSources and references

The basic idea: rules that are worked out, not written

Imagine you had to write a program that recognizes a cat in a photo. Writing out every rule by hand (“two pointy ears, whiskers…”) is almost impossible: cats come in every shape, color, and posture. Modern AI gets around the problem: we show it millions of photos labeled “cat” or “not cat,” and it adjusts its own internal settings until it can tell them apart.

This adjustment process is called training (see how an AI learns). This article focuses instead on what you end up with: the model itself, and what happens when you use it.

The artificial neuron: a small calculator

The artificial neuron is inspired (very loosely) by the neurons of the brain. The first mathematical model dates from 1943 (Warren McCulloch and Walter Pitts). Here is what it does:

Diagram of an artificial neuron: several inputs multiplied by weights, added together, passed through an activation function, to produce an output
Diagram 1 — An artificial neuron: inputs, weights, sum, activation, output.
  1. It receives inputs

    These are numbers: for example, the brightness of the pixels in an image, or the numerical representation of a word.

  2. It weighs each input

    Each input is multiplied by a weight, which indicates its importance. A high weight means “this signal matters a lot”; a negative weight means “this signal pushes in the opposite direction.”

  3. It adds everything up

    It computes the sum of the weighted inputs, then adds a small offset called the bias.

  4. It applies an activation function

    A simple rule decides how strong the signal passed on will be (for example: “pass nothing on if the sum is negative”). Without this step, stacking neurons would be pointless: the result would remain a simple linear operation.

The neural network: layers that extract meaning

Neurons are organized in layers. The first receives the raw data, the last gives the result, and the intermediate layers (called “hidden” layers) gradually transform the information. This is what we call deep learning: “deep” refers to the number of layers.

Diagram of a three-layer neural network: input, hidden layers, output, using the example of recognizing a cat
Diagram 2 — A neural network: information flows from left to right and is transformed at each layer.

In a network that recognizes images, you generally see a progression: the first layers detect simple features (edges, patches of color), the next ones shapes (eyes, ears, wheels), and the last ones whole objects (face, cat, car). Nobody programmed this hierarchy: it emerges from training.

Parameters: where the model’s “knowledge” lives

The full set of a model’s weights and biases is called its parameters. These are the “tuning knobs” that training adjusts. A model contains neither a database of facts nor sentences memorized in readable form: its knowledge is spread across billions of numbers.

Bar chart of the number of parameters in GPT-1, GPT-2, and GPT-3: about 117 million, 1.5 billion, and 175 billion
Diagram 3 — The number of parameters in GPT models grew roughly 1,500-fold in two years (logarithmic scale).

The more parameters a model has, the more nuance it can capture, as long as there is enough data and computing power to train it. But bigger is not always better: smaller but better-trained models can outperform larger ones, and they cost less to run.

Training and inference: two moments not to be confused

The life of an AI model is divided into two very different phases.

TrainingInference (use)
What happens?The model adjusts its parameters based on dataThe model, now frozen, produces a response to your request
When?Once (or in large phases), before it is put into serviceEvery time a user makes a request
DurationWeeks or months for a large modelA few seconds
CostVery high: thousands of GPUs and millions of dollars for the largest modelsLow per request, but enormous in total with millions of users
Does the model learn from my conversation?Yes, but that is precisely what happens at this stageNo: during a conversation, it does not change its parameters (even though it can keep the context of the discussion)

Not all neural networks are alike

The network’s architecture is adapted to the type of data:

  • Convolutional networks (CNNs): specialized in images. Popularized by Yann LeCun starting in the 1990s, they had their moment of glory with AlexNet in 2012.
  • Recurrent networks (RNNs, LSTMs): designed for sequences (text, sound), reading word by word. They dominated language processing until 2017.
  • Transformers: introduced in 2017, they read an entire sequence in parallel thanks to an attention mechanism. They underpin ChatGPT, Claude, Gemini, and also many image and audio models (see how ChatGPT works).
  • Diffusion models: specialized in generating images, video, and sound (see generative AI).

What to keep in mind

  • An AI doesn’t “know”: it computes. Its answer is the result of millions of numerical operations, not of conscious reasoning.
  • We don’t understand everything. Neural networks are often “black boxes”: we know how to train them, but it is hard to explain precisely why they give a particular answer. Interpretability is an active area of research.
  • A model reflects its data. If it was trained on biased or incomplete data, its answers will be too (see limits and risks).

Frequently asked questions

How does a neural network work in simple terms?

A neural network is a series of layers of small computing units. Each unit adds up weighted signals and passes a result on to the next layer. By adjusting the weights based on examples, the network learns to recognize patterns: a face, a word, a sound.

What is a parameter in AI?

A parameter is a number (a weight or a bias) inside the model, adjusted during training. An AI’s knowledge is spread across its parameters: GPT-3 has 175 billion of them.

What is the difference between training and inference?

Training is the long, costly phase in which the model adjusts its parameters based on data. Inference is the use of the already-trained model to respond to a request.

Does an AI store information like a database?

No. A model does not contain searchable files: its knowledge is encoded as numbers spread across its parameters, which is why it can mix up or invent details.

Why is AI called a “black box”?

Because even when we know a neural network’s architecture and data, it is very hard to explain why it produces exactly a given answer in a given case. Model interpretability is an active field of research.

Sources and references

  1. McCulloch & Pitts, “A logical calculus of the ideas immanent in nervous activity,” 1943
  2. Rumelhart, Hinton & Williams, “Learning representations by back-propagating errors,” Nature, 1986
  3. Wikipedia — Large language model
  4. Wikipedia — GPT-3
  5. Wikipedia — Artificial neural network

Independent editorial content. The facts, dates and figures cited rely on the sources listed at the end of the page; this content is for information only and does not constitute professional advice.