Language modelsGuide 6 of 9

How Does ChatGPT Work? Large Language Models Explained

10 min readUpdated on By the CheblAI editorial team
Read inFrançaisEnglish中文

In brief

  • ChatGPT, Claude and Gemini are large language models (LLMs): they generate text by predicting, token after token, the most likely continuation of a sentence.
  • Their strength comes from the Transformer architecture (2017) and its attention mechanism, which lets the model connect each word to every other word in the context.
  • Because they predict plausible text rather than check facts, they can produce “hallucinations”: wrong answers stated with confidence.
On this page
The principle: predicting the next wordTokens: how a machine reads textThe Transformer and attention: the engine of LLMsThe context window: the model’s working memoryFrom language model to assistantWhy ChatGPT sometimes gets it wrong: hallucinationsThe main assistants todayFrequently asked questionsSources and references

The principle: predicting the next word

A large language model (LLM) is a neural network trained on a colossal amount of text with one goal that seems modest: predict what comes next. Given “The cat sleeps on the,” it estimates the probability of each possible next word (“couch,” “roof,” “rug”…), then picks one.

The chosen word is then added to the sentence, and the process starts over: entire paragraphs emerge from this repetition. Everything you see, whether an answer, a poem or a computer program, is produced one small piece at a time.

Diagram of next-word prediction: the sentence “The cat sleeps on the” yields probabilities for couch, roof, rug and other words
Diagram 1 — A language model estimates the probability of each possible continuation (illustrative values).

Tokens: how a machine reads text

A computer only understands numbers. So text is first split into small pieces called tokens, which can be whole words, parts of words or punctuation marks. Each token is assigned a number in a vocabulary of several tens of thousands of entries.

Each token is then converted into a long list of numbers called an embedding. These numbers place the word on a kind of map of meaning: related words (“dog” and “cat”) end up in neighboring spots. This idea was popularized by word2vec (Tomas Mikolov, Google, 2013).

In practice, a token represents about three-quarters of a word in English; French, Chinese and other languages generally need more tokens for the same text. Models bill and limit usage in tokens.

The Transformer and attention: the engine of LLMs

Before 2017, language models read sentences word by word, with limited memory. On June 12, 2017, eight Google researchers published “Attention Is All You Need,” which introduced the Transformer architecture.

Its central idea is the attention mechanism: to understand a word, the model looks at all the other words in the context and decides which ones matter most. In “The bank refused the loan because it was cautious,” attention helps link “it” to “bank” rather than to “loan.”

Diagram of the attention mechanism: the word “it” is connected by links of varying thickness to the other words in the sentence, especially “bank”
Diagram 2 — Attention: each word weighs the importance of the other words in the sentence (simplified diagram).

Two advantages explain this architecture’s success. First, it processes all the words in parallel, which makes full use of GPUs and allows much larger models to be trained. Second, it handles long-range dependencies well: linking a word at the start of a text with one much further along. The paper has been cited more than 250,000 times and ranks among the most cited of the 21st century.

The context window: the model’s working memory

A model can only “see” a limited amount of text at a time: its context window. It includes your question, the conversation history and the answer in progress. GPT-3 (2020) was limited to 2,048 tokens, or a few pages; recent models accept far more, up to hundreds of thousands of tokens (entire books) for some.

Beyond this window, the model “forgets” what came before. It has no permanent memory of your conversations, unless the tool adds a separate memory feature.

From language model to assistant

A pretrained model on its own just continues a text. To turn it into an assistant, it is taught to follow instructions (fine-tuning) and then to favor the answers humans judge helpful (RLHF), as explained in how AI learns. ChatGPT launched on November 30, 2022, based on GPT-3.5, which was trained this way.

Today, assistants add capabilities on top of the base model:

  • Web search: the assistant queries a search engine and cites sources (this is the principle behind Perplexity AI).
  • Tools: running code, analyzing a file, doing calculations, calling an app.
  • Multimodality: understanding and producing images, sound, even video.
  • Reasoning: some models “think” longer before answering, producing intermediate steps, which improves results on complex problems.

Why ChatGPT sometimes gets it wrong: hallucinations

A hallucination is an answer that sounds fluent and credible but is false or made up. It follows directly from how these models work, as described above: the model is trained to produce plausible text, not to verify a fact. When it lacks information, it can fill the gap with a believable answer.

This is not just an anecdote. In 2023, two American lawyers filed court decisions invented by ChatGPT with a federal court in New York (the Mata v. Avianca case); the judge imposed a $5,000 sanction. In February 2024, a Canadian tribunal held Air Canada liable for incorrect information its chatbot gave about a bereavement fare.

The main assistants today

Several families of models share the market. They evolve quickly; here are the main ones, with their makers:

AssistantMakerWhat stands out
ChatGPTOpenAI (United States)The best known; 900 million weekly users announced in early 2026
ClaudeAnthropic (United States)Built on “Constitutional AI”; popular for long texts and code
GeminiGoogleIntegrated into Google services; multimodal
Mistral Le ChatMistral AI (France)A European player with open models
Perplexity AIPerplexityAn answer engine that cites its sources

For a closer comparison, use our comparison tool or browse the directory. The beginner’s guide explains how to choose.

Frequently asked questions

How does ChatGPT work in simple terms?

ChatGPT is a large language model: a neural network trained to predict the next word (or token) in a text. By repeating that prediction word after word, it produces entire answers. It was then fine-tuned with human examples and feedback to behave like an assistant.

What is a token in AI?

A token is a small piece of text (a word, part of a word or a punctuation mark) converted into a number so the model can process it. In French, a word often counts as one or two tokens.

What is the Transformer architecture?

It is a neural network architecture published by Google researchers on June 12, 2017. It uses an attention mechanism that connects each word to every other word in the context, and it processes text in parallel. It underpins most large language models.

Why does ChatGPT make things up?

Because it is designed to produce plausible text, not to verify facts. When it lacks information, it can generate an answer that sounds believable but is false: this is what is called a hallucination. That is why important information should always be checked.

Does ChatGPT have access to the internet?

The model itself doesn’t browse the internet: its knowledge stops at a training cutoff date. But ChatGPT and other assistants can be equipped with a web search feature that lets them look up recent information and cite sources.

Sources and references

  1. Vaswani et al., “Attention Is All You Need,” arXiv, 2017
  2. Wikipedia — Large language model
  3. Wikipedia — ChatGPT
  4. Wikipedia — Hallucination (artificial intelligence)
  5. Wikipedia — Mata v. Avianca, Inc.
  6. Mikolov et al., “Efficient Estimation of Word Representations in Vector Space” (word2vec), 2013
  7. Brown et al. (OpenAI), “Language Models are Few-Shot Learners” (GPT-3), 2020

Independent editorial content. The facts, dates and figures cited rely on the sources listed at the end of the page; this content is for information only and does not constitute professional advice.