In brief
- Building a modern AI takes three resources: data, computing power (GPUs) and teams of researchers and engineers.
- The process follows clear stages: data collection and cleaning, pretraining, fine-tuning, alignment with human preferences, evaluation, then deployment.
- The cost is considerable for the largest models, and the electricity used by data centers could more than double by 2030, according to the International Energy Agency.
On this page
Overview: the life cycle of an AI model
Creating an assistant like ChatGPT, Claude, Gemini or Mistral isn’t a matter of “programming” intelligence. It is an industrial project with several stages, involving hundreds of people and gigantic infrastructure.
Let’s look at each stage in detail.
Step 1: gathering and preparing the data
A language model is trained on enormous volumes of text. For GPT-3 (2020), the model was trained on about 300 billion “tokens,” drawn from a corpus of roughly 500 billion; more than half came from Common Crawl, a public copy of a large part of the web, supplemented by books, Wikipedia pages and other corpora.
But raw data can’t be used as it is. It has to be:
- filtered to remove low-quality content, duplicates and spam;
- stripped, as far as possible, of personal information and illegal content;
- balanced across sources and languages;
- handled lawfully: the use of copyrighted works to train AI is the subject of several lawsuits, such as the one filed by The New York Times against OpenAI and Microsoft in December 2023.
Step 2: pretraining, the heaviest phase
The model learns to predict the next token across the entire corpus. This is the most compute-hungry stage, and the one that gives the model its general knowledge of language and culture.
It relies on GPUs (graphics processing units), processors originally designed for video games that can perform thousands of small calculations in parallel, which suits neural networks perfectly. In 2012, AlexNet ran on just two consumer graphics cards. Today, a large model uses tens of thousands of GPUs in data centers, for weeks or months.
According to orders of magnitude published in the literature (these are estimates, since companies don’t always disclose their real costs): about $50,000 for GPT-2 (2019), about $8 million for PaLM (2022), and hundreds of millions of dollars for the largest recent models. In January 2025, DeepSeek claimed to have trained its V3 model for about $5.6 million in compute, a figure that, according to observers, covers only part of the total cost.
“Scaling laws” observed by researchers show that performance improves steadily as you increase the model’s size, the amount of data and the computing power together. These observations, published notably by OpenAI (2020) and then by DeepMind with the Chinchilla model (2022), have driven the race toward ever-larger models.
Steps 3 and 4: fine-tuning and alignment
At the end of pretraining, the model knows how to “continue a text,” but it isn’t an assistant yet: if you ask it a question, it might answer with another question, or carry on like a forum post. Two stages make it useful and safe:
- Supervised fine-tuning: the model is retrained on thousands to tens of thousands of examples of good conversations, written by humans.
- Alignment, notably through RLHF (reinforcement learning from human feedback): reviewers rank answers, and the model is adjusted to produce responses that are more helpful, honest and harmless. Some labs add written principles (Constitutional AI, at Anthropic).
These stages cost far less than pretraining, but they are decisive for how good the assistant feels to use.
Step 5: evaluation and safety testing
Before releasing a model, teams put it through benchmarks: batteries of standardized tests (math, code, general knowledge, reading comprehension, and more). They also run red teaming: experts deliberately try to trick the model to find its weaknesses (dangerous answers, data leaks, bias).
Step 6: deployment and continuous improvement
The model is then made available: through a chat website, an app, or an API that developers build into their software. Every answer requires computing power (known as inference), which is why some services are paid or limit the number of requests.
Once a model is live, providers monitor how it is used, fix flaws and regularly release new versions. Some companies publish their models’ “weights,” which anyone can then download and run themselves: these are called open-weight models, like those from Mistral AI or DeepSeek. Others keep their models closed and offer them only through their own services.
The energy and environmental cost
Data centers use electricity to train models, but also to answer millions of daily requests. According to the International Energy Agency (IEA), data centers consumed about 415 TWh in 2024, or around 1.5% of the world’s electricity. The IEA projects a doubling, to about 945 TWh in 2030, with AI as one of the main drivers of the increase.
The impact varies widely depending on the energy source, the cooling and the efficiency of the models. Researchers are working to reduce this cost: smaller models, more efficient chips, compression techniques. For users, there are simple steps too: choose a model that fits your task rather than the biggest one, and avoid unnecessary requests.
Who builds AI?
The development of the largest models is dominated by a handful of players: OpenAI (founded in December 2015), Google DeepMind, Anthropic (founded in 2021), Meta, Mistral AI (France, 2023), DeepSeek (China, 2023) and others. According to Stanford’s 2026 AI Index, more than 90% of notable models come from industry rather than academia, and both the United States and China play a major role.
You can discover and compare the assistants from these players in our AI tools directory, for example ChatGPT, Claude, Gemini or Mistral Le Chat.
Frequently asked questions
How do you create an artificial intelligence?
You gather large amounts of data, train a neural network on it using computing power (GPUs), fine-tune it with human examples and then human feedback, evaluate it, and finally deploy it through an app or an API.
How much does it cost to train an AI?
It varies enormously: from a few thousand dollars for a small model to hundreds of millions for the largest. As a published order of magnitude, about $50,000 for GPT-2 (2019) and about $8 million for PaLM (2022). The real figures for recent models are rarely public.
Why are GPUs used for AI?
Because these processors, designed for graphics, can perform thousands of calculations at the same time. Training a neural network consists of a huge number of similar operations, which are perfectly suited to parallel computing.
Can you build your own AI?
Yes, on a small scale. You can train a small model on your own data, or more simply adapt an existing open model (fine-tuning) on a computer or a cloud service. Training a very large model from scratch remains reserved for organizations with considerable resources.
Does AI use a lot of electricity?
Collectively, yes: according to the IEA, data centers accounted for about 1.5% of the world’s electricity in 2024 (415 TWh) and could reach about 945 TWh in 2030. A single request uses little, but the total is high with millions of users.
Sources and references
- IEA — Energy and AI (2025)
- Stanford HAI — AI Index Report 2026
- Wikipedia — GPT-3
- Wikipedia — Large language model (training costs and scaling laws)
- Wikipedia — DeepSeek
- Hoffmann et al. (DeepMind), “Training Compute-Optimal Large Language Models” (Chinchilla), 2022
- Kaplan et al. (OpenAI), “Scaling Laws for Neural Language Models,” 2020