BuildingGuide 5 of 9

How Is Artificial Intelligence Built?

10 min readUpdated on By the CheblAI editorial team
Read inFrançaisEnglish中文

In brief

  • Building a modern AI takes three resources: data, computing power (GPUs) and teams of researchers and engineers.
  • The process follows clear stages: data collection and cleaning, pretraining, fine-tuning, alignment with human preferences, evaluation, then deployment.
  • The cost is considerable for the largest models, and the electricity used by data centers could more than double by 2030, according to the International Energy Agency.
On this page
Overview: the life cycle of an AI modelStep 1: gathering and preparing the dataStep 2: pretraining, the heaviest phaseSteps 3 and 4: fine-tuning and alignmentStep 5: evaluation and safety testingStep 6: deployment and continuous improvementThe energy and environmental costWho builds AI?Frequently asked questionsSources and references

Overview: the life cycle of an AI model

Creating an assistant like ChatGPT, Claude, Gemini or Mistral isn’t a matter of “programming” intelligence. It is an industrial project with several stages, involving hundreds of people and gigantic infrastructure.

Six-step diagram: data, pretraining, fine-tuning, alignment, evaluation, deployment
Diagram 1 — The six main stages of building an AI assistant.

Let’s look at each stage in detail.

Step 1: gathering and preparing the data

A language model is trained on enormous volumes of text. For GPT-3 (2020), the model was trained on about 300 billion “tokens,” drawn from a corpus of roughly 500 billion; more than half came from Common Crawl, a public copy of a large part of the web, supplemented by books, Wikipedia pages and other corpora.

But raw data can’t be used as it is. It has to be:

  • filtered to remove low-quality content, duplicates and spam;
  • stripped, as far as possible, of personal information and illegal content;
  • balanced across sources and languages;
  • handled lawfully: the use of copyrighted works to train AI is the subject of several lawsuits, such as the one filed by The New York Times against OpenAI and Microsoft in December 2023.

Step 2: pretraining, the heaviest phase

The model learns to predict the next token across the entire corpus. This is the most compute-hungry stage, and the one that gives the model its general knowledge of language and culture.

It relies on GPUs (graphics processing units), processors originally designed for video games that can perform thousands of small calculations in parallel, which suits neural networks perfectly. In 2012, AlexNet ran on just two consumer graphics cards. Today, a large model uses tens of thousands of GPUs in data centers, for weeks or months.

According to orders of magnitude published in the literature (these are estimates, since companies don’t always disclose their real costs): about $50,000 for GPT-2 (2019), about $8 million for PaLM (2022), and hundreds of millions of dollars for the largest recent models. In January 2025, DeepSeek claimed to have trained its V3 model for about $5.6 million in compute, a figure that, according to observers, covers only part of the total cost.

“Scaling laws” observed by researchers show that performance improves steadily as you increase the model’s size, the amount of data and the computing power together. These observations, published notably by OpenAI (2020) and then by DeepMind with the Chinchilla model (2022), have driven the race toward ever-larger models.

Steps 3 and 4: fine-tuning and alignment

At the end of pretraining, the model knows how to “continue a text,” but it isn’t an assistant yet: if you ask it a question, it might answer with another question, or carry on like a forum post. Two stages make it useful and safe:

  • Supervised fine-tuning: the model is retrained on thousands to tens of thousands of examples of good conversations, written by humans.
  • Alignment, notably through RLHF (reinforcement learning from human feedback): reviewers rank answers, and the model is adjusted to produce responses that are more helpful, honest and harmless. Some labs add written principles (Constitutional AI, at Anthropic).

These stages cost far less than pretraining, but they are decisive for how good the assistant feels to use.

Step 5: evaluation and safety testing

Before releasing a model, teams put it through benchmarks: batteries of standardized tests (math, code, general knowledge, reading comprehension, and more). They also run red teaming: experts deliberately try to trick the model to find its weaknesses (dangerous answers, data leaks, bias).

Step 6: deployment and continuous improvement

The model is then made available: through a chat website, an app, or an API that developers build into their software. Every answer requires computing power (known as inference), which is why some services are paid or limit the number of requests.

Once a model is live, providers monitor how it is used, fix flaws and regularly release new versions. Some companies publish their models’ “weights,” which anyone can then download and run themselves: these are called open-weight models, like those from Mistral AI or DeepSeek. Others keep their models closed and offer them only through their own services.

The energy and environmental cost

Data centers use electricity to train models, but also to answer millions of daily requests. According to the International Energy Agency (IEA), data centers consumed about 415 TWh in 2024, or around 1.5% of the world’s electricity. The IEA projects a doubling, to about 945 TWh in 2030, with AI as one of the main drivers of the increase.

Two-bar chart: data center electricity consumption, 415 TWh in 2024 and about 945 TWh projected in 2030
Diagram 2 — Worldwide data center electricity consumption (source: IEA, Energy and AI report, 2025).

The impact varies widely depending on the energy source, the cooling and the efficiency of the models. Researchers are working to reduce this cost: smaller models, more efficient chips, compression techniques. For users, there are simple steps too: choose a model that fits your task rather than the biggest one, and avoid unnecessary requests.

Who builds AI?

The development of the largest models is dominated by a handful of players: OpenAI (founded in December 2015), Google DeepMind, Anthropic (founded in 2021), Meta, Mistral AI (France, 2023), DeepSeek (China, 2023) and others. According to Stanford’s 2026 AI Index, more than 90% of notable models come from industry rather than academia, and both the United States and China play a major role.

You can discover and compare the assistants from these players in our AI tools directory, for example ChatGPT, Claude, Gemini or Mistral Le Chat.

Frequently asked questions

How do you create an artificial intelligence?

You gather large amounts of data, train a neural network on it using computing power (GPUs), fine-tune it with human examples and then human feedback, evaluate it, and finally deploy it through an app or an API.

How much does it cost to train an AI?

It varies enormously: from a few thousand dollars for a small model to hundreds of millions for the largest. As a published order of magnitude, about $50,000 for GPT-2 (2019) and about $8 million for PaLM (2022). The real figures for recent models are rarely public.

Why are GPUs used for AI?

Because these processors, designed for graphics, can perform thousands of calculations at the same time. Training a neural network consists of a huge number of similar operations, which are perfectly suited to parallel computing.

Can you build your own AI?

Yes, on a small scale. You can train a small model on your own data, or more simply adapt an existing open model (fine-tuning) on a computer or a cloud service. Training a very large model from scratch remains reserved for organizations with considerable resources.

Does AI use a lot of electricity?

Collectively, yes: according to the IEA, data centers accounted for about 1.5% of the world’s electricity in 2024 (415 TWh) and could reach about 945 TWh in 2030. A single request uses little, but the total is high with millions of users.

Sources and references

  1. IEA — Energy and AI (2025)
  2. Stanford HAI — AI Index Report 2026
  3. Wikipedia — GPT-3
  4. Wikipedia — Large language model (training costs and scaling laws)
  5. Wikipedia — DeepSeek
  6. Hoffmann et al. (DeepMind), “Training Compute-Optimal Large Language Models” (Chinchilla), 2022
  7. Kaplan et al. (OpenAI), “Scaling Laws for Neural Language Models,” 2020

Independent editorial content. The facts, dates and figures cited rely on the sources listed at the end of the page; this content is for information only and does not constitute professional advice.