Generative AIGuide 7 of 9

Generative AI: How an AI Creates Images, Videos and Sound

9 min readUpdated on By the CheblAI editorial team
Read inFrançaisEnglish中文

In brief

  • Generative AI creates new content: text, images, voices, music, video and code. It learns the patterns in millions of examples, then generates plausible variations.
  • For images, the dominant technique is diffusion: the model learns to remove noise, then starts from random noise and gradually “denoises” it until it becomes an image guided by your description.
  • Synthetic content (“deepfakes”) raises questions about copyright, consent and disinformation, which European regulation is starting to address.
On this page
Generative AI: Producing Content Instead of Just Sorting ItThe First Generation: GANs (2014)Diffusion: The Art of Removing NoiseVoice, Music, Video: The Other FieldsTechnical Limits of Generated Images and VideoCopyright, Consent and DeepfakesFrequently asked questionsSources and references

Generative AI: Producing Content Instead of Just Sorting It

“Classic” AI (known as discriminative AI) answers questions like “Is this image a cat?” Generative AI does the opposite: “Draw me an astronaut cat.” It doesn’t copy an existing image. Instead, it has learned the statistical patterns of millions of images (shapes, textures, styles) and produces new ones that resemble them.

For text, this principle is described in how ChatGPT works. This article focuses on images, sound and video.

The First Generation: GANs (2014)

In 2014, researcher Ian Goodfellow and his colleagues proposed generative adversarial networks (GANs). The principle is a duel between two networks: a generator that produces fake images, and a discriminator that tries to tell real from fake. By competing against each other, the generator becomes able to produce highly realistic images.

GANs popularized synthetic faces of people who don’t exist. But they are hard to train and offer little control from a text description. They have largely been replaced by diffusion models.

Diffusion: The Art of Removing Noise

Diffusion models, popularized by research published around 2020, underpin DALL-E 2 (2022), Midjourney, Stable Diffusion (released openly in August 2022) and many video generators. Their idea is counterintuitive:

Diagram of the diffusion process: during training, an image gradually turns into noise, then the model starts from noise and recovers an image guided by text
Diagram 1 — Diffusion: learning to add noise, then to remove it in order to create.
  1. During training

    Take a real image and gradually add noise to it (like TV static) until it becomes completely random noise. The model learns to reverse each step: removing a little bit of noise.

  2. When you generate

    The model starts from pure random noise and “denoises” it step by step, guided by your text description (the prompt). An image gradually appears.

  3. The link with text

    A second model turns your description into numbers to steer the denoising. CLIP-type models, trained on millions of image-caption pairs, learn to match words with images.

Each generation starts from different noise, which is why the same prompt gives a different result every time.

Voice, Music, Video: The Other Fields

FieldHow it works (in brief)Example tools
ImageDiffusion guided by textMidjourney, Stable Diffusion, Adobe Firefly, Leonardo AI
Voice (speech synthesis)A model learns the characteristics of a voice (timbre, rhythm, intonation) so it can read any text aloudElevenLabs
MusicModels generate an audio signal from a text describing the style and lyricsSuno AI
VideoDiffusion applied to a sequence of images, with consistency over timeRunway Gen-4
Editing and subtitlesAutomatic transcription, then editing of the textDescript

Find these tools and their alternatives in the AI tools directory.

Technical Limits of Generated Images and Video

  • Inconsistent details: hands, written text, reflections and symmetry are classic weak points, even though recent models are improving fast.
  • Consistency over time: in video, keeping the same face, the same setting and the same physics throughout remains difficult.
  • Limited control: getting exactly the image you imagined often takes many attempts and a precise prompt.
  • Bias: a model reproduces the stereotypes found in its training data (for example, associating certain jobs with a gender).

Copyright, Consent and Deepfakes

Generative AI raises major legal and ethical questions, which are still being clarified:

  • Copyright: models were trained on huge collections of images and texts, often protected. Several lawsuits are under way, for example Getty Images against Stability AI (2023) and The New York Times against OpenAI (December 2023). Rules vary from country to country and are evolving.
  • Image rights and consent: cloning a person’s voice or face without their permission is a major legal and ethical problem.
  • Disinformation: deepfakes (realistic fake videos or voices) can be used to deceive, impersonate someone or harass them.
  • Transparency: the EU AI Act requires telling users when they are interacting with an AI and identifying synthetic content, notably deepfakes; these transparency obligations have applied since August 2, 2026, with a grace period until December 2, 2026 for marking content produced by systems already on the market (see limits, risks and regulation).

Frequently asked questions

How does AI create an image?

Most current generators use a diffusion model: it starts from random noise and gradually denoises it, step by step, guided by your text description, until it produces a coherent image.

Does an AI-generated image copy existing images?

The model doesn’t stitch together stored pieces of images: it generates pixels from patterns it has learned. It can, however, reproduce styles or, rarely, elements very close to its training data, which fuels the legal debates over copyright.

What is the difference between a GAN and a diffusion model?

A GAN pits two networks (a generator and a discriminator) against each other in a duel. A diffusion model learns to remove noise and generates by gradually denoising. Diffusion is more stable and more controllable today, and it dominates image generation.

Can I use an AI-generated image commercially?

It depends on the tool (its terms of use) and the country (the legal status of AI-generated works). Read the platform’s terms, and avoid any resemblance to protected works, brands or people.

What is a deepfake?

A deepfake is content (video, image or voice) generated or altered by AI to realistically imitate a real person. It can be used for creative work, but also for disinformation or impersonation.

Sources and references

  1. Goodfellow et al., “Generative Adversarial Nets,” 2014
  2. Ho, Jain & Abbeel, “Denoising Diffusion Probabilistic Models,” 2020
  3. Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models,” 2022
  4. Wikipedia — Stable Diffusion
  5. European Commission — AI Act

Independent editorial content. The facts, dates and figures cited rely on the sources listed at the end of the page; this content is for information only and does not constitute professional advice.