Diffusion Model Models & Architecture

A type of generative AI model that creates images by learning to reverse a gradual noising process, transforming random noise into coherent images guided by text prompts.

Diffusion models are the core technology behind most modern AI image generators including Stable Diffusion, DALL-E, and Midjourney.

How It Works

The process has two phases:

Training (Forward Diffusion)

  1. Start with real images
  2. Gradually add noise over many steps until the image becomes pure noise
  3. The model learns to predict the noise at each step

Generation (Reverse Diffusion)

  1. Start with pure random noise
  2. Predict and remove noise step by step
  3. Guide the denoising with text embeddings (from CLIP)
  4. Result: a coherent image matching the prompt

Why Diffusion Works

Unlike earlier approaches (GANs), diffusion models:

  • Train more stably
  • Produce more diverse outputs
  • Handle complex scenes better
  • Scale well with compute

Key Parameters

  • Sampling Steps: More steps = higher quality but slower
  • Sampler: Algorithm used for denoising (Euler, DPM++, etc.)
  • CFG Scale: How closely to follow the prompt

Most generation happens in “latent space” (a compressed representation) rather than pixel space, making it much faster.