U-Net Models & Architecture

The core neural network architecture in diffusion models that predicts and removes noise during image generation. Named for its U-shaped structure with encoder and decoder paths.

The U-Net is the heart of Stable Diffusion and similar models. It’s the component that actually “imagines” the image by learning to remove noise.

Architecture

The U-Net has a distinctive shape:

Input → Encoder → Bottleneck → Decoder → Output
         ↓                        ↑
         └──── Skip Connections ──┘

Encoder Path

  • Progressively compresses the input
  • Captures high-level features
  • Multiple downsampling blocks

Decoder Path

  • Progressively reconstructs detail
  • Uses skip connections from encoder
  • Multiple upsampling blocks

Skip Connections

  • Connect encoder layers to decoder layers
  • Preserve fine details during generation
  • The “U” in U-Net

In Diffusion Models

During generation:

  1. U-Net receives noisy latent + text embedding
  2. Predicts the noise component
  3. Noise is subtracted
  4. Repeat for N steps

Model Size

U-Net size determines model capabilities:

  • SD 1.5: ~860M parameters
  • SDXL: ~2.6B parameters
  • SD 3: Uses transformer instead