Latent Space Models & Architecture

A compressed mathematical representation where AI models process images more efficiently than working with raw pixels. Each point in latent space corresponds to a possible image.

Latent space is where the magic of AI image generation happens. Instead of working directly with millions of pixels, the model operates in this compressed representation.

Why Latent Space?

A 512x512 image has 786,432 pixels (times 3 for RGB = 2.3 million values). Working in latent space reduces this to roughly 64x64x4 = 16,384 values - a 140x reduction.

How It Works

  1. Encoding: The VAE compresses your image into latent space
  2. Processing: The U-Net works in this compressed space
  3. Decoding: The VAE expands the result back to a full image

Practical Implications

  • Generation is much faster than pixel-space models
  • Some fine details may be lost in compression
  • Latent space dimensions affect output resolution
  • SDXL uses a larger latent space than SD 1.5