VAE Models & Architecture

Variational Autoencoder - a neural network that compresses images into latent space and reconstructs them back. It acts as the encoder/decoder for diffusion models.

The VAE (Variational Autoencoder) is responsible for converting between pixel space and latent space in Stable Diffusion and similar models.

Two Functions

Encoder

  • Takes a full-resolution image
  • Compresses it to latent representation
  • Used during img2img and inpainting

Decoder

  • Takes the latent representation
  • Reconstructs a full-resolution image
  • Used at the end of every generation

VAE Quality

Different VAEs produce different results:

  • Standard VAE: Default, balanced quality
  • EMA VAE: Slightly different color rendering
  • MSE VAE: Better for faces, sometimes duller colors
  • Custom VAEs: Fine-tuned for specific styles

Common Issues

  • Washed out colors: Try a different VAE
  • Blurry details: VAE struggling with fine features
  • Color shift: VAE mismatch with the model

Swapping VAEs

In most UIs, you can load a different VAE:

  1. Download a VAE file (.safetensors)
  2. Place in VAE folder
  3. Select in settings or workflow