U-Net Models & Architecture
The core neural network architecture in diffusion models that predicts and removes noise during image generation. Named for its U-shaped structure with encoder and decoder paths.
The U-Net is the heart of Stable Diffusion and similar models. It’s the component that actually “imagines” the image by learning to remove noise.
Architecture
The U-Net has a distinctive shape:
Input → Encoder → Bottleneck → Decoder → Output
↓ ↑
└──── Skip Connections ──┘
Encoder Path
- Progressively compresses the input
- Captures high-level features
- Multiple downsampling blocks
Decoder Path
- Progressively reconstructs detail
- Uses skip connections from encoder
- Multiple upsampling blocks
Skip Connections
- Connect encoder layers to decoder layers
- Preserve fine details during generation
- The “U” in U-Net
In Diffusion Models
During generation:
- U-Net receives noisy latent + text embedding
- Predicts the noise component
- Noise is subtracted
- Repeat for N steps
Model Size
U-Net size determines model capabilities:
- SD 1.5: ~860M parameters
- SDXL: ~2.6B parameters
- SD 3: Uses transformer instead