CLIP Models & Architecture

Contrastive Language-Image Pre-training - an AI model that understands relationships between text and images. Used to guide image generation based on text prompts.

CLIP is the bridge between your text prompt and the image being generated. It translates words into a format the image model can understand.

How CLIP Works

CLIP was trained on 400 million image-text pairs to learn relationships:

  1. Text Encoder: Converts your prompt to embeddings
  2. Image Encoder: Converts images to embeddings
  3. Matching: Both share the same embedding space

Role in Image Generation

When you write a prompt:

  1. CLIP tokenizes your text (breaks into tokens)
  2. Converts tokens to embeddings
  3. These guide the U-Net during denoising
  4. Higher CFG = stronger CLIP influence

Token Limits

CLIP has a maximum token limit:

  • SD 1.5: 77 tokens
  • SDXL: 77 tokens (but two text encoders)
  • Some UIs support token extension

CLIP Skip

Some models work better skipping the final CLIP layers:

  • CLIP Skip 1: Use all layers (default)
  • CLIP Skip 2: Skip last layer (common for anime)
  • Higher values = more abstract interpretation

Practical Tips

  • Front-load important words (processed first)
  • CLIP understands concepts, not just keywords
  • Artist names and style terms work because CLIP learned them