Key Concepts
- Prompt Adherence: The degree to which a generated image accurately reflects the input text prompt.
- Aesthetics: The visual quality, physical plausibility, and artistic appeal of an image.
- FID (Fréchet Inception Distance): A metric comparing the distribution of real images to generated images in a latent space.
- ELO Rating: A system for ranking models based on pairwise comparisons, accounting for the relative strength of opponents.
- CLIP Score: A metric using contrastive language-image pre-training to measure semantic alignment between text and images.
- SSIM (Structural Similarity Index Measure): A metric evaluating the similarity of two images based on luminance, contrast, and structure.
- LPIPS (Learned Perceptual Image Patch Similarity): A deep-learning-based metric that measures perceptual similarity using pre-trained encoders.
- VLM/MLM as a Judge: Using Multimodal Large Language Models to evaluate images by providing reasoning and scores based on specific rubrics.
1. Evaluation Frameworks
The lecture categorizes evaluation into two primary buckets: Aesthetics (physical plausibility, realism) and Prompt Adherence (object presence, style, specific details).
Human-Based Evaluation
- Absolute Scale (1-5): Nuanced but prone to high variance and subjective noise.
- Binary (Good/Bad): Easier for humans, but still suffers from lack of absolute reference.
- Pairwise Comparison: The most robust human method. It reduces cognitive load by asking which of two images is better.
- ELO Rating: Used to rank models. It updates a model's score based on the "surprise" factor—winning against a strong opponent yields more points than winning against a weak one.
2. Automated Metrics (Reference-Free vs. Reference-Based)
Reference-Free Metrics (Distribution-Based)
- FID: Measures the distance between the distribution of real images and generated images using an Inception network encoder. It assumes Gaussian distributions to provide a closed-form solution for the Wasserstein distance.
- Limitation: Not perfectly reflective of human quality; sensitive to the choice of the reference dataset (e.g., 50k images).
Reference-Based Metrics (Reconstruction-Based)
Used primarily for VAEs or image-to-image tasks where a ground truth exists.
- MSE (Mean Squared Error): Pixel-wise distance. Highly sensitive to minor spatial shifts.
- PSNR (Peak Signal-to-Noise Ratio): Normalizes MSE with a logarithmic scale to better reflect human perception of light/error.
- SSIM: Compares patches based on luminance, contrast, and structure. It is more robust than MSE because it considers relative values rather than absolute pixel differences.
- LPIPS: Uses pre-trained encoders (e.g., VGG, AlexNet) to compare feature maps, aligning closely with human perceptual judgment.
3. Multimodal LLMs (MLM) as Judges
Modern evaluation shifts toward using MLMs to act as "judges" that provide both scores and rationales.
- Architecture: Modern MLMs (e.g., LLaVA) feed image and text tokens directly into a decoder-only structure, avoiding the need for complex cross-attention layers.
- Methodologies:
- TIFA (Text-to-Image Faithfulness Evaluation): Decomposes prompts into atomic yes/no questions to pinpoint specific failures.
- VQA Score: Uses the model's next-token probability distribution to answer "Does this image show [prompt]?"
- VIE Score (Visual Instruction Guided Explainable Score): Uses a rubric-based approach where the model outputs a rationale before assigning a score for semantic consistency and perceptual quality.
4. Best Practices for Evaluation
- Chain of Thought: Always prompt the judge model to output a rationale before the score to improve accuracy.
- Determinism: Set temperature to zero for evaluation tasks to ensure reproducibility.
- Position Bias: In pairwise comparisons, swap the order of images to prevent bias toward the first or second position.
- Alignment: Ensure the judge model's rubric is calibrated against human preferences before trusting it for large-scale automated evaluation.
5. Notable Benchmarks
- GenEval: Tests increasing difficulty in object counting, color, and spatial positioning.
- DPGB: Uses a logical graph of claims to evaluate dense, complex prompts in a coarse-to-fine manner.
- Long Text Bench: Focuses on OCR capabilities and text rendering within images.
- Grounded Edit Bench: Evaluates image-to-image editing tasks (e.g., background replacement).
Conclusion
Evaluation is a multi-faceted problem. While quantitative metrics like FID and SSIM provide a baseline, they are often insufficient for capturing semantic nuance. The current industry trend is moving toward MLM-as-a-judge frameworks, which offer interpretability and reasoning. However, these require careful calibration against human intuition to avoid "black box" errors. Ultimately, no single metric is perfect, and a combination of automated benchmarks and human-aligned model judging is recommended.
AI summaries can miss context or contain errors. Check important details against the original video.