Generative Adversarial Networks (GAN) Taxonomy: Minimax Game, JS Divergence Flaw, WGAN Earth Mover Distance & WGAN-GP Guide
Summary: Generative Adversarial Networks (GANs) frame generative modeling as a two-player zero-sum game between a Generator and Discriminator. This 100% exhaustive guide covers Minimax game formulation, optimal discriminator D*(x) proof, JS divergence vanishing gradient flaws, WGAN Earth Mover distance derivations, WGAN-GP Gradient Penalty, Spectral Normalization, and Pure Numpy implementations with rich SEO explanatory text.
🧭 Knowledge Map & Architecture Graph
graph TD
subgraph A["1. Minimax Game & Optimal D*"]
A1["Objective: min_G max_D V(D, G)"]
A2["Optimal Discriminator D*(x) = p_data / (p_data + p_g)"]
A3["V(D*, G) = -2log2 + 2 · JSD(p_data || p_g)"]
A1 --> A2 --> A3
end
subgraph B["2. JS Divergence Gradient Flaw"]
B1["Disjoint Supports in High Dimensions"]
B2["JSD = log 2 (Constant) → Zero Gradients"]
B1 --> B2
end
subgraph C["3. WGAN & Gradient Penalty"]
C1["Wasserstein Distance W(p_r, p_g)"]
C2["WGAN-GP: λ E[(||∇_x̂ f(x̂)||_2 - 1)²] 1-Lipschitz Constraint"]
C1 --> C2
end
subgraph D["4. Mode Collapse & Stability"]
D1["Spectral Normalization (SN-GAN)"]
D2["Metrics: IS & FID (Fréchet Inception Distance)"]
D1 --> D2
end
A --> B --> C --> D
💡 Intuition: GAN training is a counterfeit-money game: the discriminator learns to be a good judge, the generator learns to forge. With the optimal judge $D^*(x) = p_{data}/(p_{data}+p_g)$, the objective becomes $-2log 2 + 2cdot JSD$ — and in high dimensions the real and generated images live on non-overlapping manifolds, so JSD is a constant $log 2$ with zero gradient: the generator gets no signal at all. Wasserstein distance fixes this by measuring “earth-mover cost” — it stays smooth and linear ($W = theta$) even for disjoint distributions — and WGAN-GP enforces the required 1-Lipschitz critic via the penalty $lambdamathbb{E}[(|nabla_{hat x} f|_2 – 1)^2]$ instead of fragile weight clipping.
🎤 Quick Answer: “KL diverges to $+infty$ and JS is stuck at $log 2$ (zero gradient) for two point-masses at 0 and $theta$, while $W = theta$ grows linearly — that’s why WGAN trains. Mode collapse = the generator only draws digit ‘1’; mitigation: WGAN-GP, spectral normalization, Unrolled GAN. FID (not eyes) is the metric: lower is better.”
📚 Chapter 1: Pure Numpy GAN Engine
Plain-language reading (full implementations in the zh version): minimax_loss translates the original GAN objective — the discriminator is penalized for both missing real images and accepting fakes; wgan_gp_critic_loss computes the Wasserstein estimate (mean fake score − mean real score) plus the gradient penalty that pulls $|nabla_{hat x} f|_2$ toward 1 on interpolated points $hat x = epsilon x + (1-epsilon)G(z)$.
import numpy as np
class PureNumpyGANEngine:
@staticmethod
def wgan_gp_critic_loss(f_real: np.ndarray, f_fake: np.ndarray, grad_hat: np.ndarray, lambda_gp: float = 10.0) -> float:
pass
💡 Intuition: The critic is a “scoring judge” without the final Sigmoid: it outputs an unconstrained score, and the score gap between real and fake approximates the Wasserstein distance — as long as the score function stays 1-Lipschitz.
🎤 Quick Answer: “WGAN loss = $mathbb{E}[f_{fake}] – mathbb{E}[f_{real}] + lambda,GP$ with $lambda=10$ default; the GP term is computed on random interpolations between real and fake samples so the Lipschitz constraint holds everywhere, not just at the samples.”
🧠 深入探索 TalentMe 全景技术图谱与备考路线
本文选自 TalentMe AI 技术专栏与高维职业罗盘。支持双模态 Obsidian 本地私域同步、艾宾浩斯智能复习与 IDE 内嵌 AI 导师模拟面试。