EE 641 - Unit 3A
Dr. Brandon Franzke
Fall 2026
GAN Architectures
Evaluation
Variational Autoencoders
Comparing the Families
[Score Matching] Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Advances in Neural Information Processing Systems, 2019, pp. 11918–11930.
[GAN] I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems, 2014, pp. 2672–2680.
[GAN Review] I. Goodfellow, “NIPS 2016 tutorial: Generative adversarial networks,” arXiv preprint arXiv:1701.00160, 2016.
[GAN Theory] S. Arora and Y. Zhang, “Do GANs actually learn the distribution? An empirical study,” arXiv preprint arXiv:1706.08224, 2017.
[DCGAN] A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” in International Conference on Learning Representations, 2016.
[WGAN] M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in International Conference on Machine Learning, 2017, pp. 214–223.
[WGAN-GP] I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of Wasserstein GANs,” in Advances in Neural Information Processing Systems, 2017, pp. 5767–5777.
[StyleGAN] T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4401–4410.
[Evaluation] T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila, “Improved precision and recall metric for assessing generative models,” in Advances in Neural Information Processing Systems, 2019, pp. 3927–3936.
[VAE] D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in International Conference on Learning Representations, 2014.
[VAE Review] D. P. Kingma and M. Welling, “An introduction to variational autoencoders,” Foundations and Trends in Machine Learning, vol. 12, no. 4, pp. 307–392, 2019.
[ELBO] M. D. Hoffman and M. J. Johnson, “ELBO surgery: yet another way to carve up the variational evidence lower bound,” in NIPS Workshop on Advances in Approximate Bayesian Inference, 2016.
[VQ-VAE] A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Advances in Neural Information Processing Systems, 2017, pp. 6306–6315.
Two modeling targets
What each must keep
Why it is harder
Both panels come from one fit - the boundary falls out of the density, never the reverse.
The output of a trained model: samples
The target is a distribution, seen only through samples
The consequence for training
Training and evaluation both compare distributions - the reason divergences run through everything here.
Shannon’s link
Worked exactly (the figure)
Bits per dimension
Representing \(p(\mathbf{x})\) requires a functional form - the energy is the most general choice.
Energy scores configurations
The form follows from maximum entropy
No generality lost
Mass concentrates in low-energy basins - deeper basin, more mass.
\(T\) rescales energy differences - minima unmoved
The score needs no normalization
Models fix \(T = 1\) - sampling reintroduces temperature as a control.
What requires \(Z\)
What survives without \(Z\)
Counting the terms (binary image, 224 × 224 × 3)
Every EBM training method is a strategy for never evaluating this integral.
Differentiate once
What the identity costs
Estimating \(\partial F / \partial \boldsymbol{\theta}\) means sampling from the model - the requirement that shapes all EBM training.
Data \(\{\mathbf{x}_1, ..., \mathbf{x}_N\} \sim p_{\text{data}}\):
The same objective as a divergence
What forward KL demands
Which divergence a method minimizes determines which errors it tolerates.
The objective is a gap
Positive phase - push down at data
Negative phase - push up where the model puts mass
Why the \(F\) term must be there
Sampling from the model is the entire difficulty of likelihood training for EBMs.
At any maximum of the likelihood the two phases balance:
Moment matching
Richer statistics need richer energies
The energy family fixes the statistics the model can represent - the rest is invisible to training.
Two different asks of the same landscape
The mode is not where the mass is
One walk, two algorithms
An optimizer answers where probability peaks - a sampler answers where probability lives.
Monte Carlo
Producing \(\mathbf{x}^{(k)} \sim p\)
The expectation is easy given samples - producing the samples is the problem.
Build the chain around the target
Detailed balance
Metropolis-Hastings
Named kernels
Detailed balance guarantees the target distribution - nothing bounds the time to reach it.
Burn-in
Mixing time
Autocorrelation
In practice
\(K\) chain states are not \(K\) samples - budget by \(K_{\text{eff}}\).
The kernel deep EBMs use in practice - proposals follow the energy gradient
Langevin dynamics
Hamiltonian Monte Carlo
Step size trades discretization error against exploration - no setting fixes separated modes.
Escaping a basin is a rare event
High dimension compounds it
Spectral view
At image scale the chain never reaches the distribution the gradient formula assumes.
One sampler run inside every update
Published training budgets
Why no budget closes the gap
Truncation is not a shortcut to the exact gradient - it is a different, biased estimator.
Four approaches, each giving up something exact:
Truncate the chain - contrastive divergence
Match the score - score matching, diffusion models
Bound the likelihood - variational autoencoders
Drop the density - generative adversarial networks
Each trades the exact gradient for a computable one.
Boltzmann machine (Ackley, Hinton, Sejnowski 1985)
The restriction
What a deleted edge means
Deleting edges is what turns unit-at-a-time Gibbs into two block draws.
What the units are
The data model is the marginal
Role
Hidden units are a first appearance of latent variables - VAEs make them continuous and learn the inference.
Bipartite energy
Fix \(\mathbf{v}\) - the energy regroups
Conditionals
Block Gibbs sampling
Tractable conditionals - the negative phase still needs the chain to mix.
Why hidden units at all - measure against the no-hidden baseline:
Baseline: visible-only pairwise energy
Integrating out the hidden units
Marginalized hidden units produce higher-order visible interactions - without higher-order terms in \(E\).
CD-k
Gradient estimate
Chains start at the data - the estimate stays near the data distribution.
Objective (Hinton 2002)
Why the gradient is tractable
Consequences
A biased estimator, cheap enough to make RBM training practical.
Fit the score, not the density
\(\nabla_{\mathbf{x}} \log p_{\text{data}}\) is unknown
What the trace term costs
Sampling is gone - the trace term takes its place as the obstacle.
Perturb, then regress
Equivalence (Vincent 2011)
Denoising autoencoder connection
Basis of diffusion models
The trace term is gone too - denoising regression needs only forward passes on noisy data.
Energy parameterizations
Training loop (sketch)
# inner loop: short-run Langevin from the replay buffer x = buffer.sample() for t in range(60): x = x - eps/2 * grad_x(E(x)) + sqrt(eps) * randn() # outer loop: two-phase update g = grad_E(x_data).mean() - grad_E(x).mean() theta = theta - lr * g
Budget (Du and Mordatch 2019)
Persistent problems
A network energy is learned statistics - \(T(\mathbf{x})\) with millions of parameters.
Fit the field - denoising score matching
Follow the field - short-run Langevin, network energy
Same field - one method regresses onto it, the other walks it at every step.
Cost per update, at the published truncation
Memory
Beyond computation
Truncated training runs ~60× a classifier’s cost - unbiased training has no finite cost at all.
Score matching → diffusion models
Contrastive ideas → self-supervised learning
Around the intractability → the other families
The parts that survived are the parts that never needed \(Z\).
EBM: model \(p(\mathbf{x})\) explicitly
GAN: model the sampler directly
What replaces the negative phase
The trade
Sampling cost becomes a forward pass - the cost reappears as training instability.
thispersondoesnotexist.com
{width=90%}
Samples: thispersondoesnotexist.com (StyleGAN2, Karras et al. 2020)
Peak numbers (FFHQ faces, 1024 × 1024)
{width=100%}
Published GAN outputs across application domains
In 2026
Photograph-quality samples from one network pass.
Architecture (Goodfellow et al. 2014)
Reported results
What was new
{width=100%}
Original samples, MNIST and TFD (Goodfellow et al. 2014)
Direction matters
Forward: \(\text{KL}(p_{\text{data}} \| q)\) - the maximum-likelihood direction
Reverse: \(\text{KL}(q \| p_{\text{data}})\)
Reverse KL drops modes cheaply - the direction returns with the generator’s behavior.
No density → no likelihood objective - training becomes a game against a learned opponent
Zero-sum game
Best response
Minimax
Nash equilibrium
For GANs
Training = seeking a saddle point, not a minimum.
\(D\): an ordinary binary classifier
\(G\): trained through the classifier
Alternation
\(D\) is trained as an ordinary classifier - \(G\) is trained through it.
Fix \(G\), maximize over \(D\) pointwise
Reading the result
The optimal classifier’s confidence encodes where the two distributions disagree.
Substitute \(D^*\) into the value
Jensen-Shannon divergence
Saturation
With a perfect critic, adversarial training is JS minimization.
With separated modes, each KL reduces to a KL between the mixture weights
Removing a data mode - \(p_{\boldsymbol{\theta}}\)'s weight on one data mode \(\to 0\)
Adding a spurious mode - \(p_{\boldsymbol{\theta}}\) mass where \(p_{\text{data}}\) has none
Objective vs behavior
Forward KL penalizes missing modes, reverse KL penalizes spurious mass - JS bounds both.
Early training: \(G\) poor → \(D\) separates easily → \(D(G(\mathbf{z})) \to 0\)
Differentiate in logit space - \(D = \sigma(a)\), gradients reach \(G\) through \(a\)
Same fixed points, opposite failure
Under the saturating loss, the gradient to \(G\) is weakest exactly when \(G\) is worst.
Result (Arjovsky and Bottou 2017) - at optimal \(D\), the non-saturating generator update follows
Reading the right side
Consequences
The vanishing-gradient fix chose the mode-seeking divergence.
Reverse KL, applied
Two named failures
The cycle in practice
Collapse follows the objective - not an optimizer failure.
Equilibrium exists - in function space
Simultaneous updates on the simplest saddle - \(V(g, d) = g \cdot d\)
Observed GAN training
Update ratios
The saddle exists - simultaneous gradient updates do not converge to it by default.
Reference map for the section - every technique traces to a failure the dynamics predict:
| Failure | Mechanism | Treatments |
|---|---|---|
| Vanishing G gradient | D saturates, JS plateau | non-saturating loss · spectral norm · Wasserstein objective |
| Mode collapse / dropping | reverse-KL geometry | batch diversity features · unrolled D · bigger batches |
| Oscillation | simultaneous updates orbit | TTUR · EMA of G weights · lower rates |
| D overconfidence | classifier outpaces G | one-sided label smoothing · augmentation |
| Exploding D slopes | unconstrained critic | spectral norm · gradient penalties |
Targets (Salimans et al. 2016)
real_targets = 0.9 # smoothed fake_targets = 0.0 # never smoothed
Why the fake side stays at zero
Fake-side smoothing rewards exactly the samples training should reject.
Separate learning rates (Heusel et al. 2017)
Optimizer settings in practice
What the asymmetry approximates
The faster \(D\) time scale stands in for the optimal-\(D\) assumption.
One line per layer (Miyato et al. 2018)
Placement
Measured (CIFAR-10 unconditional, Inception Score)
Why a slope cap is the right constraint: the Wasserstein distance supplies the reason.
Generator (Radford et al. 2016)
Discriminator
Initialization
nn.init.kaiming_normal_(conv.weight) # He: ReLU stacks nn.init.xavier_normal_(deconv.weight) # Xavier: tanh output end
Normalization that shares statistics across the batch breaks per-sample discrimination.
Reference table - what to measure, what it means, what to change:
| Failure | Signature | Response |
|---|---|---|
| Mode collapse | sample diversity ↓ · \(D\) loss → 0 | larger batches · minibatch std-dev feature (Salimans et al. 2016; Karras et al. 2018) · unrolled \(D\) |
| Vanishing \(G\) gradient | \(|\nabla_G| \to 0\) while \(D\) accuracy → 1 | non-saturating loss · spectral norm · Wasserstein objective |
| Oscillation | loss variance grows · samples cycle between modes | lower rates · EMA of \(G\) · TTUR |
| \(D\) memorization | \(D\) train accuracy 1.0, large gap to held-out data | augmentation · one-sided smoothing |
Loss curves diagnose the game - sample quality needs its own measurements.
Support - where a distribution puts its mass
Image data is thin in pixel space
Generated data is thin by construction
Two thin sets rarely meet (Arjovsky and Bottou 2017)
The JS plateau is generic for image models - the objective needs replacing, not the optimizer.
Reading the definition
Two point masses, separation \(\alpha\)
What that buys the generator - a slope toward the data even from far away
W varies smoothly with separation - JS stops at \(\log 2\).
Kantorovich-Rubinstein duality
Lipschitz constraint - \(\|f\|_L \leq 1\):
What duality changes
Why the cap is essential
Remove the slope cap and the supremum diverges - the constraint is the metric.
Changes from the classifier game
Training loop
for step in range(iters): for _ in range(n_critic): # typically 5 d_loss = -D(x_real).mean() + D(G(z)).mean() update(D, d_loss) enforce_lipschitz(D) # clip or penalty g_loss = -D(G(z)).mean() update(G, g_loss)
What the loss now means
What stays
The critic loss estimates \(W\) - a curve that tracks sample quality.
Original WGAN enforcement
for p in D.parameters(): p.data.clamp_(-c, c) # c = 0.01
Three failures (Gulrajani et al. 2017)
What the clip actually bounds
Clipping satisfies the constraint at the cost of the function class.
WGAN-GP objective (Gulrajani et al. 2017)
Why target \(\|\nabla D\| = 1\)
Cost and settings
eps = torch.rand(b, 1, 1, 1) x_hat = eps * x_real + (1 - eps) * x_fake x_hat.requires_grad_(True) d_hat = D(x_hat) grad = autograd.grad(d_hat.sum(), x_hat, create_graph=True)[0] gnorm = grad.view(b, -1).norm(2, dim=1) gp = lam * ((gnorm - 1) ** 2).mean()
Against clipping
Constrain the slope directly, keep the weights free - the standard Wasserstein implementation.