Understanding Diffusion Models in Two Perspectives
DDPM and score-based SDEs as two routes to reverse diffusion, contrasting their shared structure and modeling differences.
Overview#
DDPM[1] and Score-Based Model[2] introduce diffusion model as a new paradigm of generative models. Since the concepts of both papers are similar, one might regard Score-Based Model[2] as only a continuous version of DDPM[1]. Two papers have slight different views, even their loss functions and implementations are similar.
This post mainly explains how formulations and objectives of two papers are different, and how they are related even with the differences.
Summary of the post#
- The objective of DDPM[1] is to minimize the surrogate of the negative log-likelihood.
- The objective of Score-Based Model[2] is to match marginal distributions of forward SDE and backward SDE/ODE.
- Even with the differences, both derivations require the score function, differential of the log of the probability density function. The score functions are parametrized by neural networks and both papaers have similar loss functions.
Maximizing Log-Likelihood#
Forward (Diffusion) Process#
The forward process is a Markov chain that gradually adds Gaussian noise to the data for
where
Backward (Denoising) Process#
The backward process is a Markov chain that gradually denoises perturbed data and it is parametrized by neural networks.
When
proof.
It is reasonable to parametrize the denoising distribution as Gaussian as long as
Note that we expect
Minimizing Surrogate of Negative Log-Likelihood#
The negative log-likelihood of data is
proof.
Using Markov properties,
proof.
Therefore, the surrogate of negative log-likelihood becomes
The surrogate of negative log-likelihood can be explictly expressed using
where
Finally, the objective function becomes
where
Matching Marginal Distributions#
Forward SDE#
For pre-defined function
where
If
Backward SDE/ODE#
Following backward SDE and ODE are known to have the same marginal distributions:
where
Since
Learning Score Function#
Since we parametrized the score function with the neural network, we can consider a loss function of
where
References#
-
Jonathan Ho et al., Denoising diffusion probabilistic models (NeurIPS, 2020)
-
Yang Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (ICLR, 2021)