Junoh Kang

Back

Slides PDF Open slides ↗

Overview#

DDPM[1] and Score-Based Model[2] introduce diffusion model as a new paradigm of generative models. Since the concepts of both papers are similar, one might regard Score-Based Model[2] as only a continuous version of DDPM[1]. Two papers have slight different views, even their loss functions and implementations are similar.

This post mainly explains how formulations and objectives of two papers are different, and how they are related even with the differences.

Summary of the post#

  1. The objective of DDPM[1] is to minimize the surrogate of the negative log-likelihood.
  2. The objective of Score-Based Model[2] is to match marginal distributions of forward SDE and backward SDE/ODE.
  3. Even with the differences, both derivations require the score function, differential of the log of the probability density function. The score functions are parametrized by neural networks and both papaers have similar loss functions.

Maximizing Log-Likelihood#

Forward (Diffusion) Process#

The forward process is a Markov chain that gradually adds Gaussian noise to the data for steps with distributions defined as follows:

where are pre-defined constants.

Backward (Denoising) Process#

The backward process is a Markov chain that gradually denoises perturbed data and it is parametrized by neural networks. When the backward distribution can be approximated as

proof.

It is reasonable to parametrize the denoising distribution as Gaussian as long as are infinitesimal. Therefore the bacward process is defined as follows:

Note that we expect to learn .

Minimizing Surrogate of Negative Log-Likelihood#

The negative log-likelihood of data is

proof.

Using Markov properties,

proof.

Therefore, the surrogate of negative log-likelihood becomes

The surrogate of negative log-likelihood can be explictly expressed using

where .

Finally, the objective function becomes

where are some constants.


Matching Marginal Distributions#

Forward SDE#

For pre-defined function and , a forward SDE perturbs the data with Gaussian noise by

where is Brownian process.

If is a solution of the forward SDE, it can be treated as a sample from the joint distribution . However, learning joint distribution is difficult and our interest is only , not . Therefore, it suffices to consider weakened objective, learning how marginal distributions evolve as changes. The evolution of the marginal distributions is goverened by the Fokker-Plank equation:

Backward SDE/ODE#

Following backward SDE and ODE are known to have the same marginal distributions:

where is the reverse-time Brownian motion.

Since and are known, the only unknown component in backward SDE/ODE is which is also known as a score function. The score function is parametrized by neural network, .

Learning Score Function#

Since we parametrized the score function with the neural network, we can consider a loss function of

where are some constants. Note that is intractable and with some tricks, the loss function changes into tractable form:

References#

  1. Jonathan Ho et al., Denoising diffusion probabilistic models (NeurIPS, 2020)

  2. Yang Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (ICLR, 2021)

Search