Designing Diffusion Models in Real World
Elucidating the Design Space of Diffusion-Based Generative Models, focusing on the rationale behind its engineering details.
Overview#
In deep learning, practical implementations are just as important as theoretical supports. Especially when proposing new paradigms, such as GANs, diffusion models, and transformers, etc., engineering skills are essential to bring the paradigms into the real world. Even if the suggested designs on diffusion models in Elucidating the Design Space of Diffusion-Based Generative Models (EDM)[1] are not optimal, the choices of the designs are theoretically or empirically supported. Learning these reasonings may help to bring your theory into the real world.
Revisit Diffusion Models#
Reformulate diffusion models#
Score-Based Model[2] defines forward SDE and marginal distribution is calculated from forward SDE. However when it comes to training, marginal distribution is more important than the SDE and therefore EDM[1] defines the marginal distribution first.
where
Then, the corresponding probability flow ODE is
Obstacles in diffusion models#
Generation by diffusion models can interpreted as solving ODE:
-
is not known and it is parametrized by a network . The inaccurate approximation on the target causes degradation. \ → Better training! -
The solution at
given boundary condition at is \
The integral is numerically calculated, which causes truncation errors. \ → Reduce truncation errors, focus on important region!
Design space of diffusion models#
Components regarding training#
- Parametrization and network preconditioning:
, , , . - Loss weighting:
. - Noise level distribution for training:
. - Augmentation
Components regarding deterministic sampling#
- Truncation-error-reducing ODE:
, . - Truncation-error-reducing algorithms: Higher-roder integrators
- Distributing truncation errors properly: Discretization
Components regarding stochastic sampling#
- Rate of replaced noises:
- Heuristics:
.
Improvements to Training#
For this section, this post assumes
Parametrization, network preconditioning, loss weighting#
Then, the relation between a score function and the ideal denoiser is
Networks in many baselines predicts either
![]() |
|---|
| Loss comparison between predicting the denoised output or the added noise. |
From the observation, EDM[1] designs the network to predict
where
1.Network inputs should have bounded range#
2.Effective training target should have bounded range#
3.Errors of network should not be amplified#
4. Effecitve weight should be uniform#
Putting 1 ~ 4 together, the expected value of the loss at each noise level is 1. Moreover, the change of effective training target according to
Noise level distribution for training#
![]() |
|---|
| Observed loss per noise level. The shaded regions represent the standard deviation over 10k random samples. EDM’s proposed training sample density is shown by the dashed red curve. |
At low noise levels, seperating the small noise components is difficult and irrelevant, whereas at high noise levels, the correct answer approaches to dataset average; EDM[1] focuses on middle range noise levels for training:
Augmentataion#
EDM[1] follows the augmentation pipiline from the GAN literature[4].

- Each agmentation is enabled with
. - Draw
from each enabled augmentation and construct transformation matrix. - Pass data through
supersampled high-quality Wavelet filters. - Construct a 9-dimensional conditioning input vector for non-leaking augmentation. This vector makes the network to perform auxiliarty tasks.
Improvements to Deterministic Sampling#
Higher-order integrators#
For
Euler method ↗#
Euler method approximates the integral by
Therefore, the total truncation error is
Heun’s method ↗#
Then, Heun’s method approximates the integral by
Therefore, the total truncation error is
Deterministic sampling algorithm for EDM[1]#
Discretization#
As long as using numerical integrators with limited computational resources, truncation errors are inevitable. In terms of obtaining ODE trajectories accurately, it is important to minimize total truncation erros. Hoever, the interests of diffusion models at generation are only the solutions at low noise levels; it is reasonable to focus on low noise levels. EDM discretizes as
Increasing
![]() |
|---|
| (a),(b) Local truncation error at different noise levels. (c) FID as a function of |
Truncation-error-reducing ODE#
Many integrators including Euler and Heun’s method have small truncation errors if
![]() |
|---|
| A sketch of ODE curvature in 1D where |
Results of deterministic sampling#
- Config B changes basic hyperparameters such as batch size, learning rate, dropout, *etc*; it disable gradient clipping
- Config C improves the expressive power of the model.
- Configs D, E, and F are explained in the previous context.
Stochastic Sampling#
SDE formulation#
EDM[1] reformulates forward and backward SDE as a sum of the probability flow ODE and a varying-rate Langevin diffusion ↗ SDE:
Role of stochasticity#
In theory, ODE and SDE have the same marginal distributions. However in practice, stochasticity in sampling often enhances the sample quality. The authors attribute the beneficial role of stochasticity to the following steps:
deviates from the ideal marginal distribution, due to the training and truncation errors.- The Langevin diffusion drives the sample towards the ideal marginal distribution.
Stochastic sampling algorithm in EDM#
Stochastic sampling algorithm in EDM[1] is executed in two steps:
- Noise injection: integrate noise into samples according to
. - Noise decay with probability flow: solve the ODE from increased noise level to desired level.
Algorithm in real world#
![]() |
|---|
| Observe the effect of Langevin diffusion in real world: there is gradual image degradation with the repeated addition and removal of noise. A random image is drawn from |
Langevin diffusion is supposed to drive the sample towards the true data distribution, however…
- For low noise levels, images drift toward oversaturated colors.
- For high noise levels, images become abstract when
.
Authors suspect that non-conservative vector field generated by parametrized denoiser violates the premises of Langevin diffusion since their analytical denoisers have not shown such degradation.
-
For low noise levels, images drift toward oversaturated colors.\
Enable stochasticity within . -
For high noise levels, images become abstract when
. \ removes too much noise because of regression towards the mean ↗, which often happens when trained.\ Inflate the standard deviation of newly added noise: . -
New noise never exceeds the noise already in the image. \
Clamp . -
Controls the overal stochasticity by
.
Results of stochastic sampling#
![]() |
|---|
| Evaluation of stochastic samplers with ablations. Red line is deterministic sampler while purple line is optimal stochastic sampler. |
References#
-
Yang Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (ICLR, 2021)
-
Yaniv Benny and Lior Wolf, Dynamic Dual-Output Diffusion Models (CVPR, 2022)
-
Tero Karras et al., Training generative adversarial networks with limited data (NeurIPS, 2020)





