<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet href="/scripts/pretty-feed-v3.xsl" type="text/xsl"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:h="http://www.w3.org/TR/html4/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Junoh Kang</title><description>AI researcher working on generative models for images, video, and financial markets.</description><link>https://junoh-kang.github.io</link><item><title>Working Effectively with AI Agents</title><link>https://junoh-kang.github.io/blog/working-effectively-with-ai-agents</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/working-effectively-with-ai-agents</guid><description>More verified, useful outcomes—for less total human attention.</description><pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Summary&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Remember what productivity means.&lt;/strong&gt; It is not more output, but more verified outcomes with less human attention.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stay attached to the Plan → Do → Verify loop.&lt;/strong&gt; Your mental model must evolve with the work; otherwise, you cannot plan or verify what comes next.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Remember the limits of being human.&lt;/strong&gt; Your bio-token budget is finite, cannot be parallelized, and is replenished only by time.&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>LOB-Bench: Orderbook Generative Model Evaluation Benchmark</title><link>https://junoh-kang.github.io/blog/lob-bench-orderbook-generative-model-evaluation-benchmark</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/lob-bench-orderbook-generative-model-evaluation-benchmark</guid><description>A review of LOB-Bench, focusing on why generated order book models need rollout-level evaluation beyond one-step loss and stylized facts.</description><pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;LOB-Bench starts from a critique of how generative limit order book models are usually evaluated. One-step prediction loss and a few stylized facts can be useful local checks, but they do not show whether a model remains realistic after it samples a trajectory. A model can pass those narrower checks and still drift away from realistic market behavior once it generates a sequence autoregressively.&lt;/p&gt;
&lt;p&gt;The benchmark therefore evaluates generated LOBSTER-compatible order book rollouts by comparing real and generated score distributions, conditional behavior, market-impact response curves, and discriminator separability. This makes it useful for diagnosing where a generated market trajectory fails: quote state, event timing, cancellation behavior, event placement, order flow, price response, or trajectory-level artifacts.&lt;/p&gt;
&lt;p&gt;The reported results make the benchmark&apos;s role concrete. LOBS5 is the strongest tested model overall, but LOB-Bench still exposes growing horizon error, conditional microstructure failures, market-impact response gaps, and discriminator separability. The useful output is therefore not a single leaderboard score, but a failure profile for generated order book trajectories.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;LOB Generation Needs a Rollout Benchmark&lt;/h2&gt;
&lt;p&gt;Limit order book data is hard to generate because it mixes discrete events, continuous prices and quantities, irregular timing, strategic interaction, and market microstructure constraints. A generator that emits message-level data must keep the book coherent over time. A generator that emits book states directly must still preserve spread, depth, liquidity, return, and order-flow behavior across a rollout.&lt;/p&gt;
&lt;p&gt;The evaluation problem is that common checks answer narrower questions. Held-out next-token cross-entropy tests whether the model predicts the next event under real history, but it does not test the distribution of sampled trajectories. Stylized facts test selected marginal patterns, but they can miss failures in timing, cancellation behavior, event placement, or response dynamics. Fragmented realism metrics can show isolated failures, but they do not by themselves produce a benchmark-level failure profile.&lt;/p&gt;
&lt;p&gt;LOB-Bench is the paper&apos;s answer to that gap. It evaluates generated samples after rollout, when the conditioning history has started to contain model-generated outputs. The target question is not just &quot;did the model predict the next event?&quot; The target question is &quot;does the generated market trajectory still look like real LOB data after the model runs?&quot;&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;What LOB-Bench Measures&lt;/h2&gt;
&lt;p&gt;LOB-Bench starts from generated, real, and conditioning sequences stored in LOBSTER-compatible CSV format. This input contract matters. If a model generates messages, those messages need to be exported or replayed into comparable message and book-state sequences. If a model generates another representation, it needs a conversion step before the default benchmark metrics apply.&lt;/p&gt;
&lt;p&gt;At a high level, the benchmark has three measurement groups.&lt;/p&gt;
&lt;p&gt;| Group | Flow | What it catches |
| --- | --- | --- |
| Group 1: score-distribution metrics | Sequence $d$ -&gt; scalar score $\Phi(d)$ -&gt; real/generated distribution comparison | Broad microstructure realism across book state, timing, lifecycle, event placement, and order flow. |
| Group 2: market-impact response functions | Event class $\pi$ -&gt; sign-adjusted mid-price response curve $R_{\pi}(l)$ -&gt; real/generated curve comparison | Whether generated trajectories preserve average price response after market events. |
| Group 3: adversarial measurement | Orderbook-state trajectory -&gt; compact state-change representation -&gt; discriminator score | Whether a learned classifier can still separate generated trajectories from real ones. |&lt;/p&gt;
&lt;h3&gt;Group 1: Score-distribution Metrics&lt;/h3&gt;
&lt;p&gt;Group 1 is the main distributional evaluation layer.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Project a sequence to a scalar.&lt;/strong&gt; A scoring function $\Phi$ maps a generated or real sequence $d$ into a scalar score $\Phi(d)$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Compare real and generated score distributions.&lt;/strong&gt; LOB-Bench uses histogram L1 distance for binned distribution mismatch and Wasserstein-1 distance for how far score values move.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Repeat the comparison conditionally.&lt;/strong&gt; The benchmark can bucket by a context variable or conditioning score, compare the target score distribution inside each bucket, and average the bucket-level errors.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Track the same error by rollout horizon.&lt;/strong&gt; Conditioning the comparison on rollout step shows whether a model stays realistic as it samples farther away from the real conditioning history.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h4&gt;Default Score Families&lt;/h4&gt;
&lt;p&gt;The concrete $\Phi$ functions are best read as a menu of probes. They are not the whole benchmark; they are the scalar projections used by Group 1.&lt;/p&gt;
&lt;p&gt;| Metric family | Examples | What it checks |
| --- | --- | --- |
| Quote and book state | Bid-ask spread, order-book imbalance, bid and ask volume, best-level volume | Whether top-of-book tightness, balance, and liquidity look realistic. |
| Message timing and lifecycle | Inter-arrival time, time-to-cancel | Whether event timing and cancellation lifetimes match real data. |
| Event placement | Limit-order depth, cancellation depth, limit-order level, cancellation level | Whether new limits and cancellations occur at realistic distances from the mid-price or book levels. |
| Trading and order flow | Volume per minute, order-flow imbalance, OFI conditional on next mid-price move | Whether trading pressure and directional order-flow patterns are preserved. |&lt;/p&gt;
&lt;h3&gt;Group 2: Market-impact Response Functions&lt;/h3&gt;
&lt;p&gt;Group 2 measures whether generated trajectories preserve average price responses around market events.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Choose an event class.&lt;/strong&gt; LOB-Bench groups events such as market orders, limit orders, and cancellations, and separates them by whether they change the mid-price.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Compute a sign-adjusted response curve.&lt;/strong&gt; For each event class, it tracks the future mid-price response across a lag grid and aligns the sign with the event&apos;s expected price-pressure direction. If $p_t$ is the mid-price at event time $t$ and $\epsilon_t$ is this event-aligned sign, the response curve for event class $\pi$ is&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
R_{\pi}(l) = \left\langle (p_{t+l} - p_t)\epsilon_t \mid \pi_t = \pi \right\rangle_T.
$$&lt;/p&gt;
&lt;ol start=&quot;3&quot;&gt;
&lt;li&gt;&lt;strong&gt;Compare real and generated curves.&lt;/strong&gt; The benchmark measures the gap between the real response curve and the generated response curve. For one event class, this is the mean absolute curve gap across lags:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
\Delta R_{\pi} = \frac{1}{L}\sum_l \left|R_{\pi}^{\mathrm{real}}(l)-R_{\pi}^{\mathrm{gen}}(l)\right|.
$$&lt;/p&gt;
&lt;p&gt;This is a stronger test than a marginal statistic. A model that matches spreads and volumes can still fail if it does not reproduce how prices react after order-flow shocks. The paper reports that LOBS5 reproduces GOOG response curves much better than a stochastic baseline in the discussed comparison.&lt;/p&gt;
&lt;h3&gt;Group 3: Adversarial Measurement&lt;/h3&gt;
&lt;p&gt;Group 3 asks whether a learned classifier can still find trajectory-level artifacts that hand-written metrics miss.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Start from orderbook-state trajectories.&lt;/strong&gt; The adversarial check uses real and generated book-state sequences.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Map state changes into a compact representation.&lt;/strong&gt; LOB-Bench represents sparse book updates through features such as mid-price change, relative price level, and quantity change.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Train a discriminator.&lt;/strong&gt; The discriminator tries to classify whether a trajectory is real or generated.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Read separability as an artifact signal.&lt;/strong&gt; If the discriminator separates real and generated samples well, then the generated data still contains detectable structure that the interpretable metrics may not fully describe.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2&gt;Positioning Current LOB Generators&lt;/h2&gt;
&lt;p&gt;LOB-Bench is most useful as a map of where current LOB generators stand. LOBS5 is the strongest tested model overall, but it is not treated as solved market simulation: the benchmark still exposes horizon error, conditional microstructure failures, response-curve gaps, and discriminator separability.&lt;/p&gt;
&lt;p&gt;The comparisons also show why the generated object matters. Message-only RWKV variants diverge quickly on price levels and book-volume-related statistics, which suggests that plausible event tokens are not enough when the book state drifts. Hand-coded or parametric baselines can match some placement statistics, such as depths and levels, while still missing imbalance, volume, timing, and market-impact behavior.&lt;/p&gt;
&lt;p&gt;This makes LOB-Bench a positioning tool rather than only a leaderboard. It helps say whether a model is mainly good at static score distributions, conditional behavior, event-response dynamics, or adversarial trajectory realism. It still does not validate a single inserted action, queue position, fill probability, order identity, or matching-engine correctness; those claims need simulator- and execution-specific checks.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Nagy et al., &lt;em&gt;LOB-Bench: Benchmarking Generative AI for Finance -- an Application to Limit Order Book Markets&lt;/em&gt; (2025), arXiv:2502.09172.&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>SFT Memorizes, RL Generalizes</title><link>https://junoh-kang.github.io/blog/sft-memorizes-rl-generalizes</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/sft-memorizes-rl-generalizes</guid><description>A review of SFT Memorizes, RL Generalizes, focusing on how supervised finetuning and verifier-based reinforcement learning behave under shifted rules and visual inputs.</description><pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;This post reviews &lt;em&gt;SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training&lt;/em&gt; by Chu et al. The paper asks a simple but important question: when we post-train a foundation model, does the model learn a transferable rule, or does it mainly become better at reproducing the rule distribution seen during training?&lt;/p&gt;
&lt;p&gt;The paper finds that &lt;strong&gt;verifier-based RL generalizes better than SFT when the evaluation rule changes&lt;/strong&gt;. In the tested GeneralPoints and V-IRL environments, supervised finetuning improves or fits the training setting but often collapses under shifted rules. PPO-style reinforcement learning with verifier feedback improves out-of-distribution rule and visual performance from the same SFT-initialized checkpoint.&lt;/p&gt;
&lt;p&gt;This does not mean that SFT is useless. In the paper&apos;s setup, RL works well after SFT has taught the model a usable answer format; &lt;strong&gt;without that SFT initialization, RL does not obtain reliable structured rewards&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Figure 1 from Chu et al. illustrates the paper&apos;s core pattern: under an OOD rule shift, RL improves while SFT collapses. |&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;The Question: Does Post-training Memorize or Generalize?&lt;/h2&gt;
&lt;p&gt;Post-training is often described as the stage that makes a pretrained model useful. SFT teaches the model to follow instructions and produce the desired answer format. RL then optimizes the model against a reward signal, which can come from human preference models, verifiers, or task-specific outcome checks.&lt;/p&gt;
&lt;p&gt;The paper reframes this familiar pipeline as a generalization question. If a model is trained on one rule and then evaluated on a related but different rule, what does each post-training method preserve? A method that really learns the underlying principle should adapt to the new instruction. A method that mainly fits the training distribution may keep applying the old rule even when the prompt states a new one.&lt;/p&gt;
&lt;h3&gt;Experiment Design: Train on One Rule, Test on Another&lt;/h3&gt;
&lt;p&gt;This is why the paper compares SFT and RL under shifted rules rather than only in-distribution accuracy. In-distribution improvement alone cannot distinguish rule learning from rule memorization. The interesting test is whether the trained behavior survives when the rule, visual input, or action convention changes.&lt;/p&gt;
&lt;p&gt;The paper uses the same high-level experimental pattern across the main settings:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Start from the same SFT-initialized checkpoint.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Scale post-training in two different ways.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SFT path: supervised examples under the training rule.&lt;/li&gt;
&lt;li&gt;RL path: PPO-style RL with verifier reward and feedback.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Evaluate after changing the rule or visual condition.&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;A small but important detail is that verification is iterative. The model can revise its answer after verifier feedback, and one GeneralPoints ablation shows that more verification iterations improve RL&apos;s OOD gains. This supports the idea that RL benefits from learning how to use feedback.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Experiments&lt;/h2&gt;
&lt;h3&gt;Rule Generalization&lt;/h3&gt;
&lt;p&gt;The rule generalization results are the clearest evidence for the paper&apos;s title. In GeneralPoints, the training rule treats J, Q, and K as 10. The shifted rule evaluates them as 11, 12, and 13. In V-IRL, the training rule uses absolute orientation actions such as north and east. The shifted rule uses relative actions such as left and right.&lt;/p&gt;
&lt;p&gt;| Task | RL OOD performance | SFT OOD performance | Metric |
| --- | ---: | ---: | --- |
| GP-L | 11.5% -&gt; 15.0% | 11.5% -&gt; 3.4% | Episode success |
| GP-VL | 11.2% -&gt; 14.2% | 11.2% -&gt; 5.6% | Episode success |
| V-IRL-L | 80.8% -&gt; 91.8% | 80.8% -&gt; 1.3% | Per-step accuracy |
| V-IRL-VL | 35.7% -&gt; 45.0% | 35.7% -&gt; 2.5% | Per-step accuracy |&lt;/p&gt;
&lt;p&gt;Under these shifts, RL improves OOD performance from the shared initial checkpoint. SFT moves in the opposite direction. The V-IRL numbers are especially striking. SFT does not merely fail to improve under the shifted action convention. It almost collapses. This supports the paper&apos;s interpretation that SFT can overfit to the training rule or output convention, while verifier-based RL can preserve more rule-following flexibility.&lt;/p&gt;
&lt;h3&gt;Visual Generalization&lt;/h3&gt;
&lt;p&gt;The paper also tests whether the RL advantage survives visual shifts. In GP-VL, the shift changes card suit colors. In V-IRL-VL, evaluation moves from New York training routes to the multi-city VLN mini benchmark. This matters because rule following alone does not guarantee robust visual recognition.&lt;/p&gt;
&lt;p&gt;| Task | RL OOD change | SFT OOD change | Visual shift |
| --- | ---: | ---: | --- |
| GP-VL | 23.6% -&gt; 41.2% | 23.6% -&gt; 13.7% | Black suits to red suits |
| V-IRL-VL | 16.7% -&gt; 77.8% | 16.7% -&gt; 11.1% | NYC routes to VLN mini benchmark |&lt;/p&gt;
&lt;p&gt;The visual results suggest that RL is not only improving final task behavior. In GP-VL, RL improves both recognition accuracy and episode success, while SFT deteriorates both. The authors hypothesize that SFT may locally overfit to frequent reasoning tokens while neglecting recognition tokens, but they leave the exact mechanism unresolved. In V-IRL-VL, verifier-backed revision appears to help the model use visual and instruction information more robustly under a new route distribution.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Why SFT Still Matters&lt;/h2&gt;
&lt;p&gt;The paper&apos;s conclusion is not &quot;skip SFT.&quot; Its main RL pipeline starts from an SFT-initialized model because the verifier needs structured outputs to assign useful rewards. When the authors apply RL directly to the base Llama-3.2-Vision-11B model on GP-L, the runs fail to improve: the base model often produces long, tangential, or unstructured answers, so the verifier cannot reliably extract task information.&lt;/p&gt;
&lt;p&gt;SFT therefore acts as a format teacher. It teaches the model to answer in a way that the verifier can parse, creating the usable initialization window for RL. This also explains why the result does not directly contradict DeepSeek-R1-style claims that strong reasoning behavior may emerge with little or no cold-start SFT in other settings. The requirement depends on the backbone, task, verifier, and output format.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Chu et al., &lt;em&gt;SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training&lt;/em&gt; (ICML 2025), arXiv:2501.17161v2.&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Post-Training of Modern LLMs</title><link>https://junoh-kang.github.io/blog/post-training-of-modern-llms</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/post-training-of-modern-llms</guid><description>Modern LLM post-training, from preference learning to reinforcement learning with verifiable rewards.</description><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;Modern LLMs are usually built in stages. Pre-training gives the model broad next-token prediction capability. Instruction tuning teaches the model to follow user instructions. Post-training then shapes the model&apos;s behavior with preference feedback, reward optimization, or verifiable task rewards.&lt;/p&gt;
&lt;p&gt;This post follows two post-training routes. The first is RLHF, where preference data is used through PPO-style reward optimization or direct preference optimization; the examples are InstructGPT and DPO. The second is RLVR, where the reward comes from checkable outcomes rather than a learned preference model; the examples are DeepSeekMath&apos;s GRPO, DeepSeek-R1-Zero, and the final DeepSeek-R1 pipeline.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;From Next-token Prediction to RLHF&lt;/h2&gt;
&lt;p&gt;The limitation of next-token prediction is that it imitates text rather than directly optimizing assistant behavior. A base model can learn fluent language from web-scale data, but helpfulness, truthfulness, and harmlessness are not direct training targets. More importantly, next-token prediction does not provide direct negative feedback on what behavior to avoid.&lt;/p&gt;
&lt;p&gt;RLHF adds a preference signal after supervised instruction tuning by increasing the probability of preferred behavior and making dispreferred behavior less likely.&lt;/p&gt;
&lt;h3&gt;InstructGPT&lt;/h3&gt;
&lt;p&gt;PPO-based RLHF is the explicit reward-model route. The RL objective is not just &quot;maximize reward.&quot; It also constrains the updated policy to stay close to a reference policy. In InstructGPT, the PPO-ptx objective adds a pretraining loss term to reduce capability regressions:&lt;/p&gt;
&lt;p&gt;$$
\max_{\phi};
\mathbb{E}&lt;em&gt;{(x,y)\sim D&lt;/em&gt;{\pi_\phi^{\mathrm{RL}}}}
\left[
r_\theta(x,y)-\beta \log
\frac{\pi_\phi^{\mathrm{RL}}(y\mid x)}
{\pi^{\mathrm{SFT}}(y\mid x)}
\right]
+
\gamma,
\mathbb{E}&lt;em&gt;{x\sim D&lt;/em&gt;{\mathrm{pretrain}}}
\left[
\log \pi_\phi^{\mathrm{RL}}(x)
\right].
$$&lt;/p&gt;
&lt;p&gt;The first term optimizes the reward model while penalizing deviation from the SFT policy. The second term is the ptx loss: it keeps part of the update anchored to pretraining data to reduce broad capability regressions.&lt;/p&gt;
&lt;h4&gt;Result&lt;/h4&gt;
&lt;p&gt;The result was strong: PPO and PPO-ptx outperformed SFT and prompted base models on labeler preference evaluations, and the 1.3B PPO-ptx model was preferred over the 175B GPT-3 baseline despite the parameter gap.&lt;/p&gt;
&lt;h4&gt;Limitation&lt;/h4&gt;
&lt;p&gt;The limitation is the reward model. It is costly to train, it is not a perfect proxy for human preference, and optimizing against it can create reward hacking or reward-model over-optimization problems.&lt;/p&gt;
&lt;h3&gt;DPO&lt;/h3&gt;
&lt;p&gt;DPO is the direct preference-optimization route inside the same preference-based framing. It starts from the same KL-regularized preference view, but removes the explicit reward-model and PPO stages. Its key observation is that the optimal reward can be represented by a policy/reference log-ratio:&lt;/p&gt;
&lt;h1&gt;$$
\hat r_\theta(x,y)&lt;/h1&gt;
&lt;p&gt;\beta \log
\frac{\pi_\theta(y \mid x)}
{\pi_{\mathrm{ref}}(y \mid x)}.
$$&lt;/p&gt;
&lt;p&gt;The resulting loss directly compares a preferred answer $y_w$ and a dispreferred answer $y_l$:&lt;/p&gt;
&lt;h1&gt;$$
L_{\mathrm{DPO}}(\pi_\theta;\pi_{\mathrm{ref}})&lt;/h1&gt;
&lt;h2&gt;-\mathbb{E}&lt;em&gt;{(x,y_w,y_l)}
\left[
\log \sigma\left(
\beta \log \frac{\pi&lt;/em&gt;\theta(y_w\mid x)}{\pi_{\mathrm{ref}}(y_w\mid x)}&lt;/h2&gt;
&lt;p&gt;\beta \log \frac{\pi_\theta(y_l\mid x)}{\pi_{\mathrm{ref}}(y_l\mid x)}
\right)
\right].
$$&lt;/p&gt;
&lt;p&gt;This is why DPO is sometimes described as RL-free. That phrase is easy to misunderstand. DPO removes the explicit reward model and PPO optimization loop, but it still uses preference pairs, a reference policy, and a KL-regularized objective.&lt;/p&gt;
&lt;h4&gt;Limitations&lt;/h4&gt;
&lt;p&gt;The original DPO evidence is also scale-limited. Its strongest reported experiments reach models up to 6B parameters, so the paper should be read as evidence for the tested settings rather than as proof that DPO replaces PPO-based RLHF at frontier scale.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;From Preference Rewards to Verifiable Rewards&lt;/h2&gt;
&lt;p&gt;Preference-based RLHF is useful, but the reward signal is expensive and subjective. Human labelers or reward models must judge completed responses. The learned reward model can also be over-optimized, especially when the policy discovers answers that score well under the reward model without being genuinely useful.&lt;/p&gt;
&lt;p&gt;RLVR changes the reward source. In math, code, and some reasoning tasks, the final answer can be checked. A math answer can be compared against the known answer. A program can be run against tests. In these settings, the training loop does not need a learned preference model for every completed answer. This distinction is the bridge from preference alignment to reasoning RL. RLVR is attractive because the reward can be less subjective. The hope is that training on logical, checkable problems can also improve the model&apos;s broader reasoning capability.&lt;/p&gt;
&lt;p&gt;| Setting | Reward signal | Natural domain |
| --- | --- | --- |
| RLHF / DPO | Human or model preference | Chat, helpfulness, safety |
| RLVR | Rule-based outcome reward | Math, code, verifiable reasoning |&lt;/p&gt;
&lt;hr&gt;
&lt;h3&gt;GRPO and DeepSeekMath&lt;/h3&gt;
&lt;p&gt;DeepSeekMath introduces Group Relative Policy Optimization, or GRPO, as a PPO-style method for mathematical reasoning. The main engineering problem is the learned value model, or critic. PPO normally uses a value function to estimate advantages. For LLM reasoning, that value model can be another large model, and math rewards often arrive only after a full solution is generated.&lt;/p&gt;
&lt;p&gt;GRPO removes the learned critic. For each question, it samples a group of outputs, scores them, normalizes the rewards inside the group, and uses those group-relative values as advantages. The GRPO objective can be remembered as a grouped PPO clipped objective with a reference-policy KL penalty:&lt;/p&gt;
&lt;p&gt;$$
\begin{aligned}
J_{\mathrm{GRPO}}(\theta)
&amp;#x26;=
\mathbb{E}&lt;em&gt;{q,,{o_i}&lt;/em&gt;{i=1}^{G}}
\Bigg[
\frac{1}{G}\sum_{i=1}^{G}
\frac{1}{|o_i|}\sum_{t=1}^{|o_i|}
\left(
\min\left(
\rho_{i,t}(\theta)A_i,,
\operatorname{clip}\left(\rho_{i,t}(\theta),1-\epsilon,1+\epsilon\right)A_i
\right)
-\beta\mathbb{D}&lt;em&gt;{\mathrm{KL}}\left(\pi&lt;/em&gt;\theta\Vert\pi_{\mathrm{ref}}\right)
\right)
\Bigg].
\end{aligned}
$$&lt;/p&gt;
&lt;p&gt;The group-relative advantage is:&lt;/p&gt;
&lt;h1&gt;$$
A_i&lt;/h1&gt;
&lt;p&gt;\frac{
r_i-\operatorname{mean}({r_1,r_2,\ldots,r_G})
}{
\operatorname{std}({r_1,r_2,\ldots,r_G})
}.
$$&lt;/p&gt;
&lt;p&gt;The core move is critic removal, not reward removal. DeepSeekMath-RL still scores sampled outputs and still uses a reference policy for KL regularization. GRPO should therefore be read as a PPO variant, not as a wholly separate reinforcement learning family.&lt;/p&gt;
&lt;h3&gt;DeepSeek-R1-Zero&lt;/h3&gt;
&lt;p&gt;R1-Zero is the cleaner RLVR case. It applies RL directly to DeepSeek-V3-Base with verifiable rewards. AIME 2024 is a good example: the final answer is an integer from 0 to 999, so correctness can be checked by exact match. On AIME 2024, the paper reports pass@1 rising from 15.6% to 77.9%, and self-consistency with 16 samples reaching 86.7%. It also reports longer generated reasoning during training.&lt;/p&gt;
&lt;h3&gt;DeepSeek-R1&lt;/h3&gt;
&lt;p&gt;R1-Zero shows pure RLVR can induce reasoning, but with usability issues such as CoT readability and language mixing. The final DeepSeek-R1 pipeline turns that idea into a more usable multi-stage training recipe:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Cold-start SFT fixes CoT readability before RL.&lt;/li&gt;
&lt;li&gt;First RL improves reasoning while reducing language mixing.&lt;/li&gt;
&lt;li&gt;Second SFT broadens the model beyond math and code reasoning.&lt;/li&gt;
&lt;li&gt;Final RL aligns general assistant behavior with helpfulness and safety.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;Limitations of RLVR and GRPO&lt;/h3&gt;
&lt;h4&gt;GRPO normalization bias&lt;/h4&gt;
&lt;p&gt;GRPO removes the learned critic, but it does not remove all design choices from the objective. Two normalization terms are especially important:&lt;/p&gt;
&lt;p&gt;$$
\frac{1}{|o_i|}\sum_{t=1}^{|o_i|}L^{\mathrm{PPO}}_{i,t}
\qquad
A_i =
\frac{r_i-\operatorname{mean}(r)}
{\operatorname{std}(r)}.
$$&lt;/p&gt;
&lt;p&gt;The first term is length normalization. Because the token-level loss is averaged by output length, a long incorrect output can receive a weaker per-token penalty. If a model expects to be wrong, lengthening the answer can dilute the penalty signal.&lt;/p&gt;
&lt;p&gt;The second term is reward standard-deviation normalization. Dividing by $\operatorname{std}(r)$ can overweight questions where the sampled answers have low reward variance. This can happen for very easy questions where most samples are correct, or for very hard questions where most samples are wrong.&lt;/p&gt;
&lt;p&gt;Dr. GRPO is motivated by these caveats: it removes both the length normalization and the reward standard-deviation normalization terms.&lt;/p&gt;
&lt;h4&gt;RLVR outcome-level reward&lt;/h4&gt;
&lt;p&gt;RLVR usually gives an outcome-level reward. It can check whether the final answer is right, but it does not automatically explain which reasoning step was wrong. It is still a sparse signal for improving the reasoning process itself.&lt;/p&gt;
&lt;h3&gt;References&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Ernest Ryu, &lt;a href=&quot;https://ernestryu.com/courses/RL-LLM.html&quot;&gt;&lt;em&gt;RL for LLMs&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Ouyang et al., &lt;em&gt;Training language models to follow instructions with human feedback&lt;/em&gt; (2022).&lt;/li&gt;
&lt;li&gt;Rafailov et al., &lt;em&gt;Direct Preference Optimization: Your Language Model is Secretly a Reward Model&lt;/em&gt; (2023).&lt;/li&gt;
&lt;li&gt;Shao et al., &lt;em&gt;DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models&lt;/em&gt; (2024).&lt;/li&gt;
&lt;li&gt;Guo et al., &lt;em&gt;DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning&lt;/em&gt; (2025).&lt;/li&gt;
&lt;li&gt;Liu et al., &lt;em&gt;Understanding R1-Zero-Like Training: A Critical Perspective&lt;/em&gt; (2025).&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>One-step Generation in the Post Diffusion Era</title><link>https://junoh-kang.github.io/blog/one-step-generation-in-the-post-diffusion-era</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/one-step-generation-in-the-post-diffusion-era</guid><description>Training-based routes toward one-step generation for diffusion and flow models, covering distillation, Consistency Models, CTM, MeanFlow, DMD, and Drifting Models.</description><pubDate>Thu, 12 Mar 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;Diffusion models are still one of the major paradigms in image and video generation. They produce high-quality samples and remain a strong default for many generative tasks. However, their sampling process is inherently iterative: generation usually requires many denoising steps or numerical solver evaluations. This inference-time iteration becomes a practical bottleneck when we care about fast generation, large-scale serving, or long-form generation.&lt;/p&gt;
&lt;p&gt;This post gives a high-level review of training-based approaches toward one-step or few-step generation: methods that try to compress, replace, or bypass the inference-time trajectory through distillation, consistency constraints, average-velocity regression, distribution matching, or training-time distribution evolution.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Drifting Models&lt;/strong&gt; are discussed in more detail than the others because they make this contrast especially explicit. Diffusion and flow models evolve samples during inference, while Drifting Models interpret training itself as a distribution-evolution process: as the generator parameters change, the generated distribution drifts toward the data distribution. The purpose of the post is still broader than Drifting Models alone: to organize several attempts at one-step generation and clarify what each method depends on.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Generative Modeling&lt;/h2&gt;
&lt;h3&gt;Pushforward Formulation&lt;/h3&gt;
&lt;p&gt;At a high level, generative modeling can be written as a pushforward problem. We start from a simple prior distribution, such as a Gaussian noise distribution $p_\epsilon$, and learn a network $f_\theta$ that transforms it into a generated distribution&lt;/p&gt;
&lt;p&gt;$$
q_\theta = (f_\theta)&lt;em&gt;# p&lt;/em&gt;\epsilon \approx p_{\mathrm{data}}.
$$&lt;/p&gt;
&lt;h3&gt;Diffusion / Flow Models: Inference-Time Iteration&lt;/h3&gt;
&lt;p&gt;Diffusion and flow models make the pushforward problem tractable by decomposing a complex transformation into many small steps. Instead of learning the full noise-to-data map in one shot, they introduce an intermediate time variable and learn a score, denoising direction, or velocity field that tells each sample which direction to move along the trajectory.&lt;/p&gt;
&lt;p&gt;In both cases, sampling can be viewed as solving learned SDE or ODE dynamics that move each sample from noise toward data at inference time. Numerically solving these dynamics requires iterative updates. This multi-step inference-time sampling is the bottleneck that motivates one-step generation.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Training-Based Attempts Toward One-Step Generation&lt;/h2&gt;
&lt;p&gt;This section reviews training-based attempts to obtain one-step or few-step generators. The methods differ mainly in what training signal they use: teacher trajectories, trajectory consistency, average velocity, score-based distribution matching, or training-time distribution evolution.&lt;/p&gt;
&lt;h3&gt;Progressive Distillation&lt;/h3&gt;
&lt;p&gt;Progressive distillation trains a student one-step transition to match two consecutive teacher transitions:&lt;/p&gt;
&lt;p&gt;$$
f_\theta(z_t, t \to s)
\approx
f_\eta\left(f_\eta(z_t, t \to u), u \to s\right),
\qquad t &gt; u &gt; s.
$$&lt;/p&gt;
&lt;h3&gt;Consistency Models&lt;/h3&gt;
&lt;p&gt;The key idea of Consistency Models is to learn a function whose output is consistent along the same probability-flow ODE trajectory. Instead of predicting only the next denoising step, the model maps noisy samples at different times to the same clean endpoint.&lt;/p&gt;
&lt;p&gt;The consistency constraint is:&lt;/p&gt;
&lt;p&gt;$$
f_\theta(x_t,t)
\approx
f_\theta(x_{t^\prime},t^\prime),
\qquad
x_t,;x_{t^\prime} \text{ on the same trajectory}.
$$&lt;/p&gt;
&lt;p&gt;The practical question is how to obtain such paired points. Consistency Models use two main training setups: Consistency Distillation (CD) and Consistency Training (CT). The pair construction differs as follows:&lt;/p&gt;
&lt;p&gt;$$
\begin{cases}
\text{Consistency Distillation (CD):} &amp;#x26; x_{t^\prime} = \operatorname{ODESolver}(x_t, t \to t^\prime; s_\phi), \
\text{Consistency Training (CT):} &amp;#x26; x_t = \alpha_t x_0 + \sigma_t \epsilon,\quad
x_{t^\prime} = \alpha_{t^\prime} x_0 + \sigma_{t^\prime} \epsilon.
\end{cases}
$$&lt;/p&gt;
&lt;p&gt;CD uses a pretrained score model $s_\phi$ and an ODE solver to move along the probability-flow trajectory, while CT constructs two noisy versions of the same clean sample using the same noise.&lt;/p&gt;
&lt;p&gt;After training, one-step sampling is a single evaluation from terminal noise:&lt;/p&gt;
&lt;p&gt;$$
x_0 = f_\theta(x_T,T).
$$&lt;/p&gt;
&lt;h4&gt;Limitation&lt;/h4&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Unstable training.&lt;/strong&gt; Consistency Models are highly sensitive to curriculum design, such as how the time discretization is scheduled and how difficult trajectory pairs are introduced during training.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Indirect training signal.&lt;/strong&gt; The constraint $f_\theta(x_t,t) = f_\theta(x_{t^\prime},t^\prime)$ only says that two outputs should match. By itself, this admits trivial solutions, so practical training needs boundary conditions and other engineering tricks to make the shared output correspond to a real clean sample.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Dependence on diffusion teacher in CD.&lt;/strong&gt; Consistency distillation still requires a pretrained diffusion or score model to construct reliable same-trajectory pairs.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;CTM&lt;/h3&gt;
&lt;p&gt;Consistency Trajectory Models (CTM) extend Consistency Models from endpoint consistency to trajectory consistency. Instead of only mapping a noisy point to the clean endpoint, CTM learns a transition map&lt;/p&gt;
&lt;p&gt;$$
G_\theta(x_t,t,s)
$$&lt;/p&gt;
&lt;p&gt;that moves a sample from time $t$ to another time $s$ along the probability-flow ODE trajectory.&lt;/p&gt;
&lt;p&gt;The trajectory consistency constraint is:&lt;/p&gt;
&lt;p&gt;$$
G_\theta(x_t,t,s)
\approx
G_\theta(G_\theta(x_t,t,u),u,s).
$$&lt;/p&gt;
&lt;p&gt;That is, a direct jump from $t$ to $s$ should match a composed path through an intermediate time $u$.&lt;/p&gt;
&lt;p&gt;CTM also has to handle special cases of the transition map. For example,&lt;/p&gt;
&lt;p&gt;$$
\begin{cases}
G_\theta(x_t,t,t)=x_t, &amp;#x26; \text{identity at the same time}, \
G_\theta(x_t,t,0)\approx x_0, &amp;#x26; \text{endpoint prediction}.
\end{cases}
$$&lt;/p&gt;
&lt;p&gt;After training, one-step sampling is:&lt;/p&gt;
&lt;p&gt;$$
x_0 = G_\theta(x_T,T,0).
$$&lt;/p&gt;
&lt;p&gt;Because the target time $s$ is an input, the same model can also be used for multi-step sampling by composing shorter transitions. In that sense, CTM tries to learn the probability-flow ODE trajectory itself, not only the final endpoint.&lt;/p&gt;
&lt;h4&gt;Limitation&lt;/h4&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Additional constraints for special cases.&lt;/strong&gt; CTM is more flexible than CM because it models transitions between arbitrary times, but that flexibility introduces boundary cases such as $s=t$ and $s=0$ that need to be handled by the parameterization and loss design.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Training stability.&lt;/strong&gt; CTM still has an indirect training signal: the model is trained by making direct and composed paths agree, unlike MeanFlow&apos;s velocity-regression-style objective.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Not self-contained.&lt;/strong&gt; Strong practical performance requires an additional GAN component.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;MeanFlow&lt;/h3&gt;
&lt;p&gt;The key idea of MeanFlow is to learn the average displacement over a time interval, rather than the instantaneous velocity at each time. Flow Matching learns a local velocity field $v(x_t,t)$ and then integrates it with an ODE solver. MeanFlow asks whether the model can learn the interval-level velocity directly.&lt;/p&gt;
&lt;p&gt;The average velocity is defined as:&lt;/p&gt;
&lt;h1&gt;$$
u(z_t,r,t)&lt;/h1&gt;
&lt;p&gt;\frac{1}{t-r}
\int_r^t v(z_\tau,\tau),d\tau.
$$&lt;/p&gt;
&lt;p&gt;The training signal comes from the MeanFlow identity, which connects this average velocity to the instantaneous velocity:&lt;/p&gt;
&lt;h1&gt;$$
v(z_t,t)&lt;/h1&gt;
&lt;p&gt;\partial_z u \cdot v
+
\partial_t u.
$$&lt;/p&gt;
&lt;p&gt;The derivative term is computed with a Jacobian-vector product (JVP), so the method can train the average-velocity model with a regression-style objective rather than a trajectory-consistency constraint.&lt;/p&gt;
&lt;p&gt;After training, one-step sampling applies the learned average displacement over the full interval:&lt;/p&gt;
&lt;p&gt;$$
x_1 = \epsilon + u_\theta(\epsilon,0,1).
$$&lt;/p&gt;
&lt;h4&gt;Limitation&lt;/h4&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Additional constraint for the special case.&lt;/strong&gt; When the interval collapses, MeanFlow needs the average velocity to match the instantaneous velocity:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
u_\theta(x_t,t,t)=v(x_t,t).
$$&lt;/p&gt;
&lt;h3&gt;DMD&lt;/h3&gt;
&lt;p&gt;The key idea of Distribution Matching Distillation (DMD) is to train a one-step generator by directly matching the generated distribution to the data distribution. Instead of enforcing agreement along an ODE trajectory, DMD minimizes a KL divergence between the noised fake distribution and the noised real distribution:&lt;/p&gt;
&lt;p&gt;$$
\int_0^T w(t),
D_{\mathrm{KL}}!\left(q_t^{\mathrm{fake}} ,|, p_t^{\mathrm{real}}\right),dt.
$$&lt;/p&gt;
&lt;p&gt;The important observation is that the gradient of this KL objective can be expressed using score functions:&lt;/p&gt;
&lt;h2&gt;$$
\nabla_\theta
\int_0^T w(t),
D_{\mathrm{KL}}!\left(q_t^{\mathrm{fake}} ,|, p_t^{\mathrm{real}}\right),dt
\approx
\int_0^T w&apos;(t),
\mathbb{E}_{\hat{x}&lt;em&gt;t}
\left[
\epsilon&lt;/em&gt;{\mathrm{real}}(\hat{x}_t,t)&lt;/h2&gt;
&lt;p&gt;\epsilon_{\mathrm{fake}}(\hat{x}_t,t)
\right]
\frac{\partial \hat{x}_t}{\partial \theta}
,dt.
$$&lt;/p&gt;
&lt;p&gt;Here, the real score and fake score tell the generator how the fake distribution should move relative to the real distribution. After training, sampling is simply a single generator forward pass:&lt;/p&gt;
&lt;p&gt;$$
\hat{x}&lt;em&gt;0 = g&lt;/em&gt;\theta(\epsilon),
\qquad
\epsilon \sim p_\epsilon.
$$&lt;/p&gt;
&lt;h4&gt;Limitation&lt;/h4&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Not self-contained.&lt;/strong&gt; DMD requires a pretrained diffusion model to provide the score signal, and the one-step generator also needs a reasonably good initialization.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heavy overhead before training.&lt;/strong&gt; The method requires extra data-generation work, such as constructing paired noise-image data for distillation.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;Comparing One-step Methods&lt;/h3&gt;
&lt;p&gt;| Method | Objective Type | ODE Dep. | Self-Contained |
|:--|:--|:--:|:--:|
| Diffusion Models | Regression | -- | Yes |
| Consistency Models | Consistency constraint | Yes | Partial |
| CTM | Consistency constraint | Yes | Partial |
| MeanFlow | Regression (velocity) | Yes | Yes |
| DMD | Regression (KL-divergence) | No | No |&lt;/p&gt;
&lt;p&gt;One thing I want to highlight from this table is the objective type. Consistency-based objectives are elegant, but the supervision is indirect: the model is asked to make two paths or two time points agree. Regression-style objectives are often easier to reason about and can be more stable in practice because they provide a more explicit target. This is somewhat orthogonal to the main topic of one-step generation, but it matters when comparing these methods as training recipes.&lt;/p&gt;
&lt;p&gt;This is one reason MeanFlow and DMD are interesting in this list. MeanFlow turns flow learning into regression on an average velocity, while DMD turns distribution matching into a score-based regression signal. They are not free of assumptions, but their objectives are closer to direct regression than pure consistency constraints.&lt;/p&gt;
&lt;p&gt;The remaining question is then not simply &quot;can we generate in one step?&quot; These methods already show several ways to do that. The more useful question is what kind of training signal makes one-step generation stable, self-contained, and less dependent on an inference-time trajectory or a pretrained diffusion teacher.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Drifting Models&lt;/h2&gt;
&lt;h3&gt;Training-Time Distribution Evolution&lt;/h3&gt;
&lt;p&gt;Drifting Models take a different perspective on where the generative dynamics happen. In diffusion and flow models, the generated distribution evolves along pseudo-time while sampling. In Drifting Models, the generated distribution evolves during training.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Drifting Models shift the distribution evolution from inference-time sampling to training-time updates. |&lt;/p&gt;
&lt;h3&gt;Method&lt;/h3&gt;
&lt;h4&gt;Drift&lt;/h4&gt;
&lt;p&gt;Suppose the generator at training step $i$ is $f_i$. For a fixed noise sample $\epsilon$, the generated point is $x_i = f_i(\epsilon)$. After one update, the generator becomes $f_{i+1}$ and the same noise maps to $x_{i+1} = f_{i+1}(\epsilon)$. The difference&lt;/p&gt;
&lt;p&gt;$$
\Delta x_i = f_{i+1}(\epsilon) - f_i(\epsilon)
$$&lt;/p&gt;
&lt;p&gt;is the sample&apos;s training-induced movement, which is why the method is called &quot;drifting.&quot; Since parameter updates already move generated samples during training, Drifting Models make this movement explicit and govern it with a drift field $V_{p,q}$, where $p$ is the data distribution and $q$ is the current model distribution.&lt;/p&gt;
&lt;h4&gt;Drift Field&lt;/h4&gt;
&lt;p&gt;The drift field $V_{p,q}(x)$ determines how a generated sample should move when the current model distribution is $q$ and the target distribution is $p$. The desired equilibrium is clear: when $p=q$, samples should stop drifting. To encode this, the field is designed to be anti-symmetric:&lt;/p&gt;
&lt;p&gt;$$
V_{p,q}(x) = -V_{q,p}(x).
$$&lt;/p&gt;
&lt;p&gt;If $p=q$, anti-symmetry implies $V_{p,q}(x)=0$. The converse is more delicate. A zero drift field does not generally guarantee that $p=q$, so the theory relies on expressive kernels and additional identifiability assumptions. This is one of the important gaps in the method.&lt;/p&gt;
&lt;h4&gt;Training Objective&lt;/h4&gt;
&lt;p&gt;The training objective is derived from the fixed-point relation at equilibrium:&lt;/p&gt;
&lt;h1&gt;$$
f_{\hat{\theta}}(\epsilon)&lt;/h1&gt;
&lt;p&gt;f_{\hat{\theta}}(\epsilon)
+
V_{p,q_{\hat{\theta}}}!\left(f_{\hat{\theta}}(\epsilon)\right).
$$&lt;/p&gt;
&lt;p&gt;Drifting Models turn this relation into a stop-gradient regression target:&lt;/p&gt;
&lt;h1&gt;$$
L(\theta; \textcolor{#B509AC}{q_\theta})&lt;/h1&gt;
&lt;h2&gt;\mathbb{E}&lt;em&gt;{\epsilon}
\left[
\left|
\underbrace{f&lt;/em&gt;\theta(\epsilon)}_{\text{prediction}}&lt;/h2&gt;
&lt;p&gt;\underbrace{
\operatorname{sg}
\left(
f_\theta(\epsilon)
+
V_{p,\textcolor{#B509AC}{q_\theta}}(f_\theta(\epsilon))
\right)
}_{\text{frozen target}}
\right|^2
\right].
$$&lt;/p&gt;
&lt;p&gt;The important part is the dependence on $\textcolor{#B509AC}{q_\theta}$. The objective changes with the current generated distribution, so training is not just fitting a pre-defined target field; it repeatedly updates the target according to the model&apos;s own distribution.&lt;/p&gt;
&lt;h3&gt;Drifting Field Design&lt;/h3&gt;
&lt;p&gt;The drift field is built from an attraction-repulsion interaction. For a generated point $x$, positive samples $y^+$ come from the data distribution and negative samples $y^-$ come from the generated distribution. A kernel $k(x,y)$ measures similarity, and the resulting drift has the form of an attraction toward real samples and a repulsion away from generated samples:&lt;/p&gt;
&lt;p&gt;Formally, the drifting field is defined as&lt;/p&gt;
&lt;p&gt;$$
V_{p,q}(x)
\propto
\mathbb{E}_{y^+ \sim p,; y^- \sim q}
\left[
k(x,y^+) k(x,y^-)(y^+ - y^-)
\right].
$$&lt;/p&gt;
&lt;p&gt;This construction satisfies the anti-symmetry condition $V_{p,q}=-V_{q,p}$. The form is close in spirit to contrastive learning: similar real samples pull the generated point, while similar fake samples provide repulsion.&lt;/p&gt;
&lt;h3&gt;Implementation Details&lt;/h3&gt;
&lt;p&gt;Two practical details matter for applying the drift objective to image generation: feature-space drift and CFG-conditioned sampling.&lt;/p&gt;
&lt;h4&gt;Encoder-Based Drift&lt;/h4&gt;
&lt;p&gt;Doing the attraction-repulsion computation directly in raw pixel space is not enough. Pixel distances do not capture semantic similarity well in high-dimensional image domains, so the practical version computes similarity in feature space using a frozen self-supervised encoder $\phi$, such as MoCo, SimCLR, or latent-MAE features. In effect, the kernel becomes&lt;/p&gt;
&lt;h1&gt;$$
k_\phi(x,y)&lt;/h1&gt;
&lt;p&gt;k(\phi(x),\phi(y)).
$$&lt;/p&gt;
&lt;p&gt;The generator can still output pixels or latents, and the encoder is used only during training to construct the drift target. At inference time, the encoder is not evaluated. This is why I would describe Drifting Models as not ODE-dependent, but not completely assumption-free: they remove the diffusion teacher and inference-time trajectory, but introduce a strong dependence on representation and kernel design.&lt;/p&gt;
&lt;h4&gt;Classifier-Free Guidance&lt;/h4&gt;
&lt;p&gt;For conditional generation, CFG is introduced by changing how the positive and negative samples are drawn:&lt;/p&gt;
&lt;h1&gt;$$
\begin{aligned}
y^+ &amp;#x26;\sim p_{\mathrm{data}}(\cdot \mid c), \
y^- &amp;#x26;\sim \tilde{q}(\cdot \mid c)&lt;/h1&gt;
&lt;p&gt;(1-\gamma)q_\theta(\cdot \mid c)
+
\gamma p_{\mathrm{data}}(\cdot \mid \phi_{\mathrm{null}}),
\end{aligned}
$$&lt;/p&gt;
&lt;p&gt;$\phi_{\mathrm{null}}$ denotes the null condition. During training, the drift field constructed from these positive and negative samples is used to train the $\gamma$-conditioned generator $f_\theta(\epsilon,c,\gamma)$. During sampling, the drift field is not evaluated. The trained generator directly takes the condition and guidance scale:&lt;/p&gt;
&lt;p&gt;$$
x = f_\theta(\epsilon,c,\gamma),
\qquad
\epsilon \sim p_\epsilon.
$$&lt;/p&gt;
&lt;p&gt;So CFG is not applied by mixing two predictions at every denoising step, as in diffusion sampling. It is learned during training and exposed as an input to the one-step generator.&lt;/p&gt;
&lt;h3&gt;Results&lt;/h3&gt;
&lt;h4&gt;Importance of Anti-Symmetry&lt;/h4&gt;
&lt;p&gt;The anti-symmetric attraction-repulsion structure is not cosmetic. When the balance is broken, FID degrades sharply, which suggests that the field design is central to the method.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Breaking the anti-symmetric attraction-repulsion design substantially degrades FID. |&lt;/p&gt;
&lt;h4&gt;Ablations on Encoders&lt;/h4&gt;
&lt;p&gt;The feature encoder is another major factor. The ablation shows that representation choice has a large effect on FID, reinforcing that Drifting Models depend strongly on the feature space used to compute the kernel.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| The feature encoder choice is a major part of the method&apos;s empirical behavior. |&lt;/p&gt;
&lt;h3&gt;Limitation&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Theoretical gap.&lt;/strong&gt; A vanishing drift field, $V \to 0$, does not by itself guarantee $p=q$; the identifiability argument remains heuristic.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Kernel and representation dependency.&lt;/strong&gt; The field is computed through $k(\phi(x),\phi(y))$, so performance depends heavily on the encoder and kernel design.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Design optimality is not guaranteed.&lt;/strong&gt; The kernel, drifting field, and architecture are still design choices with room for improvement.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Scalability is still unverified.&lt;/strong&gt; The presented setting does not yet cover generation beyond $256 \times 256$ or text conditioning, and the contrastive-like objective benefits from large batch sizes.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;Positioning&lt;/h3&gt;
&lt;p&gt;| Method | Objective Type | ODE Dep. | Self-Contained |
|:--|:--|:--:|:--:|
| Diffusion Models | Regression | -- | Yes |
| Consistency Models | Consistency constraint | Yes | Partial |
| CTM | Consistency constraint | Yes | Partial |
| MeanFlow | Regression (velocity) | Yes | Yes |
| DMD | Regression (KL-divergence) | No | No |
| &lt;strong&gt;Drifting Models&lt;/strong&gt; | Regression (drift) | No | Partial |&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Luhman and Luhman, &lt;em&gt;Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed&lt;/em&gt; (2021).&lt;/li&gt;
&lt;li&gt;Salimans and Ho, &lt;em&gt;Progressive Distillation for Fast Sampling of Diffusion Models&lt;/em&gt; (ICLR 2022).&lt;/li&gt;
&lt;li&gt;Song et al., &lt;em&gt;Consistency Models&lt;/em&gt; (ICML 2023).&lt;/li&gt;
&lt;li&gt;Kim et al., &lt;em&gt;Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion&lt;/em&gt; (ICLR 2024).&lt;/li&gt;
&lt;li&gt;Geng et al., &lt;em&gt;Mean Flows for One-Step Generative Modeling&lt;/em&gt; (NeurIPS 2025).&lt;/li&gt;
&lt;li&gt;Yin et al., &lt;em&gt;One-step Diffusion with Distribution Matching Distillation&lt;/em&gt; (CVPR 2024).&lt;/li&gt;
&lt;li&gt;Deng et al., &lt;em&gt;Generative Modeling via Drifting&lt;/em&gt; (arXiv:2602.04770).&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Path Signature: Useful Feature for Timeseries</title><link>https://junoh-kang.github.io/blog/path-signature-useful-feature-for-timeseries</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/path-signature-useful-feature-for-timeseries</guid><description>Study note on path signature in the literature of rough path theory. This post introduces the definition, algebraic structure, and probabilistic interpretation of signatures, bridging rough path theory and modern machine learning, based primarily on *A Primer on the Signature Method in Machine Learning*.</description><pubDate>Sun, 01 Feb 2026 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;(Motivation) Random Variable: Polynomial Features&lt;/h2&gt;
&lt;h3&gt;Definition (Words / Multi-Indices)&lt;/h3&gt;
&lt;p&gt;A word (or multi-index) is a finite sequence $I:=i_1 \ldots i_k$ with $i_1,\ldots,i_k \in {1,\ldots,d}$. The length of $I$ is $|I|:=k$. The set ${1,\ldots,d}$ is the alphabet. Denote $\mathcal{W}$ by the set of all words.&lt;/p&gt;
&lt;h3&gt;Definition (Polynomials for RV)&lt;/h3&gt;
&lt;p&gt;For a vector $x=(x^1,\ldots,x^d)$ and a multi-index $I=(i_1,\ldots,i_k)$, define $x^{I}=x^{(i_1,\ldots,i_k)}:=x^{i_1}\ldots x^{i_k}$. The collection ${x^I}_{I\in \mathcal{W}}$ is called the polynomials of $x$.&lt;/p&gt;
&lt;h3&gt;Theorem (Stone-Weierstrass, Carleman)&lt;/h3&gt;
&lt;p&gt;Under suitable integrability conditions on $\mathbb{P}$, $\mathbb{E}&lt;em&gt;{X \sim \mathbb{P}}[X^I] = \mathbb{E}&lt;/em&gt;{X \sim \mathbb{Q}}[Y^I] ~~~\forall I \in \mathcal{W} ~~~~~\Leftrightarrow~~~~~ \mathbb{P} = \mathbb{Q}$&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Paths: Iterated Integrals and Signature&lt;/h2&gt;
&lt;h3&gt;Definition (Path-space)&lt;/h3&gt;
&lt;p&gt;Let $\mathcal{P}([a,b], \mathbb{R}^d)$ denote the space of continuous bounded variation paths $X=(X^1,\ldots,X^d):[a,b] \rightarrow \mathbb{R}^d.$&lt;/p&gt;
&lt;h3&gt;Definition (Iterated integrals)&lt;/h3&gt;
&lt;p&gt;For $X \in \mathcal{P}([a,b], \mathbb{R}^d)$ and word $I=i_1 \ldots i_k$, define the iterated integral $S(X)_{a,\cdot}^I \in \mathcal{P}([a,b], \mathbb{R})$ inductively by&lt;/p&gt;
&lt;p&gt;$$
S(X)&lt;em&gt;{a,t}^{I}=\int_a^t S(X)&lt;/em&gt;{a,s}^{i_1 \ldots i_{k-1}} dX_s^{i_k},
$$&lt;/p&gt;
&lt;p&gt;with base case $k=0$ as $S(X)_{a,t}^{\phi}=1$. In other expression,&lt;/p&gt;
&lt;p&gt;$$
S(X)&lt;em&gt;{a,t}^{I} = \int&lt;/em&gt;{a&amp;#x3C;t_1&amp;#x3C;\ldots&amp;#x3C;t_k&amp;#x3C;t} dX_{t_1}^{i_1} \ldots dX_{t_k}^{i_k}.
$$&lt;/p&gt;
&lt;h3&gt;Definition (Signature)&lt;/h3&gt;
&lt;p&gt;The signature $S(X)&lt;em&gt;{a,b}$ of $X \in \mathcal{P}([a,b], \mathbb{R}^d)$ is the collection of real numbers indexed by words ${ S(X)&lt;/em&gt;{a,b}^I}&lt;em&gt;{I\in\mathcal{W}}=(S(X)&lt;/em&gt;{a,b}^{\phi},S(X)&lt;em&gt;{a,b}^1,\ldots,S(X)&lt;/em&gt;{a,b}^d,S(X)&lt;em&gt;{a,b}^{11}, \ldots S(X)&lt;/em&gt;{a,b}^{dd},\ldots).$&lt;/p&gt;
&lt;h4&gt;Definition&lt;/h4&gt;
&lt;p&gt;$$
S(X)&lt;em&gt;{a,b}^{(k)} = { S(X)&lt;/em&gt;{a,b}^I }_{|I|=k}
$$&lt;/p&gt;
&lt;h3&gt;Exercise (Linear path)&lt;/h3&gt;
&lt;p&gt;Let $x \in \mathbb{R}^d$ and define $X(t) = tx$ for $t\in[0,1]$. Then $S(X)&lt;em&gt;{0,1}^{i_1 \ldots i_k} = \frac{x^{i_1}\ldots x&lt;/em&gt;{i_k}}{k!}.$ sol)&lt;/p&gt;
&lt;p&gt;$$
\begin{align*}
S(X)&lt;em&gt;{0,1}^{i_1 \ldots i_k}
&amp;#x26;= \int&lt;/em&gt;{0}^{1} S(X)&lt;em&gt;{0,t}^{i_1 \ldots i&lt;/em&gt;{k-1}}dX_t^{i_k} \
&amp;#x26;= \int_0^1 \frac{x^{i_1}\ldots x_{i_{k-1}}}{(k-1)!}t^{k-1}x^{i_k} dt ~~(\because induction) \
&amp;#x26;= \frac{x^{i_1}\ldots x_{i_k}}{k!}
\end{align*}_\square
$$&lt;/p&gt;
&lt;h4&gt;Exercise (1D path)&lt;/h4&gt;
&lt;blockquote&gt;
&lt;p&gt;For $X \in \mathcal{P}([a,b],\mathbb{R})$, $S(X)_{a,b}^{1 \ldots 1} = \frac{(X_b-X_a)^k}{k!}.$sol)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;$$
\begin{align*} S(X)&lt;em&gt;{a,b}^{1^{\otimes k}}
&amp;#x26;= \int_a^b S(X)&lt;/em&gt;{a,t}^{1^{\otimes {k-1}}} dX_t \
&amp;#x26;= \int_a^b \frac{(X_t - X_a)^{k-1}}{(k-1)!} \dot{X}_t dt \
&amp;#x26;= [\frac{(X_t - X_a)^{k}}{k!}]&lt;em&gt;a^t
= \frac{(X_b - X_a)^{k}}{k!}
\end{align*}&lt;/em&gt;\square
$$&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Beyond Defaults: Is Noise Conditioning Necessary for Diffusion Models?</title><link>https://junoh-kang.github.io/blog/beyond-defaults-is-noise-conditioning-necessary-for-diffusion-models</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/beyond-defaults-is-noise-conditioning-necessary-for-diffusion-models</guid><description>A review of recent research that challenges the necessity of noise level conditioning in generative models, exploring alternative approaches to denoising and flow matching.</description><pubDate>Wed, 05 Nov 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;Traditional diffusion models explicitly specify noise level to neural network to model score function. However, recent research has begun to question whether this conditioning is truly necessary. In this post, I review two papers that challenge this fundamental assumption and propose alternative approaches:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&quot;Is Noise Conditioning Necessary for Denoising Generative Models?&quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;This paper challenges the convention by demonstrating that denoising networks can perform effectively without explicit noise level conditioning, suggesting that models may inherently learn to estimate noise levels from the input data itself.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;&quot;Equilibrium Matching: Generative Modeling with Implicit Energy-based Models&quot;&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;This work reformulates the generative modeling problem as learning an energy landscape, providing a theoretical foundation for noise-unconditioning approaches.&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>PinT algorithms for Diffusion Models</title><link>https://junoh-kang.github.io/blog/pint-algorithms-for-diffusion-models</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/pint-algorithms-for-diffusion-models</guid><description>A review of researches that accelerate diffusion models in wall clock time by parallelization.</description><pubDate>Thu, 28 Aug 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;Diffusion models require heavy computation resource due to iterative sampling strategy.
Due to their sequential sampling strategy, many accleration algorithms trade &lt;strong&gt;sample quality&lt;/strong&gt; for &lt;strong&gt;efficiency&lt;/strong&gt;.
However, there are a group of researches which trade &lt;strong&gt;compute&lt;/strong&gt; for &lt;strong&gt;time&lt;/strong&gt;.
In this post, I review two papers accelerating sampling in time, which are motivated from &lt;code&gt;PinT (Parallel in Time)&lt;/code&gt; algorithms:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Parallel Sampling of Diffusion Models&lt;/li&gt;
&lt;li&gt;Self-Refining Diffusion Samplers: Enabling Paralleization via Parareal Iterations&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Test-time scaling in Diffusion Models</title><link>https://junoh-kang.github.io/blog/test-time-scaling-in-diffusion-models</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/test-time-scaling-in-diffusion-models</guid><description>A review of researches that explore test-time scaling for diffusion models.</description><pubDate>Thu, 29 May 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;This post reviews three papers related to test-time scaling which enable optimization with restpect to diverse condtions without fine-tuning diffusion models nor training external models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-Based Decoding&lt;/li&gt;
&lt;li&gt;Test-Time Alignment of Diffusion Models without Reward Over-Optimization&lt;/li&gt;
&lt;li&gt;Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing&lt;/li&gt;
&lt;/ul&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Video Diffusion Models as World Simulators</title><link>https://junoh-kang.github.io/blog/video-diffusion-models-as-world-simulators</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/video-diffusion-models-as-world-simulators</guid><description>A review of Oasis: A Universe in a transformer, and Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion</description><pubDate>Fri, 17 Jan 2025 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;World simulators&lt;/strong&gt; are explorable and interactive systems or models that can mimic real world.
Advanced video generation models can function as world simulators, and to achieve it, they should have &lt;strong&gt;low latency&lt;/strong&gt; for input actions, and capable of &lt;strong&gt;long sequence generation&lt;/strong&gt;.
Long sequence generation includes &lt;strong&gt;capability of long generation itself&lt;/strong&gt;, &lt;strong&gt;preventing error accumulation&lt;/strong&gt;, and &lt;strong&gt;long term context preservation&lt;/strong&gt;.
This post mainly focuses on how related project &lt;strong&gt;Oasis: A Universe in a transformer&lt;/strong&gt;[1] deals with &lt;strong&gt;long sequence generation&lt;/strong&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Conventional Long Video Generation are inappropriate for World Simulator!&lt;/h2&gt;
&lt;h3&gt;Video Diffusion Models (VDMs)&lt;/h3&gt;
&lt;p&gt;||
|:--|
| Training and sampling of video diffusion models. Darker tokens has higher noise levels. |&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Videos are sequential data&lt;/strong&gt;.
However, video diffusion models are trained and inferenced to denoise tokens of same noise levels, interpreting each video clip as a single object.
This section reviews approaches to generate long videos using aforementioned VDMs.&lt;/p&gt;
&lt;h3&gt;Chunked autoregressive methods&lt;/h3&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Small $k$ (&lt;em&gt;e.g.&lt;/em&gt; $k=1$) results in high latency for each action since $f-k$ frames are output for each action. Also it tends to lose contexts.&lt;/li&gt;
&lt;li&gt;Large $k$ (&lt;em&gt;e.g.&lt;/em&gt; $k=f-1$) results in ineifficient training and inference since models learn only $f-k$ tokens, while models calculate for $f$ tokens.&lt;/li&gt;
&lt;li&gt;Chunked autoregressive methods suffers from &lt;strong&gt;quality degradation originated from error accumulation&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;h3&gt;Hierarchical methods (Multi-stage generation)&lt;/h3&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It does not fit to interactive generation since the end of the video is already determined.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Conventional approaches are not appropriate for wolrd simulator!&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Long Sequence Generation in Oasis&lt;/h2&gt;
&lt;h3&gt;Capability of long sequence generation&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Oasis&lt;/strong&gt;[1] follows &lt;strong&gt;Diffusion Forcing&lt;/strong&gt;[2] to train models for long video generation.
&lt;strong&gt;Diffusion Forcing&lt;/strong&gt; inherits advantages of Teacher Forcing and Diffusion Models: &lt;strong&gt;flexible time horizon&lt;/strong&gt; from Teacher Forcing, &lt;strong&gt;guidance at sampling&lt;/strong&gt; from Diffusion Models.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Diffusion Forcing&lt;/strong&gt; trains models to denoise &lt;strong&gt;tokens with independent noise levels&lt;/strong&gt;, and sampling noise schedules are carefully chosen depending on the purpose.
The training offers cheaper training than next-token prediction in video domain, and the complexity added by independent noise level is not excessive since the complexity is only in temporal dimension.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Training in Diffusion Forcing |&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Sampling in Diffusion Forcing |&lt;/p&gt;
&lt;h3&gt;Preventing error accumulation&lt;/h3&gt;
&lt;h4&gt;The reason of error accumulation&lt;/h4&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oasis&lt;/strong&gt;[1] and &lt;strong&gt;Diffusion Forcing&lt;/strong&gt;[2] hypothesize that the error accumulation stems from the model erroneously treating generated noisy frames as grount truth (GT), despite their inherent inaccuracies.
They interpret &lt;strong&gt;input noise levels to the models as inversely proportional to the confidence&lt;/strong&gt; in the corresponding input tokens.&lt;/p&gt;
&lt;h4&gt;Stable rollout in Diffusion Forcing&lt;/h4&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Diffusion Forcing&lt;/strong&gt; suggests to deceive models that generated clean tokens are little noisy, preventing models from believing generated tokens as GT.
However, this approach is out of distribution (OOD) inference, and there is no rule of thumb for &quot;little noisy&quot;.&lt;/p&gt;
&lt;h4&gt;Stable rollout (Another option)&lt;/h4&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;p&gt;To avoid OOD, one may suggest add little noise to generated tokens and tell models that the tokens are noisy.
However, this approach may dilute details in generated tokens.&lt;/p&gt;
&lt;h4&gt;Dynamic Noise Augmentation (DNA)&lt;/h4&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Oasis&lt;/strong&gt;[1] suggests Dynamic Noise Augmentation (DNA) to mitigate error accumulation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;For initial denoising steps, conditioning tokens (generated tokens) are moderately noised since models tend to generate low-frequency features during initial steps.&lt;/li&gt;
&lt;li&gt;For last denoising steps, noise levels of conditioning tokens gradually decreases.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Long term context preservation&lt;/h3&gt;
&lt;p&gt;||&lt;/p&gt;
&lt;p&gt;Through above approaches, &lt;strong&gt;Oasis&lt;/strong&gt; can autoregressively generate long videos without much quality degradation.
However, models do not have long time horizon memory, leading to inconsistent videos.
While there is no innovative breakthrough yet, I believe that video models with long-term memory is an important next step.&lt;/p&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://oasis-model.github.io/&quot;&gt;Etched Decard, &lt;em&gt;Oasis: A Universe in a Transformer&lt;/em&gt; (2024)&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&quot;https://boyuan.space/diffusion-forcing/&quot;&gt;Boyuan Chen et al., &lt;em&gt;Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion&lt;/em&gt; (NeurIPS, 2024)&lt;/a&gt;&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Controllabilities in Video Diffusion Models</title><link>https://junoh-kang.github.io/blog/controllabilities-in-video-diffusion-models</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/controllabilities-in-video-diffusion-models</guid><description>A review of papers that add controllabilities to video generation (DragNUWA and Boximator)</description><pubDate>Tue, 14 May 2024 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This presentation file includs videos.
You may use pdf viewer that supports video playing such as Adobe Acrobat reader.&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Designing Diffusion Models in Real World</title><link>https://junoh-kang.github.io/blog/designing-diffusion-models-in-real-world</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/designing-diffusion-models-in-real-world</guid><description>Elucidating the Design Space of Diffusion-Based Generative Models, focusing on the rationale behind its engineering details.</description><pubDate>Thu, 26 Oct 2023 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;In deep learning, practical implementations are just as important as theoretical supports.
Especially when proposing new paradigms, such as GANs, diffusion models, and transformers, &lt;em&gt;etc.&lt;/em&gt;, engineering skills are essential to bring the paradigms into the real world.
Even if the suggested designs on diffusion models in &lt;strong&gt;Elucidating the Design Space of Diffusion-Based Generative Models (EDM)&lt;/strong&gt;[1] are not optimal, the choices of the designs are theoretically or empirically supported.
Learning these reasonings may help to bring your theory into the real world.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Revisit Diffusion Models&lt;/h2&gt;
&lt;h3&gt;Reformulate diffusion models&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;Score-Based Model&lt;/strong&gt;[2] defines forward SDE and marginal distribution is calculated from forward SDE.
However when it comes to training, marginal distribution is more important than the SDE and therefore &lt;strong&gt;EDM&lt;/strong&gt;[1] defines the marginal distribution first.&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
p_t(\mathrm{x}) = s(t)^{-d}p(\mathrm{x} /s(t);\sigma(t)),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $p(\mathrm{x};\sigma) = \left&lt;a href=&quot;%5Cmathrm%7Bx%7D&quot;&gt;p_{\text{data}} * \mathcal{N}(\mathrm{0}, \sigma^2\mathrm{I})\right&lt;/a&gt;$.&lt;/p&gt;
&lt;p&gt;Then, the corresponding probability flow ODE is&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
d\mathrm{x} = \left[\dot{s}(t)/s(t) - s(t)^2 \dot{\sigma}(t)\sigma(t) \nabla_\mathrm{x} \log p(\mathrm{x}_t/s(t);\sigma(t)) \right] dt. \label{edm:ode}
\end{align}
$$&lt;/p&gt;
&lt;h3&gt;Obstacles in diffusion models&lt;/h3&gt;
&lt;p&gt;Generation by diffusion models can interpreted as solving ODE:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
d\mathrm{x} = f(\mathrm{x}_t, s(t), \sigma(t))dt.
\end{align}
$$&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;$f(\mathrm{x}&lt;em&gt;t, s(t), \sigma(t))$ is not known and it is parametrized by a network $f&lt;/em&gt;\theta(\mathrm{x}_t, s(t), \sigma(t))$.
The inaccurate approximation on the target causes degradation. \
&lt;strong&gt;→ Better training!&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The solution at $t=0$ given boundary condition at $t=T$ is \&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
\begin{align}
\mathrm{x}_0 = \mathrm{x}_T + \int_0^T f(\mathrm{x}_t, s(t), \sigma(t))dt.
\end{align}
$$&lt;/p&gt;
&lt;p&gt;The integral is numerically calculated, which causes truncation errors. \
&lt;strong&gt;→ Reduce truncation errors, focus on important region!&lt;/strong&gt;&lt;/p&gt;
&lt;h3&gt;Design space of diffusion models&lt;/h3&gt;
&lt;h4&gt;Components regarding training&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Parametrization and network preconditioning: $c_\text{skip}(\sigma)$, $c_\text{out}(\sigma)$, $c_\text{in}(\sigma)$, $c_\text{noise}(\sigma)$.&lt;/li&gt;
&lt;li&gt;Loss weighting: $\lambda(t)$.&lt;/li&gt;
&lt;li&gt;Noise level distribution for training: $\sigma \sim p_{\text{noise}}$.&lt;/li&gt;
&lt;li&gt;Augmentation&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;Components regarding deterministic sampling&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Truncation-error-reducing ODE: $s(t)$, $\sigma(t)$.&lt;/li&gt;
&lt;li&gt;Truncation-error-reducing algorithms: Higher-roder integrators&lt;/li&gt;
&lt;li&gt;Distributing truncation errors properly: Discretization ${t_i}_0^N$&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;Components regarding stochastic sampling&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Rate of replaced noises: $\beta(t)$&lt;/li&gt;
&lt;li&gt;Heuristics: $S_{\text{tmin}}, S_{\text{tmax}}, S_{\text{noise}}, S_{\text{churn}}$.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2&gt;Improvements to Training&lt;/h2&gt;
&lt;p&gt;For this section, this post assumes $s(t)=1$.&lt;/p&gt;
&lt;h3&gt;Parametrization, network preconditioning, loss weighting&lt;/h3&gt;
&lt;p&gt;$D(\mathrm{x}_t,\sigma)$ is a denoiser which minimizes $\ell_2$-norm with $\mathrm{y}$:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\mathbb{E}&lt;em&gt;{\mathrm{y} \sim p&lt;/em&gt;{\text{data}}}
\mathbb{E}||D(\mathrm{y} + \mathrm{n}) - \mathrm{y}||_2^2. \label{eq:loss}
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Then, the relation between a score function and the ideal denoiser is&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\nabla_{\mathrm{x}} \log p(\mathrm{x};\sigma) = (D(\mathrm{x};\sigma) - \mathrm{x}) / \sigma^2.
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Networks in many baselines predicts either $D(\mathrm{x},\sigma)$ or $\mathrm{n}$.
However, &lt;strong&gt;Dynamic dual-output diffusion models&lt;/strong&gt;[3] observes that predicting $D(\mathrm{x},\sigma)$ is easier for high noise level, while predicting $\mathrm{n}$ is easier for low noise level.&lt;/p&gt;
&lt;p&gt;||
|:--:|
| Loss comparison between predicting the denoised output or the added noise.|&lt;/p&gt;
&lt;p&gt;From the observation, &lt;strong&gt;EDM&lt;/strong&gt;[1] designs the network to predict $D(\mathrm{x};\sigma)$ or $\mathrm{n}$, or something in between according to the noise level.&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
D_\theta(\mathrm{x};\sigma) = c_{\text{skip}}(\sigma)\mathrm{x} + c_{\text{out}}(\sigma) F_\theta(c_{\text{in}}(\sigma)\mathrm{x}; c_{\text{noise}}(\sigma)),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $F_\theta$ is a neural network.&lt;/p&gt;
&lt;p&gt;Then the loss function ($\ref{eq:loss}$) is&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\mathbb{E}&lt;em&gt;{\sigma,\mathrm{y},\mathbf{b}}\left[
\underbrace{\lambda(\sigma)c&lt;/em&gt;{\text{out}}(\sigma)^2}&lt;em&gt;{\text{effective weight}}
||\underbrace{F&lt;/em&gt;\theta(c_{\text{in}}(\sigma)(\mathrm{y}+\mathbf{n};c_{\text{noise}}(\sigma)))}&lt;em&gt;{\text{network output}}
- \underbrace{\frac{1}{c&lt;/em&gt;{\text{out}}(\sigma)}(\mathrm{y} - c_{\text{skip}}(\sigma)(\mathrm{y}+\textbf{n}))}_{\text{effective training target}}||_2^2
\right].
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;1.Network inputs should have bounded range&lt;/h4&gt;
&lt;p&gt;$$
\begin{align}
\Rightarrow &amp;#x26;\text{Var}&lt;em&gt;{\mathrm{y},\mathbf{n}}\left[c&lt;/em&gt;{\text{in}}(\sigma)(\mathrm{y} + \mathbf{n})\right] = 1 \
\Rightarrow &amp;#x26;c_{\text{in}}(\sigma) = 1 / \sqrt{\sigma^2 + \sigma_{\text{data}}^2}\
\text{&amp;#x26; } &amp;#x26;c_{\text{noise}}(\sigma) = \log (\sigma)/4
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;2.Effective training target should have bounded range&lt;/h4&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;\Rightarrow \text{Var}&lt;em&gt;{\mathrm{y},\mathbf{n}}\left[\frac{1}{c&lt;/em&gt;{\text{out}}(\sigma)}(\mathrm{y} - c_{\text{skip}}(\sigma)(\mathrm{y}+\textbf{n}))\right] = 1 \
&amp;#x26;\Rightarrow c_{\text{out}}(\sigma)^2 = (1-c_{\text{skip}}(\sigma))^2\sigma_{\text{data}}^2 + c_{\text{skip}}(\sigma)^2 \sigma^2
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;3.Errors of network should not be amplified&lt;/h4&gt;
&lt;p&gt;$$
\begin{align}
\Rightarrow~ &amp;#x26;c_{\text{skip}}(\sigma) = \underset{c_{\text{skip}}(\sigma)}{\text{argmin}}
~c_{\text{out}}(\sigma) \
\Rightarrow~ &amp;#x26;\begin{cases}
c_{\text{skip}}(\sigma) = \sigma_{\text{data}}^2 / (\sigma^2 + \sigma_{\text{data}}^2) \
c_{\text{out}}(\sigma) = \sigma \cdot \sigma_{\text{data}} / \sqrt{\sigma^2 + \sigma_{\text{data}}^2}
\end{cases}
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;4. Effecitve weight should be uniform&lt;/h4&gt;
&lt;p&gt;$$
\begin{align}
\Rightarrow~ &amp;#x26;\lambda(\sigma)c_{\text{out}}(\sigma)^2 = 1 \
\Rightarrow~ &amp;#x26; \lambda(\sigma) = (\sigma^2 + \sigma_{\text{data}}^2)/(\sigma \cdot \sigma_{\text{data}})^2
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Putting 1 ~ 4 together, the expected value of the loss at each noise level is 1. Moreover, the change of effective training target according to $\sigma$ coincides to the observation of &lt;strong&gt;Dynamic dual-output diffusion models&lt;/strong&gt;[3].&lt;/p&gt;
&lt;h3&gt;Noise level distribution for training&lt;/h3&gt;
&lt;p&gt;||
|:--|
| Observed loss per noise level. The shaded regions represent the standard deviation over 10k random samples. EDM&apos;s proposed training sample density is shown by the dashed red curve.|&lt;/p&gt;
&lt;p&gt;At low noise levels, seperating the small noise components is difficult and irrelevant, whereas at high noise levels, the correct answer approaches to dataset average; &lt;strong&gt;EDM&lt;/strong&gt;[1] focuses on middle range noise levels for training: $\sigma \sim \mathcal{N}(-1.2, 1.2)$.&lt;/p&gt;
&lt;h3&gt;Augmentataion&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;EDM&lt;/strong&gt;[1] follows the augmentation pipiline from the GAN literature[4].
&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Each agmentation is enabled with $A_{\text{prob}}$.&lt;/li&gt;
&lt;li&gt;Draw $a_i$ from each enabled augmentation and construct transformation matrix.&lt;/li&gt;
&lt;li&gt;Pass data through $2\times$ supersampled high-quality Wavelet filters.&lt;/li&gt;
&lt;li&gt;Construct a 9-dimensional conditioning input vector for non-leaking augmentation. This vector makes the network to perform auxiliarty tasks.&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2&gt;Improvements to Deterministic Sampling&lt;/h2&gt;
&lt;h3&gt;Higher-order integrators&lt;/h3&gt;
&lt;p&gt;For $s(t)=1$ and $\sigma(t)=t$, the ODE to solve is&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
d\mathrm{x}/dt = (\mathrm{x}_t - D(\mathrm{x}_t;t))/t := f(\mathrm{x}_t,t).
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Euler_method&quot;&gt;&lt;em&gt;Euler method&lt;/em&gt;&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Euler method approximates the integral by&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\int_{t_{i}}^{t_{i-1}} f(\mathrm{x}&lt;em&gt;t,t) dt
= (t&lt;/em&gt;{i-1} - t_{i})f(\mathrm{x}&lt;em&gt;{t_i},t_i) + O(|t&lt;/em&gt;{i-1}-t_{i}|^2).
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Therefore, the total truncation error is $O(\max \lvert t_{i-1}-t_{i} \rvert)$.
Let $\hat{\mathrm{x}}&lt;em&gt;{t&lt;/em&gt;{i-1}}$ is a solution obtained by Euler method.&lt;/p&gt;
&lt;h4&gt;&lt;a href=&quot;https://en.wikipedia.org/wiki/Heun%27s_method&quot;&gt;&lt;em&gt;Heun&apos;s method&lt;/em&gt;&lt;/a&gt;&lt;/h4&gt;
&lt;p&gt;Then, Heun&apos;s method approximates the integral by&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\int_{t_{i}}^{t_{i-1}} f(\mathrm{x}&lt;em&gt;t,t) dt
= (t&lt;/em&gt;{i-1} - t_{i})(f(\mathrm{x}&lt;em&gt;{t_i},t_i)+ f(\hat{\mathrm{x}}&lt;/em&gt;{t_{i-1}},t_{i-1}))/2+ O(|t_{i-1}-t_{i}|^3).
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Therefore, the total truncation error is $O(\max{\lvert t_{i-1}-t_{i}|^2 \rvert})$.
Huen&apos;s method decreases truncation error at the cost of one additional evaluation of the network.&lt;/p&gt;
&lt;h4&gt;Deterministic sampling algorithm for &lt;strong&gt;EDM&lt;/strong&gt;[1]&lt;/h4&gt;
&lt;h3&gt;Discretization&lt;/h3&gt;
&lt;p&gt;As long as using numerical integrators with limited computational resources, &lt;strong&gt;truncation errors are inevitable&lt;/strong&gt;.
In terms of obtaining ODE trajectories accurately, it is important to minimize total truncation erros.
Hoever, the interests of diffusion models at generation are only the &lt;strong&gt;solutions at low noise levels&lt;/strong&gt;; it is reasonable to &lt;strong&gt;focus on low noise levels&lt;/strong&gt;. EDM discretizes as&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
t_{N-i} = \sigma_{i&amp;#x3C; N} = (\sigma_{\text{max}}^{1/\rho} + \frac{i}{N-1} (\sigma_{\text{min}}^{1/\rho} - \sigma_{\text{max}}^{1/\rho}))^\rho, \sigma_N = 0.
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Increasing $\rho$ results dense discretizations at low noise levels.&lt;/p&gt;
&lt;p&gt;||
|:--:|
|(a),(b) Local truncation error at different noise levels. (c) FID as a function of $\rho$.|&lt;/p&gt;
&lt;p&gt;$\rho=3$ nearly equalizes the truncation error at each step as in (a), (b). On the other hand, $\rho=7$ generates better samples as in (c).
Proper value of $\rho$ changes according to the tasks.
&lt;em&gt;e.g.&lt;/em&gt;, equalized truncation error will be better for solving ODE in both directions.&lt;/p&gt;
&lt;h3&gt;Truncation-error-reducing ODE&lt;/h3&gt;
&lt;p&gt;Many integrators including Euler and Heun&apos;s method have small truncation errors if $f(\mathrm{x}_t,t)$ has &lt;strong&gt;small curvature&lt;/strong&gt;, or is close to linear function. $s(t)$ and $\sigma(t)$ determine the shape of the ODE solution trajectories, which is closely related to linearity of the $f(\cdot)$.&lt;/p&gt;
&lt;p&gt;$$
\int_{t_{i}}^{t_{i-1}} f(\mathrm{x}&lt;em&gt;t,t) dt \approx
\begin{cases}
(t&lt;/em&gt;{i-1} - t_{i})f(\mathrm{x}&lt;em&gt;{t_i},t_i) &amp;#x26; \text{Euler method} \
(t&lt;/em&gt;{i-1} - t_{i}) (f(\mathrm{x}&lt;em&gt;{t_i},t_i)+ f(\hat{\mathrm{x}}&lt;/em&gt;{t_{i-1}},t_{i-1}))/2 &amp;#x26; \text{Heun&apos;s method}
\end{cases}
$$&lt;/p&gt;
&lt;p&gt;||
|:--|
|A sketch of ODE curvature in 1D where $p_{\text{data}}$ is two Dirac peaks at $\mathrm{x}= \pm 1$. Axis is chosen to show $\sigma \in [0,25]$ and zoom in $\sigma \in [0,1]$. (c) sketches the curvature when $s(t)=1$ and $\sigma(t)=t$. It has small curvature, while the tangent directs to the datapoints.|&lt;/p&gt;
&lt;h3&gt;Results of deterministic sampling&lt;/h3&gt;
&lt;hr&gt;
&lt;h2&gt;Stochastic Sampling&lt;/h2&gt;
&lt;h3&gt;SDE formulation&lt;/h3&gt;
&lt;p&gt;&lt;strong&gt;EDM&lt;/strong&gt;[1] reformulates forward and backward SDE as a sum of the probability flow ODE and a varying-rate &lt;a href=&quot;https://en.wikipedia.org/wiki/Langevin_dynamics&quot;&gt;&lt;em&gt;Langevin diffusion&lt;/em&gt;&lt;/a&gt; SDE:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
d\mathrm{x}&lt;em&gt;{\pm} =
\underbrace{-\dot{\sigma}(t)\sigma(t) \nabla&lt;/em&gt;\mathrm{x} \log p(\mathrm{x};\sigma(t)) dt}&lt;em&gt;{\text{probability flow ODE}}
\pm \underbrace{\underbrace{\beta(t)\sigma(t)^2 \nabla&lt;/em&gt;\mathrm{x} \log p(\mathrm{x};\sigma(t)) dt}&lt;em&gt;{\text{deterministic noise decay}}
+ \underbrace{\sqrt{2\beta(t)}\sigma(t) d\mathrm{w}&lt;em&gt;t}&lt;/em&gt;{\text{noise injection}}}&lt;/em&gt;{\text{Langevin diffusion SDE}}
\end{align}
$$&lt;/p&gt;
&lt;h4&gt;Role of stochasticity&lt;/h4&gt;
&lt;p&gt;In theory, ODE and SDE have the same marginal distributions.
However in practice, stochasticity in sampling often enhances the sample quality.
The authors attribute the beneficial role of stochasticity to the following steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;$\mathrm{x}_t$ deviates from the ideal marginal distribution, due to the training and truncation errors.&lt;/li&gt;
&lt;li&gt;The &lt;em&gt;Langevin diffusion&lt;/em&gt; drives the sample towards the ideal marginal distribution.&lt;/li&gt;
&lt;/ol&gt;
&lt;h4&gt;Stochastic sampling algorithm in EDM&lt;/h4&gt;
&lt;p&gt;Stochastic sampling algorithm in &lt;strong&gt;EDM&lt;/strong&gt;[1] is executed in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Noise injection&lt;/strong&gt;: integrate noise into samples according to $\gamma_i \geq 0$.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Noise decay with probability flow&lt;/strong&gt;: solve the ODE from increased noise level to desired level.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;Algorithm in real world&lt;/h3&gt;
&lt;p&gt;||
|:--|
|Observe the effect of &lt;em&gt;Langevin diffusion&lt;/em&gt; in real world: there is gradual image degradation with the repeated addition and removal of noise. A random image is drawn from $p(\mathrm{x};\sigma)$ and Algorithm 2 is run for a certain number of steps with $\gamma_i=\sqrt{2}-1$.|&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Langevin diffusion&lt;/em&gt; is supposed to drive the sample towards the true data distribution, however...&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;For low noise levels, images drift toward &lt;strong&gt;oversaturated colors&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;For high noise levels, images become abstract when $s_{\text{noise}}=1$.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Authors suspect that &lt;strong&gt;non-conservative vector field&lt;/strong&gt; generated by parametrized denoiser &lt;strong&gt;violates the premises of Langevin diffusion&lt;/strong&gt; since their analytical denoisers have not shown such degradation.&lt;/p&gt;
&lt;p&gt;$\Rightarrow$ Fix flaws of $D_\theta(\mathrm{x};\sigma)$ with heuristics!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;For low noise levels, images drift toward &lt;strong&gt;oversaturated colors&lt;/strong&gt;.\
$\Rightarrow$ Enable stochasticity within $t_i \in [S_{\text{tmin}}, \underline{S_{\text{tmax}}}]$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For high noise levels, images become &lt;strong&gt;abstract&lt;/strong&gt; when $S_{\text{noise}}=1$. \
$\Rightarrow$ $D_\theta(\cdot)$ removes too much noise because of &lt;a href=&quot;https://en.wikipedia.org/wiki/Regression_toward_the_mean&quot;&gt;&lt;em&gt;regression towards the mean&lt;/em&gt;&lt;/a&gt;, which often happens when $\ell_2$ trained.\
$\Rightarrow$ Inflate the standard deviation of newly added noise: $S_{\text{noise}}&gt;1$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;New noise never exceeds the noise already in the image. \
$\Rightarrow$ Clamp $\gamma_i$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Controls the overal stochasticity by $S_{\text{churn}}$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Results of stochastic sampling&lt;/h3&gt;
&lt;p&gt;||
|:--|
|Evaluation of stochastic samplers with ablations. Red line is deterministic sampler while purple line is optimal stochastic sampler.|&lt;/p&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;p&gt;&lt;a href=&quot;https://arxiv.org/abs/2206.00364&quot;&gt;Tero Karras et al., &lt;em&gt;Elucidating the Design Space of Diffusion-Based Generative Models&lt;/em&gt; (NeurIPS, 2022)&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Yang Song et al., &lt;em&gt;Score-Based Generative Modeling through Stochastic Differential Equations&lt;/em&gt; (ICLR, 2021)&lt;/p&gt;
&lt;p&gt;Yaniv Benny and Lior Wolf, &lt;em&gt;Dynamic Dual-Output Diffusion Models&lt;/em&gt; (CVPR, 2022)&lt;/p&gt;
&lt;p&gt;Tero Karras et al., &lt;em&gt;Training generative adversarial networks with limited data&lt;/em&gt; (NeurIPS, 2020)&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item><item><title>Understanding Diffusion Models in Two Perspectives</title><link>https://junoh-kang.github.io/blog/understanding-diffusion-models-in-two-perspectives</link><guid isPermaLink="true">https://junoh-kang.github.io/blog/understanding-diffusion-models-in-two-perspectives</guid><description>DDPM and score-based SDEs as two routes to reverse diffusion, contrasting their shared structure and modeling differences.</description><pubDate>Sun, 03 Sep 2023 00:00:00 GMT</pubDate><content:encoded>&lt;h2&gt;Overview&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;DDPM&lt;/strong&gt;[1] and &lt;strong&gt;Score-Based Model&lt;/strong&gt;[2] introduce diffusion model as a new paradigm of generative models.
Since the concepts of both papers are similar, one might regard &lt;strong&gt;Score-Based Model&lt;/strong&gt;[2] as only a continuous version of &lt;strong&gt;DDPM&lt;/strong&gt;[1].
Two papers have slight different views, even their loss functions and implementations are similar.&lt;/p&gt;
&lt;p&gt;This post mainly explains how formulations and objectives of two papers are different, and how they are related even with the differences.&lt;/p&gt;
&lt;h4&gt;Summary of the post&lt;/h4&gt;
&lt;ol&gt;
&lt;li&gt;The objective of &lt;strong&gt;DDPM&lt;/strong&gt;[1] is to minimize the surrogate of the negative log-likelihood.&lt;/li&gt;
&lt;li&gt;The objective of &lt;strong&gt;Score-Based Model&lt;/strong&gt;[2] is to match marginal distributions of forward SDE and backward SDE/ODE.&lt;/li&gt;
&lt;li&gt;Even with the differences, both derivations require the score function, differential of the log of the probability density function. The score functions are parametrized by neural networks and both papaers have similar loss functions.&lt;/li&gt;
&lt;/ol&gt;
&lt;hr&gt;
&lt;h2&gt;Maximizing Log-Likelihood&lt;/h2&gt;
&lt;h3&gt;Forward (Diffusion) Process&lt;/h3&gt;
&lt;p&gt;The forward process is a Markov chain that gradually adds Gaussian noise to the data for $T$ steps with distributions defined as follows:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;\mathrm{x}&lt;em&gt;t \perp\mkern-9.5mu\perp \mathrm{x}&lt;/em&gt;{0:t-1}, \
&amp;#x26;q(\mathrm{x}&lt;em&gt;0) := \mathrm{P}&lt;/em&gt;{data}(\mathrm{x}_0), \
&amp;#x26;q(\mathrm{x}&lt;em&gt;t|\mathrm{x}&lt;/em&gt;{t-1}) := \mathcal{N}(\mathrm{x}&lt;em&gt;t;\sqrt{1-\beta_t}\mathrm{x}&lt;/em&gt;{t-1}, \beta_t \mathrm{I}),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where ${\beta_t}_{t=1}^T$ are pre-defined constants.&lt;/p&gt;
&lt;h3&gt;Backward (Denoising) Process&lt;/h3&gt;
&lt;p&gt;The backward process is a Markov chain that gradually denoises perturbed data and it is parametrized by neural networks.
When $\beta_t\ll 1$ the backward distribution can be approximated as&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
q(\mathrm{x}&lt;em&gt;{t-1}|\mathrm{x}&lt;/em&gt;{t})
\approx \mathcal{N}(\mathrm{x}&lt;em&gt;{t-1}; \cfrac{1}{ \sqrt{1-\beta_t}}(\mathrm{x}&lt;/em&gt;{t} + \beta_t \nabla \log q (\mathrm{x}_t)), \beta_t \mathrm{I}).
\end{align}
$$&lt;/p&gt;
&lt;p&gt;It is reasonable to parametrize the denoising distribution as Gaussian as long as ${\beta_t}_{t=1}^T$ are infinitesimal. Therefore the bacward process is defined as follows:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;\mathrm{x}&lt;em&gt;t \perp\mkern-9.5mu\perp \mathrm{x}&lt;/em&gt;{t+1:T}, \
&amp;#x26;p(\mathrm{x}&lt;em&gt;T) := \mathcal{N}(\mathrm{x}&lt;em&gt;T; \mathrm{0}, \mathrm{I}) \
&amp;#x26;p(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}&lt;/em&gt;{t})
= \mathcal{N}(\mathrm{x}&lt;em&gt;{t-1}; \cfrac{1}{ \sqrt{1-\beta_t}}(\mathrm{x}&lt;/em&gt;{t} + \beta_t s_\theta(\mathrm{x}_t,t)), \beta_t \mathrm{I}),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;Note that we expect $s_\theta(\mathrm{x}_t,t)$ to learn $\nabla\log q(\mathrm{x}_t)$.&lt;/p&gt;
&lt;h3&gt;Minimizing Surrogate of Negative Log-Likelihood&lt;/h3&gt;
&lt;p&gt;The negative log-likelihood of data is&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\mathbb{E}&lt;em&gt;{\mathrm{x}&lt;em&gt;0 \sim q} \left[-\log p(\mathrm{x}&lt;em&gt;0)\right]
&amp;#x26;\leq \mathbb{E}&lt;/em&gt;{\mathrm{x}&lt;em&gt;0 \sim q} \mathbb{E}&lt;/em&gt;{\mathrm{x}&lt;/em&gt;{1:T|0} \sim q} \left[ \log \cfrac{q(\mathrm{x}&lt;/em&gt;{1:T}|\mathrm{x}&lt;em&gt;{0})}{p(\mathrm{x}&lt;/em&gt;{0:T})} \right].
\end{align}
$$&lt;/p&gt;
&lt;p&gt;$$
\begin{align*}
-\log p(\mathrm{x}&lt;em&gt;0)
&amp;#x26;= -\log \int p(\mathrm{x}&lt;/em&gt;{0:T}) d\mathrm{x}&lt;em&gt;{1:T} \
&amp;#x26;= -\log \int q(\mathrm{x}&lt;/em&gt;{1:T}|\mathrm{x}&lt;em&gt;{0}) \cfrac{p(\mathrm{x}&lt;/em&gt;{0:T})}{q(\mathrm{x}&lt;em&gt;{1:T}|\mathrm{x}&lt;/em&gt;{0})} d\mathrm{x}&lt;em&gt;{1:T} \
&amp;#x26;\leq -\int q(\mathrm{x}&lt;/em&gt;{1:T}|\mathrm{x}&lt;em&gt;{0}) \log \cfrac{p(\mathrm{x}&lt;/em&gt;{0:T})}{q(\mathrm{x}&lt;em&gt;{1:T}|\mathrm{x}&lt;/em&gt;{0})} d\mathrm{x}&lt;em&gt;{1:T} ~~(\because \text{Jensen})\
&amp;#x26;= \mathbb{E}&lt;/em&gt;{\mathrm{x}&lt;em&gt;{1:T|0} \sim q} \left[ \log \cfrac{q(\mathrm{x}&lt;/em&gt;{1:T}|\mathrm{x}&lt;em&gt;{0})}{p(\mathrm{x}&lt;/em&gt;{0:T})} \right].
\end{align*}
$$&lt;/p&gt;
&lt;p&gt;Using Markov properties,&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;q(\mathrm{x}_{1:T} | \mathrm{x}_0)
= q(\mathrm{x}&lt;em&gt;T | \mathrm{x}&lt;em&gt;0) \prod&lt;/em&gt;{t=2}^T q(\mathrm{x}&lt;/em&gt;{t-1} | \mathrm{x}&lt;em&gt;t, \mathrm{x}&lt;em&gt;0), \
&amp;#x26;p(\mathrm{x}&lt;/em&gt;{T:0})
= p(\mathrm{x}&lt;em&gt;T) \prod&lt;/em&gt;{t=T}^{1} p(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}_t).
\end{align}
$$&lt;/p&gt;
&lt;p&gt;$$
\begin{align*}
q(\mathrm{x}&lt;em&gt;{1:T} | \mathrm{x}&lt;em&gt;0)
&amp;#x26;= \prod&lt;/em&gt;{t=1}^{T} q(\mathrm{x}&lt;/em&gt;{t}|\mathrm{x}&lt;em&gt;{0:t-1}) \&lt;br&gt;
&amp;#x26;= \prod&lt;/em&gt;{t=1}^{T} q(\mathrm{x}&lt;em&gt;{t}|\mathrm{x}&lt;/em&gt;{t-1}) \
&amp;#x26;= q(\mathrm{x}&lt;em&gt;1|\mathrm{x}&lt;em&gt;0)\prod&lt;/em&gt;{t=2}^{T} q(\mathrm{x}&lt;/em&gt;{t}|\mathrm{x}&lt;em&gt;{t-1}, \mathrm{x}&lt;/em&gt;{0}) \
&amp;#x26;= q(\mathrm{x}&lt;em&gt;1|\mathrm{x}&lt;em&gt;0)\prod&lt;/em&gt;{t=2}^{T} \frac{q(\mathrm{x}&lt;/em&gt;{t},\mathrm{x}&lt;em&gt;{t-1}| \mathrm{x}&lt;/em&gt;{0})}{q(\mathrm{x}&lt;em&gt;{t-1}| \mathrm{x}&lt;/em&gt;{0})} \
&amp;#x26;= q(\mathrm{x}&lt;em&gt;1|\mathrm{x}&lt;em&gt;0)\prod&lt;/em&gt;{t=2}^{T} \frac{q(\mathrm{x}&lt;/em&gt;{t}|\mathrm{x}&lt;em&gt;{0})q(\mathrm{x}&lt;/em&gt;{t-1}| \mathrm{x}&lt;em&gt;{t},\mathrm{x}&lt;/em&gt;{0})}{q(\mathrm{x}&lt;em&gt;{t-1}| \mathrm{x}&lt;/em&gt;{0})} \
&amp;#x26;= q(\mathrm{x}&lt;em&gt;T | \mathrm{x}&lt;em&gt;0) \prod&lt;/em&gt;{t=2}^T q(\mathrm{x}&lt;/em&gt;{t-1} | \mathrm{x}&lt;em&gt;t, \mathrm{x}&lt;em&gt;0), \
p(\mathrm{x}&lt;/em&gt;{T:0})
&amp;#x26;= p(\mathrm{x}&lt;em&gt;T) \prod&lt;/em&gt;{t=T}^{1} p(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}&lt;em&gt;{T:t})\
&amp;#x26;= p(\mathrm{x}&lt;em&gt;T) \prod&lt;/em&gt;{t=T}^{1} p(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}_t).
\end{align*}
$$&lt;/p&gt;
&lt;p&gt;Therefore, the surrogate of negative log-likelihood becomes&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;D_{KL}(q(\mathrm{x}_T|\mathrm{x}_0) || p(\mathrm{x}_T))
+ \mathbb{E}&lt;em&gt;q\left[-\log p(\mathrm{x}&lt;em&gt;0|\mathrm{x}&lt;em&gt;1)\right] \nonumber \
&amp;#x26;+ \sum&lt;/em&gt;{t=2}^T D&lt;/em&gt;{KL}(q(\mathrm{x}&lt;/em&gt;{t-1} | \mathrm{x}_t, \mathrm{x}&lt;em&gt;0) || p(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}_t)).
\end{align}
$$&lt;/p&gt;
&lt;p&gt;The surrogate of negative log-likelihood can be explictly expressed using&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;p(\mathrm{x}&lt;em&gt;{t-1}|\mathrm{x}&lt;/em&gt;{t}) = \mathcal{N}(\mathrm{x}&lt;em&gt;{t-1}; \cfrac{1}{ \sqrt{1-\beta_t}}(\mathrm{x}&lt;/em&gt;{t} + \beta_t s_\theta(\mathrm{x}&lt;em&gt;t,t)), \beta_t \mathrm{I}), \
&amp;#x26;q(\mathrm{x}&lt;/em&gt;{t-1}|\mathrm{x}&lt;em&gt;{t}, \mathrm{x}&lt;/em&gt;{0})
= \mathcal{N}(\mathrm{x}&lt;em&gt;{t-1}; \cfrac{1}{ \sqrt{1-\beta_t}}(\mathrm{x}&lt;/em&gt;{t} + \beta_t \nabla \log q (\mathrm{x}&lt;em&gt;t|\mathrm{x}&lt;/em&gt;{0})), \frac{1-\bar\alpha_{t-1}}{1-\bar\alpha_t}\beta_t \mathrm{I}),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $\bar\alpha_t = \prod_{s=1}^t (1-\beta_s)$.&lt;/p&gt;
&lt;p&gt;Finally, the objective function becomes&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\sum_{t=1}^T \mathbb{E}&lt;em&gt;{\mathrm{x}&lt;em&gt;0}\mathbb{E}&lt;/em&gt;{\mathrm{x}&lt;/em&gt;{t}|\mathrm{x}&lt;em&gt;{0}} \left[ \lambda_t ||s&lt;/em&gt;\theta(\mathrm{x}_t,t) - \nabla \log q(\mathrm{x}_t|\mathrm{x}_0)||_2^2 \right],
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $\lambda_t$ are some constants.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Matching Marginal Distributions&lt;/h2&gt;
&lt;h3&gt;Forward SDE&lt;/h3&gt;
&lt;p&gt;For pre-defined function $f:\mathbb{R}^{h\times w \times 3}\times \mathbb{R} \rightarrow \mathbb{R}^{h\times w \times 3}$ and $g:\mathbb{R} \rightarrow \mathbb{R}$, a forward SDE perturbs the data with Gaussian noise by&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
d\mathrm{x}_t = f(\mathrm{x}_t,t)dt + g(t)d\mathrm{w}_t, ~~\text{and}~~ \mathrm{x}&lt;em&gt;0 \sim \mathrm{P}&lt;/em&gt;{data},
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $\mathrm{w}_t$ is Brownian process.&lt;/p&gt;
&lt;p&gt;If ${\mathrm{x}&lt;em&gt;t}&lt;/em&gt;{t=0}^T$ is a solution of the forward SDE, it can be treated as a sample from the joint distribution ${p_t}_{t=0}^T$.
However, learning joint distribution is difficult and our interest is only $\mathrm{x}_0$, not ${\mathrm{x}&lt;em&gt;t}&lt;/em&gt;{t=0}^T$.
Therefore, it suffices to consider weakened objective, learning how marginal distributions evolve as $t$ changes.
The evolution of the marginal distributions is goverened by the &lt;strong&gt;Fokker-Plank equation&lt;/strong&gt;:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\partial_t p_t = - \nabla_x (f \cdot p_t ) + \frac{1}{2} \mathrm{tr}(g^T ~\nabla_x^2p_t~ g).
\end{align}
$$&lt;/p&gt;
&lt;h3&gt;Backward SDE/ODE&lt;/h3&gt;
&lt;p&gt;Following backward SDE and ODE are known to have the same marginal distributions:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
&amp;#x26;d\mathrm{x}_t = \left[ f(\mathrm{x}_t,t)dt - g^2(t) \nabla \log p_t(\mathrm{x}_t)  \right]dt + g(t)d\bar{\mathrm{w}}_t
, ~~\text{and}~~ \mathrm{x}_T \sim \mathcal{N}(\mathrm{0}, \mathrm{I}), \
&amp;#x26;d\mathrm{x}_t = \left[ f(\mathrm{x}_t,t)dt - \frac{1}{2} g^2(t) \nabla \log p_t(\mathrm{x}_t)  \right]dt
, ~~\text{and}~~ \mathrm{x}_T \sim \mathcal{N}(\mathrm{0}, \mathrm{I}),
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $\bar{\mathrm{w}}_t$ is the reverse-time Brownian motion.&lt;/p&gt;
&lt;p&gt;Since $f(\cdot, \cdot)$ and $g(\cdot)$ are known, the only unknown component in backward SDE/ODE is $\nabla \log p_t (\cdot)$ which is also known as a score function.
The score function is parametrized by neural network, $s_\theta(\mathrm{x}_t,t)$.&lt;/p&gt;
&lt;h3&gt;Learning Score Function&lt;/h3&gt;
&lt;p&gt;Since we parametrized the score function with the neural network, we can consider a loss function of&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\int_{0}^{T} \lambda_t \mathbb{E}_{\mathrm{x}&lt;em&gt;t} \left[ ||s&lt;/em&gt;\theta(\mathrm{x}_t,t) - \nabla \log p_t(\mathrm{x}_t)||_2^2 \right] dt,
\end{align}
$$&lt;/p&gt;
&lt;p&gt;where $\lambda_t$ are some constants.
Note that $\nabla\log p_t(\mathrm{x}_t)$ is intractable and with some tricks, the loss function changes into tractable form:&lt;/p&gt;
&lt;p&gt;$$
\begin{align}
\int_{0}^{T} \lambda_t \mathbb{E}&lt;em&gt;{\mathrm{x}&lt;em&gt;0}\mathbb{E}&lt;/em&gt;{\mathrm{x}&lt;/em&gt;{t}|\mathrm{x}&lt;em&gt;{0}} \left[ ||s&lt;/em&gt;\theta(\mathrm{x}&lt;em&gt;t,t) - \nabla \log p&lt;/em&gt;{t|0}(\mathrm{x}_t|\mathrm{x}_0)||_2^2 \right] dt.
\end{align}
$$&lt;/p&gt;
&lt;h2&gt;References&lt;/h2&gt;
&lt;p&gt;Jonathan Ho et al., &lt;em&gt;Denoising diffusion probabilistic models&lt;/em&gt; (NeurIPS, 2020)&lt;/p&gt;
&lt;p&gt;Yang Song et al., &lt;em&gt;Score-Based Generative Modeling through Stochastic Differential Equations&lt;/em&gt; (ICLR, 2021)&lt;/p&gt;</content:encoded><h:img src="undefined"/><enclosure url="undefined"/></item></channel></rss>