ReLOBGen: Replayable Limit Order Book Message Generation

1ECE, Seoul National University 2Department of Statistics, Purdue University 3IPAI, Seoul National University
Preprint, 2026
Closed-loop market simulation with an LOB message generator and post-processing

In closed-loop market simulation, the LOB simulator uses each generated message Mt to update the current resting-order set from 𝒪t−1 to 𝒪t, an operation we call replay. Existing generators may produce messages that cannot be replayed and therefore require post-processing: correction, or rejection and resampling. Post-processing can distort the distribution of replayed messages and increases inference cost. ReLOBGen instead generates replayable messages by construction.

100%
replayability in 500-message rollouts
without correction or rejection
2.7–3.6×
faster per replayed message
than the LOBS5 baseline
35–49%
lower distance to real distributions
than LOBS5, measured by LOB-Bench

Abstract

We propose ReLOBGen, a method for generating limit order book (LOB) messages that are replayable by construction. Replayability is required for closed-loop market simulation, yet existing LOB message generators may produce non-replayable raw messages, i.e., messages inconsistent with the current market state. These generators therefore rely on post-hoc correction or rejection followed by resampling, which may alter the replayed message distribution or increase inference cost. ReLOBGen instead ensures replayability during generation: it selects the referenced order from the resting orders in the current LOB and then generates the remaining message fields to be consistent with that order and the market state. For realistic reference selection, ReLOBGen samples from a learned distribution over eligible resting orders, efficiently computed from cached order representations and a context-dependent query. It then enforces the consistency of the remaining fields by masking out invalid tokens. Together, these components enable efficient generation of realistic messages without post-hoc correction or resampling. In 500-message rollouts, ReLOBGen achieves 100% replayability, improves market realism, particularly for top-of-book statistics and the relative prices of LOB messages, and provides a 2.7–3.6× speedup per replayed message over the LOBS5 baseline.

Method

Each message Mt = (Et, Rt, Xt) has three parts: an event, a reference order, and an event order. An add event inserts the event order Xt into the resting-order set 𝒪t−1 as a new resting order; any other event applies Xt to an existing resting order, the reference order Rt.

Mt is replayable if it satisfies both of the following conditions:

  • Reference eligibility A non-add event references an eligible resting order: Rt ∈ 𝒞t(Et, 𝒪t−1) ⊂ 𝒪t−1.
  • Event-order compatibility The event order is consistent with the event, the reference, and the current LOB.

ReLOBGen guarantees both by construction, generating each message as event → reference order → event order.

a Reference-first message representation

LOBS5 reference-last token order versus ReLOBGen reference-first token order

The reference order Rt comes before the event order Xt and carries its current remaining size, so it can constrain the fields generated after it.

b Resting-order selection

Query-key scoring over cached resting-order keys

The reference is selected from eligible resting orders by scoring cached order keys against a context query.

c Inference-time masking

Validity mask applied to logits before sampling an event-order token

Invalid token values are masked given the selected reference and the market state.

Replayability and runtime

Post-processing rates, replayability violations, and runtime. 1,000 rollouts per stock (GOOG, INTC), each with 500 replayed messages on the JAX-LOB simulator. Runtime is measured on an NVIDIA RTX A6000 GPU and covers generation, post-processing, and simulator updates. Aborted baseline rollouts are excluded; ReLOBGen needs no restarts. Lower is better.

Ticker Method Post-processing Replayability violations Runtime (ms)
CorrectionRejectionReferenceEvent-orderPer replayed msg.Per attempt
GOOGLOBS5 (ref.-last)13.5%25.7%39.1%0.4%602 ± 164447 ± 119
LOBS5 (ref.-first)11.7%27.6%39.5%0.1%591 ± 145426 ± 100
ReLOBGen0.0%0.0%0.0%0.0%227 ± 12227 ± 12
INTCLOBS5 (ref.-last)9.7%16.3%25.6%0.5%1216 ± 7261018 ± 606
LOBS5 (ref.-first)11.5%13.2%24.5%0.2%1058 ± 850921 ± 744
ReLOBGen0.0%0.0%0.0%0.0%341 ± 150341 ± 150
  • 100% replayable.
    • ReLOBGen needs no correction or rejection, whereas the LOBS5 baselines correct or reject 25–39% of raw generation attempts.
  • 2.7–3.6× faster per replayed message than LOBS5 (ref.-last).
    • Fewer forward passes: a non-add message takes 7–8 forward passes instead of 17, because the reference order is selected in one step rather than decoded token by token, and fields with only one valid value are filled without a forward pass.
    • No resampling: every attempt is replayed, so no computation is wasted on rejected attempts.

Rollout realism

LOB-Bench metric-group summary of unconditional L1 distances. Overall averages all metrics, and each group column averages its metrics. Ref.-first + uniform is also replayable but picks the reference uniformly from the eligible resting orders. Lower is better, and bold marks the lowest value.

Ticker Method Overall Metric group
StateTimesVolumesDepthsLevelsTrades
GOOGLOBS5 (ref.-last)0.270.350.160.250.400.210.23
Ref.-first + uniform0.240.310.200.220.230.200.27
ReLOBGen0.170.240.150.180.200.110.18
INTCLOBS5 (ref.-last)0.220.140.140.200.320.170.27
Ref.-first + uniform0.300.100.350.270.360.340.32
ReLOBGen0.110.050.210.160.070.040.15

Market impact and error accumulation (GOOG)

Legend: Real, ReLOBGen, LOBS5 (ref.-last)
Price response after market orders without an immediate mid-price change
Price response after market orders with an immediate mid-price change
L1 distance over successive 100-message intervals for ask volume
L1 distance over successive 100-message intervals for ask cancellation depth

Left: lagged price responses after market orders without (MO0) and with (MO1) an immediate mid-price change; LOBS5 responses stay largely flat, while ReLOBGen tracks the real data more closely. Right: L1 distance over successive 100-message intervals, where ReLOBGen shows lower error.

All 21 unconditional LOB-Bench metrics Per-metric unconditional L1 distances on GOOG and INTC

Per-metric unconditional L1 distances with 99% bootstrap confidence intervals.

  • Better realism than LOBS5.
    • Largest gains in top-of-book statistics (State) and in where orders are submitted and canceled relative to the book (Depths, Levels).
    • ReLOBGen reproduces how prices respond to market orders, which LOBS5 largely misses.
  • Better realism than uniform selection.
    • Replayability alone does not make rollouts realistic: which eligible order a message references matters.
    • Learned selection picks more realistic cancellation targets, improving all five cancel-related metrics.

BibTeX

@article{kang2026relobgen,
  title={ReLOBGen: Replayable Limit Order Book Message Generation},
  author={Kang, Junoh and Lee, Kiseop and Han, Bohyung},
  journal={arXiv preprint},
  year={2026}
}