Hierarchical Continuous Diffusion Language Models

Augmenting Continuous Diffusion Language Models with Discrete Token Guidance

How HC-DLM samples · simulated example

The latent plans, the tokens guide, at every step

The HC-DLM reverse process A continuous latent trajectory from x_T to x_0 runs along the bottom. At each step tokens k_t are read out of the latent x_t, re-noised, and condition the next latent transition to x_{t-1}. Only the latent persists from step to step. Latent space ··· ··· xT xt xt−1 x0 kT kt kt−1 k0 q(kt ∣ kt−1) q(xt ∣ xt−1) Read out pθ(kt ∣ xt) Feed back pφ(xt−1 ∣ xt, kt)
xt
kt
still noise at this step revised this step

Abstract

Two complementary gaps, closed together

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded.

To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update.

On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B.

Method

Continuous planning, discrete guidance

Noise corrupts tokens and latent independently, so training stays simple. Generation is where the two meet: at every step the model reads a token draft kt out of the latent xt, and that draft guides the next latent update.

$$p_{\theta,\phi}(k_{0:T},\,x_{0:T}) \;=\; p(x_T)\;p_\theta(k_0\mid x_0)\prod_{t=1}^{T} \underbrace{p_\theta(k_t\mid x_t)}_{\text{read out tokens}}\;\;\underbrace{p_\phi(x_{t-1}\mid x_t,\,k_t)}_{\substack{\text{token-guided}\\ \text{latent denoising}}}$$
Training pipeline. An encoder maps clean tokens k_0 to a latent x_0. Independent forward kernels produce the noisy pair (x_t, k_t). The token-conditioned denoiser predicts the clean latent (loss J_cont), the token predictor decodes x_0 to tokens (loss J_recon), and the encoder is regularized by an entropy term J_ent.
One principled objective, trained end to end. A single variational bound on the token likelihood splits into three terms: reconstruction for the token readout, token-guided flow matching for the latent denoiser, and an entropy term that keeps the encoder from collapsing.

What is new

Each level covers the other’s blind spot

Discrete diffusion decodes tokens in parallel, but samples each one independently. Continuous diffusion plans all tokens in one shared latent, but never checks that plan against real tokens until the very end. HC-DLM keeps both strengths: tokens are planned jointly in the latent, and the latent is guided by tokens at every step.

Four reverse-step diagrams. (a) Discrete diffusion: tokens k_t go to k_{t-1} through a learned predictor, with no latent. (b) Continuous diffusion: the latent x_t goes to x_{t-1} through p_phi(x_{t-1} | x_t); tokens are decoded only at t = 0. (c) Hybrid diffusion: a discrete chain and a continuous chain advance together through p_phi(k_{t-1}, x_{t-1} | x_t, k_t), each conditioning the other. (d) HC-DLM: tokens are read out of x_t by p_theta and condition the latent update p_phi(x_{t-1} | x_t, k_t).
One reverse step, four designs. Discrete diffusion updates tokens directly, one marginal at a time. Continuous diffusion denoises a latent that is blind to tokens and decodes only at the end. Hybrid diffusion keeps a self-contained discrete chain and attaches a continuous signal to it. In HC-DLM the two levels talk at every step: tokens are read out of the latent, then guide its next update.
Discrete diffusione.g. MDM, LLaDA Continuous diffusione.g. Diffusion-LM, Plaid Hybrid discrete–continuouse.g. CADD, CCDD HC-DLM (ours)
Token dependence within a step ✕independent marginals ✓ ◐via conditioning only ✓
Tied to tokens at every step ✓ ✕only at t = 0 ✓ ✓
Tokens revisable at every step ◐uniform kernel only ✓ ◐uniform kernel only ✓readout from xt

Results

Gains on reasoning, planning and language

From Sudoku and Countdown to open-domain text, HC-DLM leads discrete, continuous and hybrid diffusion baselines of the same size on Hard Sudoku, Countdown and LM1B. Parameter counts exclude token embeddings.

Sudoku · Reasoning
72.41%Hard split accuracy · 6M params ↑ higher is better
  • +1.68 over CCDD
  • +22.53 over masked diffusion

Hard Sudoku accuracy (%)

HC-DLM72.41
CCDD70.73
MDM (top-prob. margin)49.88
Latent DM (no token guidance)24.74

Hard puzzles require solving strategies never seen in training, which is where globally consistent decoding pays off. A 7× larger autoregressive model reaches only 32.57% here.

Countdown · Planning
37.52%CD5 accuracy · 6M params ↑ higher is better
  • +12.17 over CCDD
  • +16.22 over masked diffusion

Countdown CD5 accuracy (%)

HC-DLM37.52
CCDD25.35
MDM (top-prob.)21.4
MDM (top-prob. margin)21.3

Best at the 6M scale on both CD4 and CD5, and the lead grows with the planning horizon: +3.2 over CCDD on CD4, +12.2 on CD5.

LM1B · Language
75.5Generative perplexity · 118M params ↓ lower is better
  • 1.8 lower than Plaid
  • 16.7 lower than LangFlow

Generative perplexity

HC-DLM75.5
Plaid77.3
LangFlow92.2
Duo97.6
MDM103.9

Lowest generative perplexity among all diffusion models evaluated, ahead of every discrete baseline and both continuous models that wait until the end to decode.

Full result tables
Sudoku. Accuracy (%) on the Easy and Hard splits.
MethodParamsEasyHard
Autoregressive
ARM (w/o ordering)42M9.73–
ARM (with ordering)42M87.1832.57
Discrete diffusion
MDM (vanilla)6M6.883.62
MDM (top-prob.)6M18.519.44
MDM (top-prob. margin)6M89.4949.88
Hybrid discrete–continuous
CCDD6M94.6570.73
HC-DLM (ours)6M94.2172.41
Countdown. Accuracy (%) with four (CD4) or five (CD5) numbers. Best overall in bold, best at 6M underlined.
MethodParamsCD4CD5
Autoregressive
GPT-2 Scratch6M31.94.3
GPT-2 Scratch85M45.85.1
GPT-2 Scratch303M41.34.5
Stream-of-Search250M54.2–
LLaMA7B41.16.7
LLaMA13B51.17.4
Discrete diffusion
VDM85M73.416.3
D3PM85M83.127.6
RDM85M87.045.8
MDM (vanilla)6M6.70.3
MDM (top-prob.)6M47.421.4
MDM (top-prob. margin)6M50.821.3
Hybrid discrete–continuous
CCDD6M81.1825.35
HC-DLM (ours)6M84.4137.52
LM1B. Unconditional generation, sequence length 128, generative perplexity under GPT-2 Large.
MethodParamsGen. PPL ↓
Autoregressive
Transformer108M66.7
Discrete diffusion
MDM116M103.9
SEDD116M115.9
Duo116M97.6
Continuous diffusion
Plaid109M77.3
LangFlow117M92.2
HC-DLM (ours)118M75.5
Ground truth–40.4

Autoregressive and masked-diffusion Sudoku rows follow Kim et al. (2025); non-MDM Countdown rows follow Ye et al. (2024); LM1B baselines are the models retrained by the LangFlow authors. MDM and CCDD on Sudoku and Countdown are reproduced at a matched 6M scale under the same protocol.

Ablations

Neither level works alone

Drop the token guidance and you are left with a plain latent diffusion model; drop the latent and you are left with a purely discrete chain. Both fall far behind the full model, most of all on Hard Sudoku, where global constraints bite. Tables and the first two plots report Sudoku accuracy (%).

Component ablation. Latent DM removes token guidance from the denoiser; MDM removes the latent altogether.
MethodCont. latentToken guidanceEasyHard
MDM (top-prob. margin)✕✕89.4949.88
Latent DM (w/o token guidance)✓✕50.4624.74
HC-DLM (ours)✓✓94.2172.41
Discrete forward kernel. A uniformly corrupted sequence is still a complete hypothesis; masks carry no information for the scaffold.
Noise typeToken orderingEasyHard
Absorbing-stateRandom69.6043.74
Absorbing-stateTop-prob.74.7247.50
Absorbing-stateTop-prob. margin75.5948.25
Uniform-stateParallel update94.2172.41
Line chart of Sudoku accuracy of the decoded intermediate prediction against normalized step t/T. HC-DLM rises from about 20% to above 90% by t/T = 0.6; the latent diffusion model without discrete conditioning rises late and plateaus near 57%.
Solutions emerge early. Accuracy of the intermediate prediction x̂0 decoded at each step.
Line chart of Hard Sudoku accuracy against the number of denoising steps from 10 to 100. HC-DLM stays between 52% and 76%; MDM falls from 56% to 20% as steps decrease.
Holds up with fewer steps. Hard Sudoku accuracy against step count; MDM degrades sharply as steps shrink.
Line chart of LM1B generative perplexity and entropy against the classifier-free guidance scale w from 1 to 5. Perplexity falls from about 129 at w = 1 to a minimum near 71 around w = 3.5 and rises slightly after; entropy stays between 4.0 and 4.3 throughout.
Tunable token guidance. Stronger classifier-free guidance sharply cuts LM1B perplexity while entropy holds steady.

Citation

BibTeX