- +1.68 over CCDD
- +22.53 over masked diffusion
Hard puzzles require solving strategies never seen in training, which is where globally consistent decoding pays off. A 7× larger autoregressive model reaches only 32.57% here.
Augmenting Continuous Diffusion Language Models with Discrete Token Guidance
How HC-DLM samples · simulated example
Abstract
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded.
To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update.
On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B.
Method
Noise corrupts tokens and latent independently, so training stays simple. Generation is where the two meet: at every step the model reads a token draft kt out of the latent xt, and that draft guides the next latent update.
What is new
Discrete diffusion decodes tokens in parallel, but samples each one independently. Continuous diffusion plans all tokens in one shared latent, but never checks that plan against real tokens until the very end. HC-DLM keeps both strengths: tokens are planned jointly in the latent, and the latent is guided by tokens at every step.
| Discrete diffusione.g. MDM, LLaDA | Continuous diffusione.g. Diffusion-LM, Plaid | Hybrid discrete–continuouse.g. CADD, CCDD | HC-DLM (ours) | |
|---|---|---|---|---|
| Token dependence within a step | ✕independent marginals | ✓ | ◐via conditioning only | ✓ |
| Tied to tokens at every step | ✓ | ✕only at t = 0 | ✓ | ✓ |
| Tokens revisable at every step | ◐uniform kernel only | ✓ | ◐uniform kernel only | ✓readout from xt |
Results
From Sudoku and Countdown to open-domain text, HC-DLM leads discrete, continuous and hybrid diffusion baselines of the same size on Hard Sudoku, Countdown and LM1B. Parameter counts exclude token embeddings.
Hard puzzles require solving strategies never seen in training, which is where globally consistent decoding pays off. A 7× larger autoregressive model reaches only 32.57% here.
Best at the 6M scale on both CD4 and CD5, and the lead grows with the planning horizon: +3.2 over CCDD on CD4, +12.2 on CD5.
Lowest generative perplexity among all diffusion models evaluated, ahead of every discrete baseline and both continuous models that wait until the end to decode.
| Method | Params | Easy | Hard |
|---|---|---|---|
| Autoregressive | |||
| ARM (w/o ordering) | 42M | 9.73 | – |
| ARM (with ordering) | 42M | 87.18 | 32.57 |
| Discrete diffusion | |||
| MDM (vanilla) | 6M | 6.88 | 3.62 |
| MDM (top-prob.) | 6M | 18.51 | 9.44 |
| MDM (top-prob. margin) | 6M | 89.49 | 49.88 |
| Hybrid discrete–continuous | |||
| CCDD | 6M | 94.65 | 70.73 |
| HC-DLM (ours) | 6M | 94.21 | 72.41 |
| Method | Params | CD4 | CD5 |
|---|---|---|---|
| Autoregressive | |||
| GPT-2 Scratch | 6M | 31.9 | 4.3 |
| GPT-2 Scratch | 85M | 45.8 | 5.1 |
| GPT-2 Scratch | 303M | 41.3 | 4.5 |
| Stream-of-Search | 250M | 54.2 | – |
| LLaMA | 7B | 41.1 | 6.7 |
| LLaMA | 13B | 51.1 | 7.4 |
| Discrete diffusion | |||
| VDM | 85M | 73.4 | 16.3 |
| D3PM | 85M | 83.1 | 27.6 |
| RDM | 85M | 87.0 | 45.8 |
| MDM (vanilla) | 6M | 6.7 | 0.3 |
| MDM (top-prob.) | 6M | 47.4 | 21.4 |
| MDM (top-prob. margin) | 6M | 50.8 | 21.3 |
| Hybrid discrete–continuous | |||
| CCDD | 6M | 81.18 | 25.35 |
| HC-DLM (ours) | 6M | 84.41 | 37.52 |
| Method | Params | Gen. PPL ↓ |
|---|---|---|
| Autoregressive | ||
| Transformer | 108M | 66.7 |
| Discrete diffusion | ||
| MDM | 116M | 103.9 |
| SEDD | 116M | 115.9 |
| Duo | 116M | 97.6 |
| Continuous diffusion | ||
| Plaid | 109M | 77.3 |
| LangFlow | 117M | 92.2 |
| HC-DLM (ours) | 118M | 75.5 |
| Ground truth | – | 40.4 |
Autoregressive and masked-diffusion Sudoku rows follow Kim et al. (2025); non-MDM Countdown rows follow Ye et al. (2024); LM1B baselines are the models retrained by the LangFlow authors. MDM and CCDD on Sudoku and Countdown are reproduced at a matched 6M scale under the same protocol.
Ablations
Drop the token guidance and you are left with a plain latent diffusion model; drop the latent and you are left with a purely discrete chain. Both fall far behind the full model, most of all on Hard Sudoku, where global constraints bite. Tables and the first two plots report Sudoku accuracy (%).
| Method | Cont. latent | Token guidance | Easy | Hard |
|---|---|---|---|---|
| MDM (top-prob. margin) | ✕ | ✕ | 89.49 | 49.88 |
| Latent DM (w/o token guidance) | ✓ | ✕ | 50.46 | 24.74 |
| HC-DLM (ours) | ✓ | ✓ | 94.21 | 72.41 |
| Noise type | Token ordering | Easy | Hard |
|---|---|---|---|
| Absorbing-state | Random | 69.60 | 43.74 |
| Absorbing-state | Top-prob. | 74.72 | 47.50 |
| Absorbing-state | Top-prob. margin | 75.59 | 48.25 |
| Uniform-state | Parallel update | 94.21 | 72.41 |
Citation