Continuous Diffusion Language Models Make a Comeback

Continuous Diffusion Language Models (CDLM's)

Continuous Diffusion Language Models Make a Comeback

After years of dominance by autoregressive models and discrete diffusion, continuous diffusion for language is seeing a resurgence. This post explores the history, from early attempts in 2021-2022 to the 'continuous extinction' around late 2023, and explains why continuous methods are now returning. It covers key technical ingredients like embedding strategies, loss functions, noise schedules, and self-conditioning, and discusses recent research that may close the efficiency gap with autoregressive models.

The transition from 2023 to 2024 is quite stark!
  1. janalsncm

    > [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022

    I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.

    (The only exception I will make is encoder-decoder models which now are often done by decoder-only.)

    But what made it go mainstream was RL. RLHF at first, then other improvements like DPO that were less of a pain in the ass to set up. Adding diffusion on top of that would be an even bigger pain in the ass.

    Before ChatGPT there really wasn’t much of a concept of pre-training and post-training. It was all pre-training. Post training was what made the bots conversational and not just “continuing the thing you wrote to them”.

    So in short, diffusion never took off because it was just a more complicated way to generate tokens, and the real problem was getting tokens in the right distribution.

  2. 2001zhaozhao

    I would love to see models that can think at different rates and also output a thinking scratchpad alongside output text instead of before all output.

    Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead.

  3. p1esk

    It’s refreshing to read something not AI generated.

  4. vatsachak

    I feel like there is still low hanging fruit on the auto regressive LLMs; the encoder

  5. Marchant_hq

    CDLMs sound promising for smoother, more coherent text generation. Excited to see how they tackle the token-level discontinuities.

More from this day

2026-08-30