Unified multimodal dLLM · few-step distillation

OMNI-DIFFUSION-DISTILL

Few-Step Distillation of Unified Multimodal Diffusion Large Language Models

1Simon Fraser University · 2Meta
*Work done during an internship at Meta. †Work done during full-time employment at Meta. Corresponding authors.
18.2×faster text-to-image
128 → 8 steps
34.9×faster image editing
128 → 8 steps
21.2×faster understanding
512 → 64 steps
0.828GenEval at 8 steps
teacher at 128 steps: 0.867

TL;DR One few-step student for every task of a unified discrete diffusion LLM. It generates an image in 8 steps instead of 128 and answers a question in 64 steps instead of 512, while staying close to the full-step teacher.

Overview of tasks supported by the few-step student
Motivation

Why fewer steps break a unified dLLM

Cutting decoding steps fails differently in each modality. Text tokens committed in the same step cannot see each other, so the output starts repeating. Images must absorb classifier-free guidance, and fitting the sharpened guided target collapses the student's entropy.

Repetition under parallel decoding and entropy collapse under guidance distillation
Left: decoding more text tokens in parallel increases repetition. Right: a student distilled toward the low-entropy guided target keeps losing entropy; entropy-matched guidance keeps it high.
01 · Acceleration

Few steps, real wall-clock speedups

Each square is one output token. The teacher commits a few tokens per step and runs extra guidance passes; the student commits many at once without guidance. Timing below is scaled from measured A100 latency.

Teacher
step 00.0 s
Omni-Diffusion-Distill
18.2× faster
step 00.0 s

Latency per output on one A100

Teacher (full-step)Omni-Diffusion-Distill

Understanding decodes block by block with an exact KV cache, so its speedup exceeds the 8× step reduction.

Benchmark scores across decoding budgets
Quality across decoding budgets. The student stays near the full-step teacher as steps shrink, while the teacher run at few steps degrades.
03 · Comparison

Against the teacher and other distillation methods

Same prompt and seed for every method. Distilled models use 8 image steps or 64 text steps; the teacher uses its full budget of 128 or 512 steps.

Benchmark overview

Radar chart of text-to-image benchmarks Radar chart of open-ended understanding benchmarks
04 · Method

Two stages, two modality-specific fixes

Two-stage distillation pipeline
Stage 1

Off-policy trajectory replay

The student learns to skip decoding steps by matching teacher predictions cached along offline trajectories, with no teacher forward pass during training.

Stage 2

On-policy refinement

The student is refined on states from its own rollouts, so it learns to recover from the errors it actually makes at inference.

Text

Pairwise collision penalty

Tokens committed in the same step cannot see each other. Penalizing the probability that two of them pick the same token cuts repetition loops from 74.4% to 7.7%.

Image

Entropy-matched guidance

Matching a guidance-sharpened target collapses the student's entropy. Rescaling the target to a fixed fraction of the conditional entropy keeps the student diverse and editable.

05 · Citation

BibTeX