Paper page - dKV-Cache: The Cache for Diffusion Language Models

A KV-cache-like mechanism, delayed KV-Cache, accelerates diffusion language models’ inference without significantly degrading performance.

Diffusion Language Models (DLMs) have been seen as a promising competitor for
autoregressive language models. However, diffusion language models have long
been constrained by slow inference. A core challenge is that their
non-autoregressive architecture and bidirectional attention preclude the
key-value cache that accelerates decoding. We address this bottleneck by
proposing a KV-cache-like mechanism, delayed KV-Cache, for the denoising
process of DLMs. Our approach is motivated by the observation that different
tokens have distinct representation dynamics throughout the diffusion process.
Accordingly, we propose a delayed and conditioned caching strategy for key and
value states. We design two complementary variants to cache key and value
step-by-step: (1) dKV-Cache-Decode, which provides almost lossless
acceleration, and even improves performance on long sequences, suggesting that
existing DLMs may under-utilise contextual information during inference. (2)
dKV-Cache-Greedy, which has aggressive caching with reduced lifespan, achieving
higher speed-ups with quadratic time complexity at the cost of some performance
degradation. dKV-Cache, in final, achieves from 2-10x speedup in inference,
largely narrowing the gap between ARs and DLMs. We evaluate our dKV-Cache on
several benchmarks, delivering acceleration across general language
understanding, mathematical, and code-generation benchmarks. Experiments
demonstrate that cache can also be used in DLMs, even in a training-free manner
from current DLMs.

Source link

What's Hot

Paper page – EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

Stability AI is working on a licensing marketplace for creators

Alibaba’s Qwen-MT Promises Smarter, Cheaper Translations Across 92 Languages

Paper page – dKV-Cache: The Cache for Diffusion Language Models

Paper page – EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

Paper page – Hierarchical Budget Policy Optimization for Adaptive Reasoning

Paper page – DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

Artist Loses Final Appeal in Case of Apologising for ‘Fishrot Scandal’

US Appeals Court Overturns $8.8 M. Trademark Judgement For Yuga Labs

Old Masters ‘Making a Comeback’ in London: Morning Links

Bill Proposed To Apply Anti-Money Laundering Regulations to Art Market

Paper page – EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

Stability AI is working on a licensing marketplace for creators

Alibaba’s Qwen-MT Promises Smarter, Cheaper Translations Across 92 Languages

What's Hot

Paper page – dKV-Cache: The Cache for Diffusion Language Models

Related Posts

Subscribe to Updates