TPLA: Tensor Parallel Latent Attention For Efficient Disaggregated Prefill & Decode Inference - Takara TLDR

Multi-Head Latent Attention (MLA), introduced in DeepSeek-V2, compresses
key-value states into a low-rank latent vector, caching only this vector to
reduce memory. In tensor parallelism (TP), however, attention heads are
computed across multiple devices, and each device must load the full cache,
eroding the advantage of MLA over Grouped Query Attention (GQA). We propose
Tensor-Parallel Latent Attention (TPLA): a scheme that partitions both the
latent representation and each head’s input dimension across devices, performs
attention independently per shard, and then combines results with an
all-reduce. TPLA preserves the benefits of a compressed KV cache while
unlocking TP efficiency. Unlike Grouped Latent Attention (GLA), every head in
TPLA still leverages the full latent representation, maintaining stronger
representational capacity. TPLA is drop-in compatible with models pre-trained
using MLA: it supports MLA-style prefilling and enables efficient
tensor-parallel decoding without retraining. Applying simple orthogonal
transforms — e.g., the Hadamard transform or PCA — before TP slicing further
mitigates cross-shard interference, yielding minimal accuracy degradation. By
reducing the per-device KV cache for DeepSeek-V3 and Kimi-K2, we achieve 1.79x
and 1.93x speedups, respectively, at a 32K-token context length while
maintaining performance on commonsense and LongBench benchmarks. TPLA can be
implemented with FlashAttention-3, enabling practical end-to-end acceleration.

Source link

What's Hot

Tesla Partners With DeepSeek, ByteDance For Advanced In-Car AI Assistant | Electric Vehicles

End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning – Takara TLDR

Elon Musk’s xAI sues Apple and OpenAI, alleging anticompetitive collusion

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill \& Decode Inference – Takara TLDR

End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning – Takara TLDR

Do What? Teaching Vision-Language-Action Models to Reject the Impossible – Takara TLDR

InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles – Takara TLDR

Amy Sherald Speaks Out About Government Censorship at the Smithsonian

Dealers Living Like Collectors, Egypt’s Tourism and More: Morning Links

Mütter Museum in Philadelphia Announces New Policy for Human Remains

Inigo Philbrick, Art Dealer Convicted of Fraud, Appears in BBC Film

Tesla Partners With DeepSeek, ByteDance For Advanced In-Car AI Assistant | Electric Vehicles

End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning – Takara TLDR

Elon Musk’s xAI sues Apple and OpenAI, alleging anticompetitive collusion

What's Hot

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill \& Decode Inference – Takara TLDR

Related Posts

Subscribe to Updates