NLP Bascis
MQA, GQA, MLA
MHA already works well, but people have come up with several ways to reduce the computational and memory cost of attention. Some of the major approaches are MQA, GQA, and MLA.
MQA and GQA share a similar idea: use fewer (W_K) and (W_V) projections. In other words, instead of letting every attention head have its own K and V, multiple heads share them.
The main idea of MLA is to let every head to have its own K and V, but stores it in a compressed latent representation.
In this post, MQA, GQA, and MLA will be explained.
MQA

Multi-Query Attention means that every head has its own Q, while all heads share a single K and V.
GQA

Every head has its own Q, while K and V are shared within groups of heads. For example, every 4 heads share one K and V pair. So if there are total 8 heads, there are being only 2Ks and 2 Vs.
MLA

Multi-Head Latent Attention (MLA) takes a different approach.
Instead of simply reducing the number of K and V representations, MLA keeps the information needed to construct them in a compressed latent representation.
In the DeepSeek paper, MLA is introduced to address a natural concern:
“MQA and GQA reduce the number of K and V representations, but wouldn’t that hurt the Transformer’s performance?”
To mitigate this, MLA works roughly as follows.
Like standard MHA, each head can still obtain its own Q, K, and V. However, instead of storing the full K and V tensors in the KV cache during inference, MLA stores a compressed latent representation:
where (W_{DKV}) is the down-projection matrix.
When K and V are needed, they can conceptually be projected back to their original dimensions:
where (W_{UK}) and (W_{UV}) are the up-projection matrices.
By storing (C_{KV}) instead of the full K and V tensors, MLA significantly reduces the KV-cache memory footprint and memory I/O during inference.
Here, one question naturally comes up:
“Wait, doesn’t MLA actually take more time because it has to scale K and V down and then scale them back up?”
The DeepSeek authors handle this by avoiding the explicit reconstruction of the full K and V tensors during attention.
Standard Attention
For standard attention:
For one head, suppose:
Then:
So the time complexity is:
where (L) is the sequence length.
MLA
Multi-Head Latent Attention (MLA) takes a different approach.
Instead of simply reducing the number of (K) and (V) representations, MLA stores the information needed to construct them in a compressed latent representation.
In the DeepSeek paper, MLA is introduced to address a natural concern:
“MQA and GQA reduce the number of (K) and (V) representations, but wouldn’t that hurt the Transformer’s performance?”
To mitigate this, MLA works roughly as follows.
Like standard MHA, each head can still obtain its own (Q), (K), and (V). However, instead of storing the full (K) and (V) tensors in the KV cache during inference, MLA stores a compressed latent representation:
where (W_{DKV}) is the down-projection matrix.
When (K) and (V) are needed, they can conceptually be projected back to their original dimensions:
where (W_{UK}) and (W_{UV}) are the up-projection matrices.
By storing (C_{KV}) instead of the full (K) and (V) tensors, MLA significantly reduces the KV-cache memory footprint and memory I/O during inference.
Here, one question naturally comes up:
“Wait, doesn’t MLA actually take more time because it has to scale (K) and (V) down and then scale them back up?”
The DeepSeek authors handle this by avoiding the explicit reconstruction of the full (K) and (V) tensors during attention.
Standard Attention
For standard attention:
For one head, suppose:
and
Then:
So the time complexity is:
where (L) is the sequence length.
MLA
In MLA, (K) can be written as:
Therefore:
Using the transpose rule:
so:
Instead of first reconstructing the full (K) tensor, MLA can calculate:
The first multiplication is:
where:
and
Therefore:
with time complexity:
The second multiplication is:
with:
and time complexity:
So the total arithmetic cost in this simplified per-head view is:
The main advantage of MLA, however, is not simply that its FLOP complexity is always lower than standard MHA. The main benefit is that it does not need to store and repeatedly read the full (K) and (V) tensors from the KV cache.
Instead, it stores the much smaller (C_{KV}), while the up-projection matrices can be absorbed into the surrounding attention computation.
This significantly reduces KV-cache memory usage and memory I/O during autoregressive inference.