NLP Bascis
How Positional Encoding Evolved in Transformers: From Sinusoidal Encoding to RoPE
I’m going to explain how the way we incorporate positional information when calculating attention has evolved: from the original fixed positional encoding, to relative positional encoding, and finally to RoPE (Rotary Positional Encoding).
original
We can define the process of calculating attention in the transformer as follows
- Obtain Q, K, and V by incorporating word embeddings, positional information, and the matrices. In the paper, they encapsulate step 1 as the function .
Here in , we add sinusoidal positional information to the word embeddings and then perform matrix multiplication with .
where each positional encoding is defined as follows:
- With the obtained Q, K, and V, calculate the attention scores as follows.
Since is and is , the ‘ transposed @ ’ can be written as follows:
relative -1
The authors of Self-Attention with Relative Position Representations (Shaw et al., 2018) came up with relative positional encoding, which measure the relative positional encoding. The main idea of relative pos encoding is this:
At that time, fixed positional encoding had a problem: regardless of context, we were training the transformer to treat position as having the same meaning all the time. Because of this, the model couldn’t easily generalize to longer sequences than those seen during training, since absolute positions beyond the training range were not encoded (whereas with relative encoding, it can!).
To handle this issue, they came up with relative positional encoding, which provides positional information by letting the transformer know how far apart two tokens are, rather than relying on their raw positions. This significantly improved the performance of LLMs and made it possible to generalize beyond the training range.
Here is more details on how they applied relative pos encoding.
-
replaced in third and fourth term into a trainable vector and .
=> To make the attention score independent of the query’s specific location. using two different learnable vectors , allows the model to learn two distinct types of positional bias -
replaced in second and fourth term into . => Because we wanted to apply relative position to the key
-
and replaced into tilda where its get matrix multiplied with positional embeddings => becomes an expert at processing content, becomes an expert at processing relative positions.
these changes resulted in
You can also understand in this way:
This is content-to-content attention.
It measures attention between the content at position and the content at position .
There is no positional information here, so nothing needs to change.
This is query-content to key-position attention.
Originally, it uses the key’s absolute position .
For relative position, we want the key position relative to the query:
So now the attention depends on the distance between the query and key.
This is query-position to key-content attention.
The problem is that is the query’s absolute position.
If we keep , the attention score still depends on whether the query is at position 5, 20, 100, etc.
So we remove that absolute-position dependence:
where is a learned vector shared across positions.
Now this term represents a global attention bias toward key content, independent of the query’s absolute position.
This is query-position to key-position attention.
We want this to depend only on relative position.
So:
and:
This gives:
So this term now represents attention bias based on relative distance , independent of the query’s absolute position.
The main idea is simply:
to make the key position relative to the query,
and
to remove the query’s absolute position from the attention score.
relative -2
Later, prople found that the last three terms in relative-1 didn’t contribute much when calculating attention, so researchers replace those three blocks with .
relative -3
Authors of the T5 paper found out that second and third term in the original actually don’t contribute much when calculating the attention, so they simply replaced thosr terms with , and replaced the fourth term with .
why replaec W with U?
Because Ws are Content Matrices. Their entire purpose is to learn the best way to project the token’s content or meaning ( and ) into the query and key spaces. They are optimized to work with semantic information. So they introduced new matrices U which wil be optimized for projecting positional information.
relative -4
In the original term, replace the all fixed positional encoding with relative positional encoding and then dropped the last term(position-position).
RoPE
<background of RoPE>To improve positional encoding, instead of adding positional information, RoPE rotates the word embedding multiplied with W matrices( and ) So that dot product between two adjacent word embeddings will get higher score and vice versa.
For simplicity, let’s assume the embedding dimension is just 2, so the word embedding for a token is . To mathematically describe rotation, we can treat this 2D vector as a complex number.
Btw Rotating a complex vector can indeed be expressed like this
Then, to compute the attention, we take the dot product between two rotated embeddings ( and ). Since the dot product between two complex vector and is , we do multiplication of and )
Now let’s scale this up: when the embedding dimension is larger than 2, we split the embedding into blocks. Each block is treated as a 2D vector and rotated separately. So for the full embedding, the rotation can be expressed like this:
where
And for computational efficiency, it can be simplified as follows:
Quick Final Summary of RoPE
In the RoPE(Rotatry Positional Encoding), we rotate word embeddings in following ways:
Divied each token’s word embeddings into emedding_dim / 2 number of pairs(each pair just two blocks). Then we rotate each of these pairs by
where is defined as
The value of , degrades fast as increases, as a result fornt part of word embeddings got rotated faster(high frequency) and rear part of word embeddings get rotated much slower(low frequency)? So Why did they designed RoPE in this way?
Front Part (High Frequencies) = Millimeter Markings
- The first few pairs of dimensions rotate very fast. A change in position from to causes a large change in their rotation. This is crucial for understanding local grammar, syntax, and phrasing (e.g., “New York” vs. “York New”).
Rear Part (Low Frequencies) = Meter & Kilometer Markings
- The dimensions toward the end rotate very, very slowly. A change in position from to causes almost no change in their rotation. It takes a large jump, like from to , to see a meaningful angular change.

