Seungjun Lee

NLP Bascis

RL Based Post-Training Methods(Policy Gradient, PPO, RLOO, and GRPO)

This blog summarizes the main RL methods used in LLM post-training and the notation commonly used in their objectives.

A useful high-level picture is:

Policy Gradient→REINFORCE + Baseline / RLOO→TRPO / PPO→GRPO: group-based baseline + PPO-style clipping\text{Policy Gradient} \rightarrow \text{REINFORCE + Baseline / RLOO} \rightarrow \text{TRPO / PPO} \rightarrow \text{GRPO: group-based baseline + PPO-style clipping}

The core idea across all of them is simple:

Generate outputs, score them, and update the model so that better outputs become more likely and worse outputs become less likely.


Notations

oio_i

The ii-th generated response.

Example:

Prompt: Solve 2 + 2

o_1 = "The answer is 4"
o_2 = "The answer is 5"

If we generate GG responses, then:

o1,o2,…,oGo_1, o_2, \dots, o_G

ai,ta_{i,t}

Token tt inside response ii.

For:

o_i = "The answer is 4"

the tokens could be:

["The", " answer", " is", " 4"]

∣oi∣|o_i|

The number of tokens in response ii.

si,ts_{i,t}

The context before generating token ai,ta_{i,t}.

It contains:

original prompt+previously generated tokens\text{original prompt} + \text{previously generated tokens}

For example, when generating "4" in:

Prompt: Solve 2 + 2
Response so far: The answer is

then si,ts_{i,t} is roughly:

Solve 2 + 2 + The answer is

RiR_i

The reward assigned to response oio_i.

Conceptually:

Ri=RewardFunction(prompt,oi)R_i = \text{RewardFunction}(\text{prompt}, o_i)

For example:

Correct answer   -> R = 1
Incorrect answer -> R = 0

The reward does not always have to come from a reward model. It can also come from a rule-based verifier, exact answer matching, test execution, etc.

πθ\pi_\theta

The current policy being trained.

For LLMs:

πθ=the current language model\pi_\theta = \text{the current language model}

The parameters θ\theta are updated during training.

πold\pi_{\text{old}}

A frozen snapshot of the policy before the current update.

It is used to measure how much the current policy has changed.

πref\pi_{\text{ref}}

A frozen reference model, often the original SFT model before RL training.

Its purpose is to keep the RL-trained policy from drifting too far away.

πθ=model currently changing\boxed{ \pi_\theta = \text{model currently changing} } πold=short-term snapshot\boxed{ \pi_{\text{old}} = \text{short-term snapshot} } πref=long-term reference model\boxed{ \pi_{\text{ref}} = \text{long-term reference model} }

ρi,t\rho_{i,t}

The policy probability ratio:

ρi,t=πθ(ai,t∣si,t)πold(ai,t∣si,t)\rho_{i,t} = \frac{ \pi_\theta(a_{i,t}\mid s_{i,t}) }{ \pi_{\text{old}}(a_{i,t}\mid s_{i,t}) }

It measures how much the probability of one specific token changed.

  • ρ=1\rho=1: probability did not change
  • ρ=1.5\rho=1.5: token became 50% more likely
  • ρ=0.8\rho=0.8: token became 20% less likely

DKLD_{KL}

KL divergence measures how different the old policy and the current policy are.

For a state ss, compare the action probabilities assigned by πold\pi_{\text{old}} and πθ\pi_\theta:

DKL(πold(⋅∣s)∥πθ(⋅∣s))=∑aπold(a∣s)log⁡πold(a∣s)πθ(a∣s)D_{KL} \left( \pi_{\text{old}}(\cdot\mid s) \parallel \pi_\theta(\cdot\mid s) \right) = \sum_a \pi_{\text{old}}(a\mid s) \log \frac{ \pi_{\text{old}}(a\mid s) }{ \pi_\theta(a\mid s) }

The sum considers every possible action aa. For an LLM, the actions are the possible next tokens. If πold\pi_{\text{old}} and πθ\pi_\theta assign similar probabilities, the KL divergence is small. If their probabilities differ substantially, it is larger.

ρ=change for one selected token/action\boxed{ \rho = \text{change for one selected token/action} } DKL=overall difference between policy distributions\boxed{ D_{KL} = \text{overall difference between policy distributions} }

μG\mu_G

The average reward in a group of generated responses.

Suppose:

R=[1,1,0,0]R=[1,1,0,0]

Then:

μG=1+1+0+04=0.5\mu_G = \frac{1+1+0+0}{4} = 0.5

So:

μG=average score of the group\boxed{ \mu_G = \text{average score of the group} }

σG\sigma_G

The standard deviation of rewards inside the group.

It measures how spread out the rewards are.

For a group containing GG rewards:

σG=1G∑i=1G(Ri−μG)2\sigma_G = \sqrt{ \frac{1}{G} \sum_{i=1}^{G} (R_i-\mu_G)^2 }

Compare:

[0.5,0.5,0.5,0.5][0.5,0.5,0.5,0.5]

and:

[1,1,0,0][1,1,0,0]

Both have mean 0.50.5, but the first group has no spread while the second has more variation.

σG=how different the rewards are from each other\boxed{ \sigma_G = \text{how different the rewards are from each other} }

AiA_i or A^i\hat A_i

The advantage.

It tells the model whether a response was better or worse than some baseline.

  • Ai>0A_i>0: increase the probability of the response
  • Ai<0A_i<0: decrease the probability of the response

In GRPO:

Ai=Ri−μGσGA_i = \frac{R_i-\mu_G}{\sigma_G}

sg⁡[x]\operatorname{sg}[x]

sg means stop-gradient.

The value of xx is used normally in the forward pass, but no gradient flows through it during backpropagation.

forward:
x --------> loss

backward:
x | STOP <-------- loss
sg⁡[x]=treat x as a constant during backpropagation\boxed{ \operatorname{sg}[x] = \text{treat }x\text{ as a constant during backpropagation} }

clip⁡(x,l,h)\operatorname{clip}(x,l,h)

Clipping makes sure that xx stays within the range from the lower bound ll to the upper bound hh:

  • If x<lx<l, set it to ll.
  • If x>hx>h, set it to hh.
  • If l≤x≤hl\le x\le h, leave it unchanged.

Equivalently:

clip⁡(x,l,h)={l,x<lx,l≤x≤hh,x>h\operatorname{clip}(x,l,h) = \begin{cases} l, & x<l \\ x, & l\le x\le h \\ h, & x>h \end{cases}

For example:

clip⁡(1.5,0.8,1.2)=1.2\operatorname{clip}(1.5,0.8,1.2)=1.2 clip⁡(0.6,0.8,1.2)=0.8\operatorname{clip}(0.6,0.8,1.2)=0.8 clip⁡(1.0,0.8,1.2)=1.0\operatorname{clip}(1.0,0.8,1.2)=1.0

PPO and GRPO use clipping to stop the policy from changing too much in one update.


Methods

1. Policy Gradient

Policy Gradient uses the simplest RL idea:

If an action leads to high reward, make it more likely next time.

For an LLM:

L=−Rilog⁡πθ(oi)L = -R_i \log \pi_\theta(o_i)

At token level:

L=−Ri∑tlog⁡πθ(ai,t∣si,t)L = -R_i \sum_t \log \pi_\theta(a_{i,t}\mid s_{i,t})

If Ri>0R_i>0, the generated response becomes more likely.

If Ri<0R_i<0, it becomes less likely.

reward×log probability\boxed{ \text{reward} \times \text{log probability} }

From this basic idea, later methods improve either how the reward is compared against a baseline or how much the policy is allowed to change in one update.

2. REINFORCE + Baseline

Plain Policy Gradient uses raw reward.

REINFORCE + baseline instead asks:

Was this reward better or worse than what we expected?

The objective becomes:

L=−(Ri−b)log⁡πθ(oi)L = -(R_i-b)\log \pi_\theta(o_i)

where bb is the baseline.

Example:

b=0.6b=0.6

If Ri=0.9R_i=0.9:

Ri−b=0.3R_i-b=0.3

The response was better than expected.

If Ri=0.2R_i=0.2:

Ri−b=−0.4R_i-b=-0.4

The response was worse than expected.

Policy Gradient: use R\boxed{ \text{Policy Gradient: use }R } REINFORCE + baseline: use R−b\boxed{ \text{REINFORCE + baseline: use }R-b }

The baseline reduces variance and makes training more stable.

RLOO follows this same idea, but calculates the baseline from a group of answers generated for the same prompt.

3. RLOO

RLOO stands for REINFORCE Leave-One-Out.

It is REINFORCE + baseline with a specific group-based baseline: each response is compared against the other responses generated for the same prompt.

A simplified main objective is:

LRLOO=−1G∑i=1GAi∑t=1∣oi∣log⁡πθ(ai,t∣si,t)L_{\text{RLOO}} = -\frac{1}{G} \sum_{i=1}^{G} A_i \sum_{t=1}^{|o_i|} \log \pi_\theta(a_{i,t}\mid s_{i,t})

where GG is the number of responses generated for one prompt, oio_i is response ii, and AiA_i is the advantage assigned to that response. Each AiA_i weights the log-probability of the response that produced it.

RLOO calculates AiA_i using the mean reward of the other responses as its baseline:

bi=1G−1∑j≠iRjb_i = \frac{1}{G-1} \sum_{j\ne i}R_j

where bib_i is the leave-one-out baseline for response ii. The corresponding advantage is:

Ai=Ri−biA_i=R_i-b_i

For example, suppose one prompt produces G=4G=4 responses with rewards:

R=[1.0,0.8,0.2,0.0]R=[1.0,0.8,0.2,0.0]

For response 1:

b1=0.8+0.2+0.03=0.333b_1 = \frac{0.8+0.2+0.0}{3} = 0.333

Then:

A1=1.0−0.333=0.667A_1 = 1.0-0.333 = 0.667

For response 3:

b3=1.0+0.8+0.03=0.6b_3 = \frac{1.0+0.8+0.0}{3} = 0.6

so:

A3=0.2−0.6=−0.4A_3 = 0.2-0.6 = -0.4

Therefore, AiA_i is applied only to response oio_i:

  • If Ai>0A_i>0, increase the probability of the tokens in oio_i.
  • If Ai<0A_i<0, decrease the probability of the tokens in oio_i.
  • If Ai≈0A_i\approx 0, make little or no update from oio_i.

The advantages are treated as fixed values during this policy update; the gradient changes πθ\pi_\theta, not the computed AiA_i values.

RLOO = compare each response against the other responses in the group\boxed{ \text{RLOO = compare each response against the other responses in the group} }

A major benefit is that it does not require a separate value model.

4. TRPO

TRPO stands for Trust Region Policy Optimization.

Like the previous methods, TRPO uses an advantage AA to decide which actions should become more or less likely. Its main additional concern is preventing the policy from changing too much in one update.

Its main idea is:

Improve the policy, but do not let it move too far in one update.

It approximately maximizes:

ρA\rho A

while enforcing:

DKL(πold∥πθ)≤δD_{KL} ( \pi_{\text{old}} \| \pi_\theta ) \le \delta

So TRPO has two jobs:

  1. Improve actions with positive advantage.
  2. Keep the new policy close to the old policy.

The problem is that enforcing the KL constraint is computationally complicated.

5. PPO

PPO stands for Proximal Policy Optimization.

It tries to get the same basic behavior as TRPO, but in a simpler way.

Instead of explicitly solving a constrained KL optimization problem, PPO clips the probability ratio.

The core objective is:

min⁡(ρA,  clip⁡(ρ,1−ϵ,1+ϵ)A)\min \left( \rho A,\; \operatorname{clip}(\rho,1-\epsilon,1+\epsilon)A \right)

For example, with:

ϵ=0.2\epsilon=0.2

the preferred range is roughly:

0.8≤ρ≤1.20.8 \le \rho \le 1.2

So PPO says:

Make good actions more likely and bad actions less likely, but do not change their probabilities too aggressively in one update.

TRPO: control change with an explicit KL constraint\boxed{ \text{TRPO: control change with an explicit KL constraint} } PPO: control change mainly through clipping\boxed{ \text{PPO: control change mainly through clipping} }

6. GRPO

GRPO stands for Group Relative Policy Optimization.

GRPO combines the group-based comparison idea used by RLOO with the policy-change control used by PPO.

A simplified main objective is:

LGRPO=−1G∑i=1G1∣oi∣∑tmin⁡(ρi,tAi,  clip⁡(ρi,t,1−ϵ,1+ϵ)Ai)L_{\text{GRPO}} = -\frac{1}{G} \sum_{i=1}^{G} \frac{1}{|o_i|} \sum_t \min \left( \rho_{i,t}A_i,\; \operatorname{clip} ( \rho_{i,t}, 1-\epsilon, 1+\epsilon ) A_i \right)

where:

  • GG is the number of responses generated for one prompt.
  • oio_i is response ii, and ∣oi∣|o_i| is its number of tokens.
  • ρi,t\rho_{i,t} is the probability ratio between πθ\pi_\theta and πold\pi_{\text{old}} for token tt.
  • ϵ\epsilon controls the clipping range [1−ϵ,1+ϵ][1-\epsilon,1+\epsilon].
  • AiA_i is the group-relative advantage for response ii.

For one prompt, GRPO generates a group of responses o1,o2,…,oGo_1,o_2,\dots,o_G and obtains their rewards R1,R2,…,RGR_1,R_2,\dots,R_G. It calculates the group mean and standard deviation:

μG=1G∑i=1GRi\mu_G = \frac{1}{G} \sum_{i=1}^{G}R_i σG=1G∑i=1G(Ri−μG)2\sigma_G = \sqrt{ \frac{1}{G} \sum_{i=1}^{G} (R_i-\mu_G)^2 }

The advantage in the main objective is then:

Ai=Ri−μGσGA_i = \frac{R_i-\mu_G}{\sigma_G}

If Ai>0A_i>0, response oio_i scored above the group average and should become more likely. If Ai<0A_i<0, it scored below the group average and should become less likely. The clipped probability ratio prevents either update from changing the policy too aggressively.

Often a KL penalty to the reference model is also included.

The important difference from traditional PPO is:

PPO: usually uses a critic/value model for advantage estimation\boxed{ \text{PPO: usually uses a critic/value model for advantage estimation} } GRPO: uses group-relative rewards instead\boxed{ \text{GRPO: uses group-relative rewards instead} }

This removes the need for a separate critic/value model.


Final Summary

Policy Gradient increases the probability of high-reward answers using L=−Rlog⁡πL=-R\log\pi; REINFORCE + baseline and RLOO improve it by comparing rewards against a baseline, TRPO and PPO additionally prevent the policy from changing too much at once, and GRPO combines a group-based baseline with PPO-style clipping.