Seungjun Lee

Paper Summaries

DeepSeek-R1

Prior Knowledge (GRPO)

1G∑i=1G(πθ(ai∣q)πθold(ai∣q)⋅Ai)−βDKL(πθ ∥ πref)\frac{1}{G} \sum_{i=1}^{G} \left( \frac{\pi_\theta(a_i \mid q)} {\pi_{\theta_{\text{old}}}(a_i \mid q)} \cdot A_i \right) - \beta D_{\mathrm{KL}} \left( \pi_\theta \,\|\, \pi_{\mathrm{ref}} \right)

where

Ai=ri−mean⁡({r1,r2,…,rG})std⁡({r1,r2,…,rG}).A_i = \frac{ r_i - \operatorname{mean}\left(\{r_1,r_2,\dots,r_G\}\right) }{ \operatorname{std}\left(\{r_1,r_2,\dots,r_G\}\right) }.

GRPO in simple terms: During training, the policy model generates (G) responses per prompt, and then each response's reward is measured. Following this, the model's weights are updated to increase the probability of generating better responses.

Here, (\pi_{\theta_{\text{old}}}) is the last snapshot of the model, (\pi_\theta) is the policy model being trained, and (\pi_{\mathrm{ref}}) is the original fixed model. (a_i) represents the response generated by the model, and (q) represents the prompt. Finally, (A_i) is the normalized reward for response (a_i).

Since large ratio values can degrade training quality, the ratio is clipped between

[1−ϵ,  1+ϵ].[1-\epsilon,\;1+\epsilon].

Additionally, to prevent the policy model's distribution from drifting too far from the original model, a KL divergence term is added.

The full GRPO objective can be written as

JGRPO(θ)=E[q∼P(Q),{oi}i=1G∼πθold(O∣q)][1G∑i=1G(min⁡(πθ(oi∣q)πθold(oi∣q)Ai,  clip⁡(πθ(oi∣q)πθold(oi∣q),1−ϵ,1+ϵ)Ai)−βDKL(πθ ∥ πref))].\mathcal{J}_{\mathrm{GRPO}}(\theta) = \mathbb{E} \left[ q \sim P(Q), \left\{o_i\right\}_{i=1}^{G} \sim \pi_{\theta_{\mathrm{old}}}(O \mid q) \right] \left[ \frac{1}{G} \sum_{i=1}^{G} \left( \min \left( \frac{\pi_\theta(o_i \mid q)} {\pi_{\theta_{\mathrm{old}}}(o_i \mid q)} A_i, \; \operatorname{clip} \left( \frac{\pi_\theta(o_i \mid q)} {\pi_{\theta_{\mathrm{old}}}(o_i \mid q)}, 1-\epsilon, 1+\epsilon \right) A_i \right) - \beta D_{\mathrm{KL}} \left( \pi_\theta \,\|\, \pi_{\mathrm{ref}} \right) \right) \right].

Paper Summary

In this paper, the authors introduced two types of models: DeepSeek-R1-Zero and DeepSeek-R1.

DeepSeek-R1-Zero

Training: Simply perform GRPO training using the following rewards.

Reward: These are rule-based rewards. They did not use outcome or process neural reward models.

  • Accuracy rewards
  • Format rewards

For mathematics, the main accuracy reward was basically:

Raccuracy={1,if the final answer is correct0,otherwiseR_{\text{accuracy}} = \begin{cases} 1, & \text{if the final answer is correct} \\ 0, & \text{otherwise} \end{cases}

They could do this because many math problems have deterministic answers that can be checked automatically.

They also had a format reward, for example, requiring the reasoning and final answer to follow the expected structure.

So roughly:

R=Raccuracy+RformatR = R_{\text{accuracy}} + R_{\text{format}}

Importantly, they did not use a neural reward model to evaluate whether the intermediate reasoning was good. The RL only received signals such as:

Did you get the final answer right?
Did you follow the required format?

Yet longer reasoning, self-reflection, and behaviors such as “wait, let me rethink this” emerged during RL training. That emergence is one of the notable results of R1-Zero.

1. Reasoning capability increased

image

2. Thinking time increased by itself during RL training

image

The number of thinking tokens increased during RL training even though the researchers never explicitly trained the model to produce longer reasoning.

3. “Wait, wait. Wait. That’s an aha moment I can flag here.”

image

The model also learned behaviors such as rethinking and self-reflection using an anthropomorphic tone.

Drawback: readability issues and language mixing.

DeepSeek-R1

To address the drawbacks of DeepSeek-R1-Zero, DeepSeek-R1 was trained using the following four-stage pipeline:

SFT→RL→SFT→RL\boxed{ \text{SFT} \rightarrow \text{RL} \rightarrow \text{SFT} \rightarrow \text{RL} }

Each stage has a different purpose.

1. First SFT — Teach Clean Reasoning

First, the model is fine-tuned on a relatively small set of high-quality, long Chain-of-Thought (CoT) examples.

The purpose is roughly:

“This is what good and readable reasoning should look like.”

This helps address some of the problems observed in R1-Zero, such as poor readability and language mixing.

2. First RL — Improve Reasoning Ability

Next, the model is trained mainly on reasoning tasks such as:

  • Mathematics
  • Coding
  • Logic

using GRPO.

The rewards are mainly rule-based:

R=Raccuracy+Rformat+Rlanguage consistencyR = R_{\text{accuracy}} + R_{\text{format}} + R_{\text{language consistency}}

For example, in mathematics:

Correct final answer → positive reward
Incorrect final answer → low or no reward

They still do not use a neural reward model to evaluate the intermediate reasoning itself.

The main goal of this stage is to improve reasoning ability.

3. Second SFT — Add General Capabilities

After the first RL stage, the RL-trained model is used to generate a large amount of reasoning data.

High-quality outputs are selected using rejection sampling, resulting in approximately:

600K reasoning examples600\text{K reasoning examples}

They additionally collect approximately:

200K non-reasoning examples200\text{K non-reasoning examples}

including tasks such as:

  • Writing
  • Translation
  • Factual QA
  • General-purpose tasks

The model is then fine-tuned on this combined dataset.

The goal is to preserve the improved reasoning ability while restoring strong general-purpose capabilities.

4. Final RL — Reasoning + General Alignment

Finally, another RL stage is performed using two types of prompts.

For reasoning tasks:

Math / Coding / Logic→Rule-based rewards\text{Math / Coding / Logic} \rightarrow \text{Rule-based rewards}

These tasks often have answers that can be automatically verified.

For general tasks:

General QA / Writing / Safety / etc.→Reward Models\text{General QA / Writing / Safety / etc.} \rightarrow \text{Reward Models}

For example:

“Write a helpful explanation of photosynthesis.”

There is no simple rule such as

R={1,correct0,wrongR = \begin{cases} 1, & \text{correct} \\ 0, & \text{wrong} \end{cases}

for evaluating such responses.

Therefore, learned reward models are used to evaluate properties such as helpfulness and safety.

Overall, the DeepSeek-R1 training pipeline can be summarized as:

SFT1:teach clean reasoning styleRL1:improve reasoning abilitySFT2:add reasoning and general capabilitiesRL2:final reasoning and alignment training\boxed{ \begin{aligned} \text{SFT}_1 &: \text{teach clean reasoning style} \\ \text{RL}_1 &: \text{improve reasoning ability} \\ \text{SFT}_2 &: \text{add reasoning and general capabilities} \\ \text{RL}_2 &: \text{final reasoning and alignment training} \end{aligned} }