Paper Summaries
DeepSeek-R1
Prior Knowledge (GRPO)
where
GRPO in simple terms: During training, the policy model generates (G) responses per prompt, and then each response's reward is measured. Following this, the model's weights are updated to increase the probability of generating better responses.
Here, (\pi_{\theta_{\text{old}}}) is the last snapshot of the model, (\pi_\theta) is the policy model being trained, and (\pi_{\mathrm{ref}}) is the original fixed model. (a_i) represents the response generated by the model, and (q) represents the prompt. Finally, (A_i) is the normalized reward for response (a_i).
Since large ratio values can degrade training quality, the ratio is clipped between
Additionally, to prevent the policy model's distribution from drifting too far from the original model, a KL divergence term is added.
The full GRPO objective can be written as
Paper Summary
In this paper, the authors introduced two types of models: DeepSeek-R1-Zero and DeepSeek-R1.
DeepSeek-R1-Zero
Training: Simply perform GRPO training using the following rewards.
Reward: These are rule-based rewards. They did not use outcome or process neural reward models.
- Accuracy rewards
- Format rewards
For mathematics, the main accuracy reward was basically:
They could do this because many math problems have deterministic answers that can be checked automatically.
They also had a format reward, for example, requiring the reasoning and final answer to follow the expected structure.
So roughly:
Importantly, they did not use a neural reward model to evaluate whether the intermediate reasoning was good. The RL only received signals such as:
Did you get the final answer right?
Did you follow the required format?
Yet longer reasoning, self-reflection, and behaviors such as “wait, let me rethink this” emerged during RL training. That emergence is one of the notable results of R1-Zero.
1. Reasoning capability increased

2. Thinking time increased by itself during RL training

The number of thinking tokens increased during RL training even though the researchers never explicitly trained the model to produce longer reasoning.
3. “Wait, wait. Wait. That’s an aha moment I can flag here.”

The model also learned behaviors such as rethinking and self-reflection using an anthropomorphic tone.
Drawback: readability issues and language mixing.
DeepSeek-R1
To address the drawbacks of DeepSeek-R1-Zero, DeepSeek-R1 was trained using the following four-stage pipeline:
Each stage has a different purpose.
1. First SFT — Teach Clean Reasoning
First, the model is fine-tuned on a relatively small set of high-quality, long Chain-of-Thought (CoT) examples.
The purpose is roughly:
“This is what good and readable reasoning should look like.”
This helps address some of the problems observed in R1-Zero, such as poor readability and language mixing.
2. First RL — Improve Reasoning Ability
Next, the model is trained mainly on reasoning tasks such as:
- Mathematics
- Coding
- Logic
using GRPO.
The rewards are mainly rule-based:
For example, in mathematics:
Correct final answer → positive reward
Incorrect final answer → low or no reward
They still do not use a neural reward model to evaluate the intermediate reasoning itself.
The main goal of this stage is to improve reasoning ability.
3. Second SFT — Add General Capabilities
After the first RL stage, the RL-trained model is used to generate a large amount of reasoning data.
High-quality outputs are selected using rejection sampling, resulting in approximately:
They additionally collect approximately:
including tasks such as:
- Writing
- Translation
- Factual QA
- General-purpose tasks
The model is then fine-tuned on this combined dataset.
The goal is to preserve the improved reasoning ability while restoring strong general-purpose capabilities.
4. Final RL — Reasoning + General Alignment
Finally, another RL stage is performed using two types of prompts.
For reasoning tasks:
These tasks often have answers that can be automatically verified.
For general tasks:
For example:
“Write a helpful explanation of photosynthesis.”
There is no simple rule such as
for evaluating such responses.
Therefore, learned reward models are used to evaluate properties such as helpfulness and safety.
Overall, the DeepSeek-R1 training pipeline can be summarized as: