Seungjun Lee

Paper Summaries

s1

TL;DR With just carefully selected 1k dataset do SFT and then do the test time scaling(just controlling thinking budget here). result was good!

Overview

OpenAI's o1 showed that language models reason better when given more compute at inference time, but it didn't reveal how, and no open replication had reproduced a clean test-time scaling curve. This paper finds a very simple recipe that does. First, fine-tune Qwen2.5-32B-Instruct on just 1,000 carefully selected reasoning examples (s1K). Second, control how long the model thinks at inference with budget forcing: cut its thinking off to shorten it, or block it from stopping and append "Wait" to lengthen it. The resulting model, s1-32B, beats o1-preview on competition math by up to 27% (relative, on AIME24), scales smoothly with more thinking tokens, and goes from 50% to 57% on AIME24 through forced extra thinking alone. Training took 26 minutes on 16 H100s.

1. Problem & Motivation

o1 proved that test-time scaling works, but the field lacked an open and simple way to reproduce it, which is the gap this paper targets.

1.1 What Test-Time Scaling Is

Most progress in language models has come from scaling train-time compute through large-scale pretraining. Test-time scaling is a newer paradigm built on top of those pretrained models: spend extra compute at inference, for example by letting the model think longer, to get better answers. OpenAI's o1 validated this idea, showing consistent accuracy gains as test-time compute increased. OpenAI described its approach as large-scale reinforcement learning, which implies substantial data and training effort.

1.2 Gap in Prior Work

Many groups tried to replicate o1 using techniques such as Monte Carlo Tree Search and multi-agent setups. DeepSeek R1 matched o1-level performance, but did so with reinforcement learning over millions of samples and several training stages. Despite all these efforts, none had openly reproduced a clear test-time scaling curve, meaning accuracy that reliably rises as you allow more thinking.

1.3 Research Question

What is the simplest approach that achieves both strong reasoning performance and test-time scaling?

2. Method

The recipe has two parts: a tiny, carefully curated training set, and a decoding-time trick that controls how long the model thinks.

2.1 Data: Building s1K

The authors assemble a large pool of hard questions with reasoning traces, then filter it down to 1,000 examples that are simultaneously high-quality, difficult, and diverse.

2.1.1 Initial Pool

They collect 59,029 questions from 16 sources, choosing datasets for quality, difficulty, and diversity. The largest source is NuminaMATH (30,660 math problems), followed by MATH, OlympicArena (olympiad questions across astronomy, biology, chemistry, computer science, geography, math, and physics), OmniMath, AGIEval (SAT and LSAT questions in English, law, and logic), and historical AIME problems from 1983 to 2021. They also create two new datasets. s1-prob has 182 problems from the probability section of Stanford's Statistics PhD qualifying exams, with handwritten solutions. s1-teasers has 23 hard brain-teasers from quantitative trading interviews.

For every question, they generate a reasoning trace and a solution using the Gemini 2.0 Flash Thinking API, producing triplets of question, trace, and solution. They then deduplicate the pool and decontaminate it against the evaluation benchmarks (MATH500, GPQA Diamond, AIME24) by removing questions with 8-gram overlap. AIME problems from 2022 and 2023 are also excluded because they were used during development.

2.1.2 Three-Stage Filtering

The pool shrinks through three stages, one per guiding principle:

59,029→API errors54,116→formatting51,581→difficulty24,496→diversity sampling1,00059{,}029 \xrightarrow{\text{API errors}} 54{,}116 \xrightarrow{\text{formatting}} 51{,}581 \xrightarrow{\text{difficulty}} 24{,}496 \xrightarrow{\text{diversity sampling}} 1{,}000

Quality. Questions that hit API errors are removed. So are samples with formatting problems, such as ASCII-art diagrams, references to images that don't exist, or inconsistent question numbering. From this cleaner pool, 384 samples from datasets the authors consider especially reliable are placed directly into the final set.

Difficulty. Two signals are used. The first is model performance: each question is attempted by Qwen2.5-7B-Instruct and Qwen2.5-32B-Instruct, with Claude 3.5 Sonnet grading each attempt against the reference solution. Any question that either model solves is dropped as likely too easy; using two models reduces the chance that an easy question slips through because of one model's random mistake. The second signal is reasoning trace length, on the assumption that harder problems need more thinking tokens. Long traces are therefore favored, not removed.

Diversity. Claude 3.5 Sonnet classifies every question into a domain using the American Mathematical Society's Mathematics Subject Classification, which covers math topics plus sciences such as physics, biology, and economics. Final sampling then spreads the 1,000 examples across about 50 domains, led by geometry (109), number theory (98), and combinatorics (75).

Notably, traces are not required to be correct. The focus is on the reasoning process rather than a perfect solution, and only 53.6% of s1K traces are graded correct.

2.1.3 Diversity Sampling Algorithm

The final selection works like a repeated two-step lottery: first pick a topic at random so every topic gets a fair turn, then pick a question from that topic in a way that strongly favors longer, and presumably harder, reasoning traces. The loop starts from the 384 samples already seeded from trusted datasets and fills the remaining 616 spots.

Step 1: Pick a domain uniformly at random. Every domain has an equal chance of being chosen, regardless of how many questions it contains. Without this step, large domains such as geometry would dominate the dataset, since they have far more questions than small ones such as astronomy. This step is what enforces diversity.

Step 2: Pick a question from that domain, weighted by trace length. The domain's remaining questions are ranked by reasoning trace length, with the longest at rank 1. Question qiq_i is then sampled with probability

P(qi)=2−ri∑j2−rjP(q_i) = \frac{2^{-r_i}}{\sum_j 2^{-r_j}}

where rir_i is the rank of qiq_i and the sum in the denominator runs over every question jj remaining in the domain. The numerator is the question's weight: since 2−r=1/2r2^{-r} = 1/2^r, the weights are 0.5, 0.25, 0.125, and so on, halving at each step down in rank. The denominator adds up all the weights in the domain, and dividing by it turns the weights into probabilities that sum to exactly 1. This step pushes toward difficulty, on the assumption that longer reasoning signals a harder problem.

For example, in a domain with four remaining questions:

Rank rrWeight 2−r2^{-r}Probability (weight ÷ 0.9375)
1 (longest)0.5≈ 53%
20.25≈ 27%
30.125≈ 13%
4 (shortest)0.0625≈ 7%

An intuitive way to read this is as lottery tickets: the longest question holds 8 tickets, the next 4, then 2, then 1. The longest question is the most likely pick, but shorter ones keep some chance, so the selection isn't completely rigid.

Step 3: Remove and repeat. The selected question is removed from the pool so it can't be picked twice, and any domain that runs out of questions is dropped from the domain list. The loop then returns to Step 1 until the dataset reaches 1,000 examples.

A note on notation: the paper itself writes only the weights, 2−rank2^{-\text{rank}}, and describes the process as weighted sampling. The explicit division by the total weight shown above is what weighted sampling does behind the scenes, written out to show the actual probabilities.

2.2 Training

s1-32B is produced by plain supervised fine-tuning with basic settings and very little compute.

2.2.1 Base Model & Format

The base model is Qwen2.5-32B-Instruct, which on math tasks generally matches or beats the larger Qwen2.5-72B-Instruct. Special delimiters separate the thinking phase from the answer phase: thinking is enclosed between <|im_start|>think and <|im_start|>answer.

2.2.2 Hyperparameters

Training runs for 5 epochs with a batch size of 16, for 315 gradient steps total, in bfloat16. The learning rate is 1×10−51 \times 10^{-5}, warmed up linearly over the first 5% of steps and then decayed to 0 with a cosine schedule. The optimizer is AdamW with β1=0.9\beta_1 = 0.9, β2=0.95\beta_2 = 0.95, and weight decay 1×10−41 \times 10^{-4}. Loss is computed only on reasoning traces and solutions, not on the questions. The sequence length is set long enough that no training sample is truncated, a choice that turns out to matter (Section 5.4).

2.2.3 Cost

Fine-tuning takes 26 minutes on 16 NVIDIA H100 GPUs, about 7 GPU-hours in total.

2.3 Budget Forcing

Budget forcing controls how long the model thinks by editing its output during generation, with no additional training.

2.3.1 Core Idea

The authors split test-time scaling methods into two families. Sequential methods let later computation build on earlier computation, as in one long chain of thought. Parallel methods run independent attempts and combine them, as in majority voting. The paper focuses on sequential scaling, reasoning that building on intermediate results should allow deeper reasoning and iterative refinement.

Normally the model itself decides when to stop thinking, by emitting the end-of-thinking delimiter. Budget forcing takes that decision away from the model.

2.3.2 Capping Thinking (Maximum Budget)

Thinking tokens are counted as they are generated. When the count reaches the limit, the end-of-thinking delimiter is inserted, optionally followed by "Final Answer:". The model is forced to switch to answering with whatever it has worked out so far.

2.3.3 Extending Thinking (Minimum Budget)

When the model tries to emit the end-of-thinking delimiter too early, that token is suppressed and the word "Wait" is appended to its reasoning instead. From the model's perspective, it has just written "Wait," so its natural continuation is to reconsider its conclusion. This can be repeated several times to push thinking further.

2.3.4 Worked Example

Question:         How many r in raspberry?
Model thinking:   ...r, a, s, p, b, e, r, r, y ... so there are 2 r's.
Model tries:      [end-of-thinking]   ← blocked
Inserted:         Wait
Model continues:  , let me re-read... r - a - s - p - b - e - r - r - y
                  First r, second r, third r. Count = 3.
                  My initial answer of 2 was incorrect.
Answer:           3

The model was about to commit to a wrong answer. The forced "Wait" gave it a second pass, and it caught its own mistake.

3. Evaluation Setup

The model is tested on three standard reasoning benchmarks against strong open and closed baselines, and the authors introduce three metrics specifically for judging test-time scaling methods.

3.1 Benchmarks

AIME24 consists of the 30 problems from the 2024 American Invitational Mathematics Examination, covering algebra, counting, geometry, number theory, probability, and other secondary-school topics. Every answer is an integer from 000 to 999. Because the model can't read images, problems with figures are given in the Asymptote vector-graphics language. MATH500 is the set of 500 competition math problems selected by OpenAI in prior work. GPQA Diamond has 198 PhD-level questions in biology, chemistry, and physics; domain experts with PhDs score 69.7% on it.

Evaluations use lm-evaluation-harness with greedy decoding (temperature 0), and accuracy is equivalent to pass@1.

3.2 Baselines

The comparison set includes the closed OpenAI o1 series; DeepSeek's open-weight r1 series; Qwen's QwQ-32B-preview; Sky-T1-32B-Preview and Bespoke-32B, which are open models with open reasoning data distilled from QwQ and r1; and Gemini 2.0 Flash Thinking, the model s1K was distilled from. Gemini's API produced "recitation errors" during evaluation, so the authors entered the 30 AIME24 questions manually in its web interface and report no Gemini scores on MATH500 or GPQA.

3.3 Test-Time Scaling Metrics

Accuracy alone isn't enough to judge a scaling method; the authors also care whether you can control compute precisely and whether more compute actually helps.

3.3.1 Setup

For a given method, they run a set of evaluations a∈Aa \in \mathcal{A} on a fixed benchmark, each with a different amount of test-time compute. This yields a piecewise linear function ff, with thinking tokens on the x-axis and accuracy on the y-axis. For example, the rightmost AIME24 point for s1-32B is f(7320)=57%f(7320) = 57\%.

3.3.2 Control

Control measures how reliably a method obeys the thinking budget it is given, which is what makes it possible to deliberately trade speed for accuracy.

Control=1∣A∣∑a∈AI(amin⁡≤a≤amax⁡)\text{Control} = \frac{1}{|\mathcal{A}|}\sum_{a \in \mathcal{A}} \mathbb{I}(a_{\min} \le a \le a_{\max})

Here aa is the number of thinking tokens a run actually used, and amin⁡a_{\min} and amax⁡a_{\max} are the minimum and maximum allowed for that run; in practice usually only amax⁡a_{\max} is set. The indicator I(⋅)\mathbb{I}(\cdot) is a yes/no check that equals 1 if the run stayed within range and 0 otherwise. Summing these over all runs and dividing by the number of runs ∣A∣|\mathcal{A}| gives the fraction of runs that respected their budget. 100% means perfect control.

For example, in token-conditional control the model was told "think for up to N tokens" across 5 runs:

Budget told (amax⁡a_{\max})Tokens actually used (aa)I\mathbb{I}
1,0247,9390
2,0487,1580
4,0968,2630
8,1927,1081
16,3847,5001

Control=0+0+0+1+15=40%\text{Control} = \frac{0 + 0 + 0 + 1 + 1}{5} = 40\%

The model used roughly 7,000 to 8,000 tokens regardless of the instruction, so it only stayed within budget when the budget happened to be large. Adding budget forcing cuts thinking off exactly at the limit, so all 5 runs comply and Control becomes 100%.

3.3.3 Scaling

Scaling measures how much accuracy improves as a method is given more thinking, computed as the average slope of the accuracy curve ff from Section 3.3.1.

Scaling=1(∣A∣2)∑a,b∈Ab>af(b)−f(a)b−a\text{Scaling} = \frac{1}{\binom{|\mathcal{A}|}{2}} \sum_{\substack{a, b \in \mathcal{A} \\ b > a}} \frac{f(b) - f(a)}{b - a}

For any two runs aa and bb, where bb used more thinking tokens, the term f(b)−f(a)b−a\frac{f(b) - f(a)}{b - a} is the slope between them: accuracy gained divided by extra tokens spent. The sum adds this slope over every pair of runs, and dividing by (∣A∣2)\binom{|\mathcal{A}|}{2}, the number of possible pairs, turns the sum into an average. "nn choose 2" equals n(n−1)2\frac{n(n-1)}{2}, so 3 runs give 3 pairs and 5 runs give 10.

As an illustrative example (made-up numbers), take three runs at 2,000, 4,000, and 6,000 tokens scoring 40%, 50%, and 56%:

PairAccuracy gainExtra tokensSlope
2,000 → 4,000102,0000.005
4,000 → 6,00062,0000.003
2,000 → 6,000164,0000.004

Scaling=0.005+0.003+0.0043=0.004\text{Scaling} = \frac{0.005 + 0.003 + 0.004}{3} = 0.004

On average, each extra 1,000 thinking tokens adds about 4 accuracy points. A positive value means more thinking helps, zero means it changes nothing, and a negative value means more thinking actually hurts; a useful method must be positive, and larger is better. Averaging over all pairs, rather than only neighboring runs, captures the overall trend so that a single noisy run can't swing the result much.

The paper's values on AIME24 are:

MethodScaling
Class-conditional control25
Budget forcing15
Step-conditional control3
Token-conditional control−24
Rejection sampling−35

The paper doesn't spell out the exact units behind these numbers, so they are best read comparatively. Rejection sampling's strongly negative value reflects the inverse scaling described in Section 5.2.5. Class-conditional control has the steepest slope, but its control is only 50% and its best accuracy is low, which is why the paper judges methods on Control, Scaling, and Performance together rather than on any one metric.

3.3.4 Performance

Performance=max⁡a∈Af(a)\text{Performance} = \max_{a \in \mathcal{A}} f(a)

This is simply the best accuracy the method achieves on the benchmark. A method that scaled forever would eventually reach 100%, but in practice every method flattens out or runs into control or context-window limits.

4. Results

Fine-tuning on 1,000 examples plus budget forcing produces a model that is competitive with o1-preview on math, shows a clean scaling curve, and is the most sample-efficient open reasoning model in the comparison.

4.1 Headline Comparison

ModelTraining examplesAIME24MATH500GPQA Diamond
o1-previewN/A44.685.573.3
o1-miniN/A70.090.060.0
o1N/A74.494.877.3
Gemini 2.0 Flash ThinkingN/A60.0N/AN/A
Qwen2.5-32B-Instruct (base)N/A26.784.049.0
QwQ-32BN/A50.090.654.5
r1≫800K79.897.371.5
r1-distill800K72.694.362.1
Sky-T117K43.382.456.8
Bespoke-32B17K63.393.058.1
s1 without budget forcing1K50.092.656.6
s1-32B1K56.793.059.6

The "up to 27%" gain over o1-preview is a relative figure from AIME24: 56.7 is about 27% higher than 44.6. On MATH500 the lead is 7.5 points. On GPQA Diamond, which tests science rather than math, o1-preview remains well ahead (73.3 vs. 59.6). s1-32B is competitive with o1-preview on math, while the full o1 and o1-mini stay stronger.

4.2 Test-Time Scaling

s1-32B's accuracy rises steadily with more thinking tokens on all three benchmarks, which is the o1-style scaling behavior the paper set out to reproduce. On AIME24, the model stopping naturally scores 50%. Forcing it to keep thinking by suppressing the end-of-thinking delimiter 2, 4, and 6 times and appending "Wait" raises this to about 57%, beyond what the model reaches on its own. Gains flatten by the sixth extension; suppressing the stop too often can push the model into repetitive loops rather than new reasoning.

4.3 Sample Efficiency

ModelTraining examplesAIME24
s1-32B1K56.7
Sky-T117K43.3
Bespoke-32B17K63.3
r1-distill-32B800K72.6

Adding just 1,000 examples lifts the base model from 26.7 to 56.7 on AIME24. r1-distill scores higher, also using only SFT, but with roughly 800 times more reasoning samples. s1-32B sits on the sample-efficiency frontier, and whether r1-level performance is reachable with only 1,000 examples remains an open question.

4.4 Distillation Quality

s1K's traces come from Gemini 2.0 Flash Thinking, which scores 60.0 on AIME24. s1-32B reaches 56.7, nearly matching its teacher after seeing only 1,000 of its traces, which suggests the distillation procedure worked well.

4.5 Sequential vs. Parallel Scaling

The base Qwen2.5-32B-Instruct was given 64 samples per question at temperature 1, with majority voting over 2, 4, 8, 16, 32, and 64 of them. Even with this extra parallel compute, it couldn't catch up with s1-32B's single, budget-forced chain of thought. This supports the authors' intuition that thinking longer (sequential) is more effective than thinking many times independently (parallel).

5. Ablations

Every component matters: all three data criteria are needed together, and budget forcing is the only compute-control method that scores well on control, scaling, and performance at once.

5.1 Data Selection

Carefully choosing 1,000 examples beats both simpler selection rules and training on the entire pool.

5.1.1 Single-Criterion Datasets

Each alternative dataset has 1,000 examples picked using only one rule. All scores in this table use budget forcing with a cap of about 30,000 thinking tokens, which is why s1K's numbers differ slightly from Table 4.1.

DatasetSelection ruleAIME24MATH500GPQA
1K-randomQuality only36.790.652.0
1K-diverseDiversity only26.791.254.6
1K-longestDifficulty only (longest traces)33.390.459.6
s1KAll three combined50.093.057.6
59K-fullEntire pool, no filtering53.392.858.1

On AIME24, the single-criterion datasets score 13 to 23 points below s1K in absolute terms, which the paper summarizes as roughly 30% worse on average. The 95% bootstrap confidence intervals for 1K-random and 1K-diverse are entirely negative, so they are confidently worse. 1K-longest does well on GPQA but still falls short overall. No single rule is enough on its own.

5.1.2 Full 59K Pool

Training on all 59K examples gives a strong model, but its AIME24 improvement over s1K (53.3 vs. 50.0) has a confidence interval spanning zero, so it is not a reliable gain. It also costs 394 H100 GPU-hours versus 7 for s1K, roughly 56 times more. This echoes earlier findings in instruction tuning (LIMA) that careful data selection matters more than quantity.

5.2 Test-Time Control Methods

Budget forcing is compared against prompt-based ways of controlling thinking length and against rejection sampling.

5.2.1 Comparison Table

Results on AIME24, where ∣A∣|\mathcal{A}| is the number of evaluation runs used to estimate each method's properties:

MethodControlScalingPerformance∣A∣\lvert\mathcal{A}\rvert
Budget forcing100%1556.75
Token-conditional control40%−2440.05
Token-conditional + budget forcing100%1340.05
Step-conditional control60%336.75
Step-conditional + budget forcing100%636.75
Class-conditional control50%2536.72
Rejection sampling100%−3540.05

Budget forcing is the only method that combines perfect control, a clearly positive slope, and the best accuracy.

5.2.2 Token-Conditional Control

A variant of the model is trained with instructions such as "Think for up to 2048 tokens" added to each training prompt, with training traces bucketed by length into powers of two. At test time the model barely follows these instructions because it can't keep track of how many tokens it has written. Told to use 1,024 tokens, it used about 7,900; told to use 16,384, it used about 7,500. Only 2 of 5 runs stayed within the limit, giving 40% control and a negative scaling slope. Adding budget forcing restores perfect control, but a hard 1,024-token cap drops AIME24 accuracy to 3.3%.

5.2.3 Step-Conditional Control

Since token counting fails, the authors try a coarser unit. Reasoning traces are split into steps at double newlines, and the model is trained with countdown markers such as "64 steps left." Without intervention, the model still ignores the target; told to use 16 steps, it used about 123. With forcing, the model respects the step count but games it by writing longer steps: about 96 tokens per step under a 16-step limit versus 56 under a 256-step limit. Total thinking barely shrinks, so the budget is effectively dodged.

5.2.4 Class-Conditional Control

Here a generic instruction is appended to the question, telling the model to think either briefly or at length. The "long" prompt does increase thinking and accuracy relative to the "short" prompt, which gives this method the highest scaling slope. But control is imprecise, and both prompts actually score lower on AIME24 (30.0% short, 36.7% long) than using no prompt at all (50.0%).

5.2.5 Rejection Sampling

This method samples at temperature 1 until a generation fits within the chosen thinking budget. Surprisingly, allowing longer generations produces lower accuracy, an inverse scaling trend. The authors' explanation is that when the model is on the right track from the start, it finishes quickly; long generations tend to be ones where it made mistakes, backtracked, and second-guessed itself. Filtering for longer outputs therefore selects the troubled runs. In one example, a question was answered correctly under a 4,000-token budget, where the model went straight to the right approach, but incorrectly under an 8,000-token budget, where it backtracked repeatedly.

5.3 Extrapolation Cue

When extending thinking twice, the choice of appended word matters:

Cue appendedAIME24MATH500GPQA
No extension50.093.057.6
Nothing (only block the stop)50.090.255.1
"Alternatively"50.092.259.6
"Hmm"50.093.059.6
"Wait"53.393.059.6

Blocking the stop without adding any word actually hurts. "Wait" gives the best results overall and is the only cue that improves AIME24, likely because it most directly signals the model to reconsider what it just concluded.

5.4 Training Sequence Length

The maximum sequence length used during fine-tuning affects both accuracy and efficiency:

4,096 tokens32,768 tokens
Training samples truncated74%0%
AIME24 (accuracy / avg. thinking tokens)30.0% / 20,72150.0% / 6,984
MATH50090.0% / 5,32491.0% / 3,268
GPQA52.5% / 6,84153.0% / 3,568

With a short limit, most samples are cut off before the answer section, so the model rarely receives gradient updates on the transition from thinking to answering. It never properly learns to wrap up and rambles at test time. Training on complete samples raises the probability of producing the answer section at any point, which yields shorter reasoning and higher accuracy. The longer setting is better on both counts.

6. Discussion

The authors argue that reasoning ability is already latent after pretraining and that minimal fine-tuning activates it; budget forcing then scales it, up to limits that parallel methods can help push past.

6.1 Why 1,000 Examples Suffice

The base model has already seen enormous amounts of reasoning during pretraining on trillions of tokens, so the capacity to reason is already present. The 1,000-example fine-tuning stage simply activates it, and budget forcing scales it further at test time. This parallels LIMA's "Superficial Alignment Hypothesis," which found that about 1,000 examples can be enough to align a model to user preferences.

6.2 Limits of Budget Forcing

Budget forcing has two main limits when pushed further. Its gains eventually flatten, partly because the model can fall into repetitive loops, and it is bounded by the model's context window. Scaling down compute, by contrast, behaves predictably and doesn't face these constraints, which is part of why the scaling curves span a wide accuracy range. The authors propose several directions for extending it: rotating through different cue strings rather than only "Wait," combining it with frequency penalties or higher temperature to avoid loops, and applying budget forcing to models trained with reinforcement learning.

6.3 Parallel Scaling as a Complement

To scale beyond these limits, the authors combine s1-32B with two parallel methods. Majority voting generates several solutions and picks the most frequent answer. REBASE is a tree search guided by a process reward model (initialized from LLaMA-34B), with its outputs aggregated by majority vote. REBASE scales better than majority voting and, in this setting, even better than sequential scaling, though it adds an extra reward-model forward pass at every step. Sequential scaling hits a wall here: when prompted to use up to 512 steps, the model exceeded its context window on 12 of the 30 AIME24 questions, causing a large performance drop. Parallel methods therefore complement sequential scaling and offer a way to keep scaling beyond a fixed context window.

6.4 Follow-Up: s1.1

Seven days after releasing s1, the authors released s1.1. It uses the same 1,000 questions and the same training procedure, but with reasoning traces regenerated by DeepSeek r1 instead of Gemini. The improvement is substantial, especially on the newer AIME 2025:

ModelExamplesMATH500GPQAAIME24AIME25
s1 (no budget forcing)1K92.656.650.026.7
s1.1 (no budget forcing)1K94.460.656.750.0
s1.1 (budget forcing, "Wait" 2×)1K95.463.656.750.0
LIMO81794.866.756.344.6

The authors also tried distilling from Claude 3.7, which performed worse than distilling from r1.

6.5 Evaluation Note: Nondeterminism

Evaluations run on vLLM, and the authors found that scores can shift noticeably between runs even with fixed seeds and greedy decoding. Causes include different batch sizes, continuing a generation, and changes in tensor parallelism. Because the model writes very long reasoning traces, tiny numerical differences snowball: generations can match exactly for thousands of tokens, then differ by a single token and end with a completely different answer. To reduce this, final evaluations are run in full precision.

7. Key Takeaways

  1. A strong reasoning model can be built with only 1,000 fine-tuning examples, provided they are selected for quality, difficulty, and diversity together.
  2. Budget forcing, which caps thinking or extends it by appending "Wait," gives precise control over test-time compute and produces the first open reproduction of o1-style scaling curves.
  3. Sequential scaling (thinking longer) beats parallel scaling (voting over many attempts) at comparable settings, though parallel methods like REBASE help push past context-window limits.
  4. Simple prompting cannot control thinking length, because models can't count tokens and will game step limits; rejection sampling for length even reverses the scaling trend.
  5. Reasoning appears to be largely latent after pretraining, so minimal fine-tuning activates it and test-time compute amplifies it.