Verifiable rewards: what a reasoning model learns from being checked
Design checkers that reward correct work and expose missing coverage.

A programming exercise offers a tempting training setup. A model proposes a function, a test suite runs, and the result supplies feedback. Unlike a review of writing style, the feedback may be mechanically checkable. The program either produces the expected output for a test or it does not. This makes coding and other verifiable tasks attractive places to study reinforcement learning for language models.
Yet a checkable result is not the same thing as a complete definition of success. Tests can omit important cases. A mathematical answer can be parsed incorrectly. A correct final value can hide an invalid argument. To understand training with verifiable rewards, follow the entire chain from the intended task to the checker, through the training update, and back to an independent evaluation.
Separate the model, the attempt, and the verifier
The model produces an attempt from a prompt. A verifier examines some property of that attempt and returns a reward. A training procedure then changes the model so that rewarded behavior becomes more likely under its objective. These are separate components, and improvements in one do not automatically repair weaknesses in another.
Imagine an invented exercise that asks for a function to return the largest valid temperature in a list. The checker tests a normal list, a list with negative values, and a list containing missing entries. Passing these examples provides evidence about those behaviors. It says nothing yet about malformed strings, numerical overflow, or the required response when every entry is missing.
What the research establishes
DeepSeekMath introduced Group Relative Policy Optimization in its mathematical reasoning work. DeepSeek-R1 studied reinforcement learning for reasoning and described training approaches that use checkable tasks. These papers are useful foundations for understanding why feedback from outcomes has become an important research direction. Their reported results belong to their specific models, datasets, and evaluation procedures.
Read the original research paper on arXiv
DeepSeek-R1: Reasoning through reinforcement learning
The practical discussion below is an original engineering framework, not a reproduction of either training recipe. In particular, a team cannot infer that adding a pass-or-fail reward to an arbitrary assistant will reproduce a research result. Starting model capability, task selection, sampling, and optimization all matter.
Outcome rewards answer a narrower question
An outcome reward asks whether a result satisfies the verifier. A process reward evaluates intermediate behavior or steps. Neither approach provides a magical view into the model's actual internal computation. A fluent written explanation may be useful to a reader without constituting a complete causal account of how the answer was produced.
For the temperature exercise, an outcome checker might compare the returned value with a reference implementation. A process-oriented assessment might also examine whether missing values were handled deliberately. The second assessment can provide additional information, but it introduces its own labeling and judgment problems. Treat the choice as a measurement design decision rather than a contest between two universal solutions.
The checker becomes part of the specification
Write down exactly what the verifier accepts. Does it compare strings, parse a number, execute a program, or inspect a structured object? Does it accept equivalent mathematical expressions? Does it distinguish a correct value with the wrong unit? A training system can only optimize the feedback it receives, including accidental preferences introduced by formatting rules.
Keep parsing failures separate from incorrect answers. Suppose a model returns a correct result inside an unsupported wrapper. That is an interface error, and it may deserve a penalty, but it should not disappear into an undifferentiated reasoning score. Separate categories make it possible to improve formatting without claiming that mathematical ability changed.
Coverage matters more than a green test indicator
Build tests around behavioral categories before collecting many near-duplicate examples. For a list-processing problem, those categories might include empty input, repeated values, valid negative numbers, and invalid entries. Then create multiple cases within each category. A large test file that repeats one easy pattern can offer less protection than a smaller set with carefully chosen boundaries.
Use hidden evaluation cases that are not reused as training feedback. If developers repeatedly inspect the hidden set and adjust the training process to it, it gradually becomes another development set. Preserve a final evaluation boundary and document when a previously held-out collection has been exposed. The boundary is a workflow property, not merely a folder name.
Relative improvement needs an absolute reference
Some training methods compare several attempts on the same problem. That comparison can distinguish better and worse responses within a group, but relative success does not guarantee that any response is good enough for use. If every candidate violates a requirement, the best candidate is still a failure under the application contract.
In a small illustrative review, imagine four attempts: one crashes, two mishandle missing values, and one passes all visible cases. The last attempt is preferable under that checker. Before calling it reliable, evaluate fresh cases and inspect whether the specification itself omitted a required behavior. Relative ranking and absolute acceptance answer different questions and should remain visible in reporting.
Reward the task without rewarding unnecessary work
A correct answer may require more computation on difficult inputs, but longer output is not automatically stronger reasoning. If the application values concise, correct results, measure accepted answers alongside output length and elapsed time. Avoid rewarding verbosity merely because it resembles examples that previously scored well.
Conversely, an aggressive length penalty can discourage useful checking. Treat efficiency as a constrained design problem: first define acceptable correctness and behavior, then examine the resources used to achieve them. Compare several budgets on representative tasks rather than optimizing a single combined number whose tradeoffs nobody can explain.
Keep execution separate from trust
When verification executes generated code, run it in an isolated environment with explicit resource limits and a narrow interface. The checker should not grant the model access to unrelated files or production services. This is ordinary containment for untrusted programs, regardless of whether the model appears cooperative in a conversation.
Also protect the evaluation records. A program that can alter expected outputs or inspect hidden answers can produce misleading success signals. Record the verifier revision with every result so that a later checker correction can be distinguished from a model improvement. Reproducible evidence requires the model and the measurement tool to be identifiable together.
Look for transfer beyond familiar exercises
A model can improve on one family of verifiable problems while remaining unchanged on another. Test variations that preserve the underlying skill but change the surface form, data representation, or combination of requirements. For the temperature example, switch from a plain list to records with units and timestamps, while making the new requirements explicit.
Do not confuse this with making the task arbitrarily harder. The purpose is to discover whether the learned behavior transfers to the intended application. Separate unfamiliar wording from genuinely new skills. When performance drops, that distinction helps determine whether the next step is better coverage, a clearer interface, or a different model.
Decide how partial credit changes behavior
A binary reward is easy to interpret, but some tasks contain separable requirements. Partial credit can help distinguish an almost-correct attempt from an unrelated answer. It can also encourage the model to collect easy points while neglecting a difficult requirement that is essential to the user. Define which conditions are mandatory before adding a graded score.
For the temperature exercise, correct handling of ordinary lists could earn diagnostic credit while invalid-input handling remains a release requirement. Report the diagnostic score, but do not let it override that requirement. This keeps training feedback informative without quietly redefining the application contract around whichever behaviors are easiest to reward.
A useful experiment produces an error map
Before training, save a baseline evaluation with failures grouped by cause. After training, run the same categories plus fresh held-out cases. Look for regressions as well as improvements. A system that solves more algebra questions but becomes less reliable at following output constraints may require a different deployment decision than its aggregate score suggests.
Include examples that a human can inspect. A table of totals can hide a checker bug or a recurring misunderstanding of the task. Short case records containing the prompt, accepted contract, attempt, and verifier result make disagreements discussable. Remove confidential inputs before sharing those records outside the evaluation team.
Verifiable rewards are powerful because they connect training to consequences that can be checked. Their value depends on the honesty of that connection. A strong system does not merely accumulate passing attempts; it makes clear what was checked, what remained untested, and whether the improvement survives a new problem that the training loop never saw.