Reinforcement Learning from Verifiable Rewards

AI and Machine Learning · Language Models · 2025 · also: RLVR · verifiable-rewards.yaml

Replaces the learned reward with a programmatic check — does the proof verify, do the tests pass, is the answer right. Works only where correctness is decidable, and there the signal cannot be gamed.

Colour is the family; a dashed line is the second member of it.

Drag to pan · scroll to zoom · click a node to open it
G grpo Group Relative Policy Optimization rlhf RLHF verifiable-rewards Reinforcement Learning from Verifiable Rewards verifiable-rewards->grpo one scalar per completion, with no critic to fit verifiable-rewards->rlhf a learned reward can be gamed; a checked answer cannot

This node

References