Replaces the learned reward with a programmatic check — does the proof verify, do the tests pass, is the answer right. Works only where correctness is decidable, and there the signal cannot be gamed.
supersedescorrects · extends
classifiesspecializes · part-of
substitutes forapproximates · alternative-to
depends onrequires · validates
Colour is the family; a dashed line is the second member of it.
Drag to pan · scroll to zoom · click a node to open it
This node
correctsfixes a defect in RLHFa learned reward can be gamed; a checked answer cannot