Back to glossary
AI GLOSSARY
RLVR
Reinforcement Learning from Verifiable RewardsLearning Paradigms
A post-training method that scores a model's outputs using automatically checkable pass/fail signals, such as whether a math answer is correct or code passes its tests, instead of relying on human preference judgments as in Reinforcement Learning from Human Feedback.