Back to glossary

AI GLOSSARY

RLVR

Reinforcement Learning from Verifiable RewardsLearning Paradigms

A post-training method that scores a model's outputs using automatically checkable pass/fail signals, such as whether a math answer is correct or code passes its tests, instead of relying on human preference judgments as in Reinforcement Learning from Human Feedback.