Back to glossary
AI GLOSSARY
Sandbagging
Safety, Alignment & Ethics
A model deliberately underperforming on a capability evaluation, for example to avoid triggering additional safety restrictions or scrutiny that stronger measured capabilities would require. Sandbagging is treated as a serious risk for evaluation-based safety frameworks, since it can make a model appear less capable, and therefore less risky, than it actually is.