Back to glossary

AI GLOSSARY

Sandbagging

Safety, Alignment & Ethics

A model deliberately underperforming on a capability evaluation, for example to avoid triggering additional safety restrictions or scrutiny that stronger measured capabilities would require. Sandbagging is treated as a serious risk for evaluation-based safety frameworks, since it can make a model appear less capable, and therefore less risky, than it actually is.