Back to glossary

AI GLOSSARY

Eval Awareness

Safety, Alignment & Ethics

A model's tendency to recognize that it is being tested or evaluated, rather than used in a normal deployment, and to adjust its behavior in response. This can undermine the validity of safety and capability evaluations, since a model that behaves differently under test conditions may not reveal how it would actually act in deployment.