Back to glossary

AI GLOSSARY

Constitutional Classifiers

Safety, Alignment & Ethics

A safeguard developed by Anthropic that uses separate classifier models, trained on a written set of principles, to detect and block harmful inputs or outputs in specific risk categories. It extends the idea behind Constitutional AI into a dedicated filtering layer around a deployed model.