Back to glossary
AI GLOSSARY
LLM-as-a-Judge
Evaluation & Performance
Using one LLM, typically prompted with grading criteria, to score, rank or evaluate the outputs of another model, as a scalable substitute for human evaluation. Widely used for benchmarking, but sensitive to biases such as Preference Leakage.