Back to glossary

AI GLOSSARY

LLM-as-a-Judge

Evaluation & Performance

Using one LLM, typically prompted with grading criteria, to score, rank or evaluate the outputs of another model, as a scalable substitute for human evaluation. Widely used for benchmarking, but sensitive to biases such as Preference Leakage.