Using a model to grade outputs against a rubric, which scales evaluation but needs checking against human grades.
You have done this if
A second model scored answers for faithfulness and you spot-checked a sample each week.
Say it in a review
We use a judge model with a written rubric, calibrated against human labels every month.
On the AI Application map Observability