FrontierAI.Engineer
Evaluation, Guardrails & Safety

Evaluation Rubric

Also known as: scoring rubric

An evaluation rubric defines the criteria and scoring scale used to judge model outputs, whether by a human annotator or an LLM judge. A rubric might specify dimensions like accuracy, completeness, tone, and groundedness, each rated on a scale such as 1–5 with anchoring examples. Explicit rubrics reduce subjectivity, improve inter-annotator agreement, and make LLM-as-judge scoring more consistent and reproducible across evaluation runs.