FrontierAI.Engineer
Evaluation, Guardrails & Safety

Inter-Annotator Agreement

Also known as: IAA, inter-rater reliability

Inter-annotator agreement quantifies how consistently different human annotators apply the same labels or scores to the same model outputs. High agreement — measured with Cohen's kappa, Krippendorff's alpha, or percent agreement — validates that the evaluation task and rubric are clear and that the resulting labels are reliable. Low agreement signals ambiguous instructions or inherently subjective evaluation criteria, and any metrics built on such labels carry significant noise.