← Evaluation, Guardrails & Safety
Inter-Annotator Agreement
Also known as: IAA, inter-rater reliability
Inter-annotator agreement quantifies how consistently different human annotators apply the same labels or scores to the same model outputs. High agreement — measured with Cohen's kappa, Krippendorff's alpha, or percent agreement — validates that the evaluation task and rubric are clear and that the resulting labels are reliable. Low agreement signals ambiguous instructions or inherently subjective evaluation criteria, and any metrics built on such labels carry significant noise.