FrontierAI.Engineer
Evaluation, Guardrails & Safety

LLM-as-Judge

Also known as: LLM judge, model-based evaluation

LLM-as-judge uses a capable language model — typically a frontier model — to score or rank the outputs of another model being evaluated. The judge reads the model's response alongside a rubric and optionally the source context, then produces a score, verdict, or comparison decision. It scales cheaply to large eval sets and can capture nuanced quality dimensions that rule-based metrics miss, but it inherits the judge model's own biases and blind spots.