456
Dictionary

LLM-as-Judge

Using a strong model to grade another model's output.

1 min readupdated 2026-07-04

/ quick answer

LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for. Using a strong model to grade another model's output.

Using a strong model to grade another model's output. LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for. In practice: GPT-4 grades 500 summaries against a rubric: relevance 1-5, faithfulness 1-5. This dictionary node is part of the Onexial knowledge graph and links to related concepts, workflows and tools below.
Definition
LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for.
Example
GPT-4 grades 500 summaries against a rubric: relevance 1-5, faithfulness 1-5.
Related Workflows
/ frequently asked

What is LLM-as-Judge?

LLM-as-judge automates eval scoring: give the judge the input, the output, and a rubric, and get a score. Cheap alternative to human labeling — with known biases (length, position) to control for.

What is an example of LLM-as-Judge?

GPT-4 grades 500 summaries against a rubric: relevance 1-5, faithfulness 1-5.

Why does LLM-as-Judge matter for AI and automation?

Using a strong model to grade another model's output. It connects to the workflows, prompts and tool stacks linked on this page, so you can move from definition to execution without leaving Onexial.

/ topics#ai#quality