Prompt
Eval Rubric Prompt
Builds a scoring rubric a grader model can apply consistently.
1 min readupdated 2026-08-01
/ quick answer
Use to define model-graded scoring for a subjective AI output. Builds a scoring rubric a grader model can apply consistently.
Builds a scoring rubric a grader model can apply consistently. Use to define model-graded scoring for a subjective AI output. Copy the prompt below, swap the bracketed variables for your own context, and run it in any capable model. This prompt node is part of the Onexial knowledge graph and links to related concepts, workflows and tools below.
Context
Use to define model-graded scoring for a subjective AI output.
Prompt
Create an evaluation rubric for the AI output described below. Return:
1. DIMENSIONS — 3–5 scoring dimensions, each with a one-line definition.
2. SCALE — a 1–5 scale per dimension with concrete descriptions of what each score looks like.
3. AUTOMATIC FAILURES — conditions that force a score of 1 regardless of other qualities.
4. GRADER PROMPT — a ready-to-use prompt that scores one output against this rubric and returns JSON.
5. THRESHOLD — the pass score, with reasoning.
Be concrete enough that two different graders would agree.
Output type: {{OUTPUT_TYPE}}
User goal: {{GOAL}}
Non-negotiables: {{REQUIREMENTS}}Example Output
1. DIMENSIONS — Accuracy: every claim is supported by the source. Completeness: answers all parts of the question...
Related Workflow
Related Tool Stacks
/ frequently asked
What does the Eval Rubric Prompt prompt do?
Use to define model-graded scoring for a subjective AI output.
Which AI models work with this prompt?
It is model-agnostic: it works with any capable general model. Replace the bracketed variables with your own context before running it.
What output should I expect?
1. DIMENSIONS — Accuracy: every claim is supported by the source. Completeness: answers all parts of the question...
↳ connected nodes
Workflow↳ linked
Build an Eval Suite Before Optimising Prompts
Stop guessing whether a change improved anything.
Tool Stack↳ linked
AI Observability Stack
Traces, cost, evals and quality drift for AI systems in production.
Dictionary↳ linked
AI Evaluation
AI evaluation is the measurement layer of an AI system: a fixed set of cases, a scoring method and a tracked pass rate you can regress against.
Dictionary↳ linked
AI Monitoring
AI monitoring is production observability for model-driven systems: traces, cost, latency, tool failures and output-quality drift.
Workflow↳ linked
Monitor an AI System in Production
See quality, cost and failure drift before your users report it.
Prompt↳ linked
AI System Threat Model Prompt
Produces a concrete threat model for an AI system with tool access.