The decision
A rubric tells reviewers what counts as success. Avoid a single score for “quality” because it mixes severe factual errors with minor style preferences. Use a few dimensions tied to the workflow and identify any failure that blocks release regardless of the average.
A worked example
A client-update rubric requires every factual claim to be supported, all required decisions to be mentioned, and conditional statements to remain conditional. Tone can be rated separately. One invented deadline is a blocking failure even if the draft is concise and pleasant.
How to put it into practice
- Name each criterion and describe a pass with an example.
- Mark critical failures that cannot be averaged away.
- Have two people independently score a small subset and discuss disagreements.
- Keep automated checks for objective rules and human review for judgments they cannot reliably cover.
A failure to plan for
A model judging another model may share the same blind spots. Use model-based review as an aid, validate it against human decisions, and retain deterministic checks where possible.
Try it on your project
Score five outputs with separate columns for factual support, completeness, preserved uncertainty, and editing effort. Compare your ratings with a second reviewer and rewrite ambiguous criteria.
Keep the next step small
Use the free demand scorecard or planning tools to make your assumptions explicit. The $19 launch kit brings the blueprint and seven editable worksheets together.