How to Grade a Single AI Response
Separate correctness, completeness and instruction following, then turn observed defects into a defensible score.
You'll learn
- Grade each requested dimension before forming an overall impression.
- Severity depends on the rubric and the effect on the user’s task.
- A correct fact does not compensate automatically for a missed constraint.
- The explanation and the submitted score must agree.
In this guide
Read the prompt as a set of obligations. Some concern the answer’s substance; others concern scope, format, sources or permitted actions. A polished response can meet one obligation and fail another.
Grade in an order that exposes errors
A single-response review
Extract the requirements
Record the question, requested format, source restrictions and any explicit exclusions. Keep these separate from preferences you would personally add.
Read the response once for its claim
Identify what the response actually says or promises. Do not start assigning scores halfway through a sentence.
Check the decisive facts
Use the supplied evidence and permitted tools. Recalculate or inspect the source where the answer depends on it.
Score the requested dimensions
Map each observed defect to the rubric. Avoid inventing dimensions or averaging scores unless instructed.
Assign severity and overall judgment
Ask what the defect prevents the user from doing, then apply the defined severity or aggregation rule.
Reconcile the rationale
Make sure the text identifies the defect that explains the score, without claiming checks you did not perform.
Keep dimensions separate
| Dimension | Question to answer |
|---|---|
| Correctness | Are the claims, calculations and inferences right under the available evidence? |
| Instruction following | Did the answer respect the explicit constraints? |
| Completeness | Does it cover what the prompt requires at the requested level of detail? |
| Relevance | Does each part help answer this request? |
| Style | Does the wording fit the audience and any specified register? |
| Safety or policy compliance | Does it satisfy the particular policy this task supplies? |
Use this grid as practice, not as a replacement for a project rubric. A real task might merge dimensions, omit style or require a specific policy label. If the rubric says a particular failure determines the overall result, follow that rule instead of taking an informal average.
A worked review of two individual responses
In the exercise below, grade each response independently first. The final A/B choice is only a compact way to show which response satisfies the practice brief; it does not replace the separate dimension judgments.
Evaluation example
- - Record decisions. - Assign follow-up responsibilities. - Preserve context for absent colleagues.
- Meeting notes record decisions, assign responsibilities and preserve context for absent colleagues.
Choose severity from consequences and rules
A typo that does not change meaning is different from a wrong date that causes the user to miss an event. A formatting defect can also be consequential when another system expects an exact output format. The same visible error can deserve different severity in different tasks.
Common mistake
Letting one strong dimension erase another
A response can be accurate and incomplete, or readable and unsupported. Record those differences before selecting an overall rating. Otherwise fluency becomes an accidental substitute for evidence.
Before you submit
- Check every explicit constraint against the response.
- Verify the claim on which the answer depends.
- Use only the rubric’s labels and scale.
- Distinguish missing evidence from a demonstrated falsehood.
- Ensure the rationale explains the chosen severity.
- Remove claims about tests or research you did not perform.
Key takeaway
A defensible score is a short chain: requirement, observation, consequence, rubric label.