RolesRatesPlatformsGuides
Get alerts

On this page

  1. Grade in an order that exposes errors
  2. Keep dimensions separate
  3. A worked review of two individual responses
  4. Choose severity from consequences and rules
  1. Home
  2. /Guides
  3. /How to Grade a Single AI Response
◆ Evaluation skillsHow-to

How to Grade a Single AI Response

Separate correctness, completeness and instruction following, then turn observed defects into a defensible score.

5 min read · Published September 9, 2026

You'll learn

  1. 01Grade each requested dimension before forming an overall impression.
  2. 02Severity depends on the rubric and the effect on the user’s task.
  3. 03A correct fact does not compensate automatically for a missed constraint.
  4. 04The explanation and the submitted score must agree.
In this guide+−
  1. Grade in an order that exposes errors
  2. Keep dimensions separate
  3. A worked review of two individual responses
  4. Choose severity from consequences and rules

Read the prompt as a set of obligations. Some concern the answer’s substance; others concern scope, format, sources or permitted actions. A polished response can meet one obligation and fail another.

Grade in an order that exposes errors

A single-response review

  1. Extract the requirements

    Record the question, requested format, source restrictions and any explicit exclusions. Keep these separate from preferences you would personally add.

  2. Read the response once for its claim

    Identify what the response actually says or promises. Do not start assigning scores halfway through a sentence.

  3. Check the decisive facts

    Use the supplied evidence and permitted tools. Recalculate or inspect the source where the answer depends on it.

  4. Score the requested dimensions

    Map each observed defect to the rubric. Avoid inventing dimensions or averaging scores unless instructed.

  5. Assign severity and overall judgment

    Ask what the defect prevents the user from doing, then apply the defined severity or aggregation rule.

  6. Reconcile the rationale

    Make sure the text identifies the defect that explains the score, without claiming checks you did not perform.

Keep dimensions separate

A compact review grid
DimensionQuestion to answer
CorrectnessAre the claims, calculations and inferences right under the available evidence?
Instruction followingDid the answer respect the explicit constraints?
CompletenessDoes it cover what the prompt requires at the requested level of detail?
RelevanceDoes each part help answer this request?
StyleDoes the wording fit the audience and any specified register?
Safety or policy complianceDoes it satisfy the particular policy this task supplies?

Use this grid as practice, not as a replacement for a project rubric. A real task might merge dimensions, omit style or require a specific policy label. If the rubric says a particular failure determines the overall result, follow that rule instead of taking an informal average.

A worked review of two individual responses

In the exercise below, grade each response independently first. The final A/B choice is only a compact way to show which response satisfies the practice brief; it does not replace the separate dimension judgments.

◆ Evaluation example

Prompt
Illustrative task. Give exactly three bullet points explaining why a team keeps meeting notes. No opening sentence.
Response A
- Record decisions. - Assign follow-up responsibilities. - Preserve context for absent colleagues.
Response B
Meeting notes record decisions, assign responsibilities and preserve context for absent colleagues.
Decision
Response A
Why
A meets the content and formatting requirements. B gives three relevant reasons but fails the explicit bullet-point format. Mark that instruction-following defect under the supplied scale. Do not invent a numeric score or call the whole answer factually wrong.

Choose severity from consequences and rules

A typo that does not change meaning is different from a wrong date that causes the user to miss an event. A formatting defect can also be consequential when another system expects an exact output format. The same visible error can deserve different severity in different tasks.

◆ Common mistake

Letting one strong dimension erase another

A response can be accurate and incomplete, or readable and unsupported. Record those differences before selecting an overall rating. Otherwise fluency becomes an accidental substitute for evidence.

Before you submit

  • →Check every explicit constraint against the response.
  • →Verify the claim on which the answer depends.
  • →Use only the rubric’s labels and scale.
  • →Distinguish missing evidence from a demonstrated falsehood.
  • →Ensure the rationale explains the chosen severity.
  • →Remove claims about tests or research you did not perform.

Key takeaway

A defensible score is a short chain: requirement, observation, consequence, rubric label.

Next in this section

How AI Evaluation Rubrics Work →

Understand dimensions, score anchors and calibration so your judgments follow the same rules as the rest of the project.

Role alerts

New roles in your fields, in one daily email, only when there are any.

Get alerts

Related guides

  • How to Compare Two AI Responses

    4 min read

  • How to Write Strong Evaluation Justifications

    5 min read

  • How AI Evaluation Rubrics Work

    4 min read

An independent platform that tracks AI expert-work roles, publishes what they pay, and compares the platforms that offer them. Free to use.

Explore
RolesRatesPlatformsGuides
Alerts
Get alertsScreening guide
Company
AboutDisclosurePrivacyContact
© 2026 SidequestNot affiliated with the platforms we list · Some links are referrals