The AI Evaluation Tasks You’ll Actually Be Given
A practical map of scoring, ranking, fact-checking, writing and reviewing tasks, including what each submission needs to contain.
You'll learn
- Identify the required deliverable before you begin judging the output.
- Scoring, ranking and rewriting solve different problems.
- Agent and multimodal tasks add evidence beyond the final text.
- Calibration examples help align reviewers on how rules apply.
In this guide
The first question is not “Is this a good answer?” It is “What am I being asked to submit?” A project may want a categorical label, separate dimension scores, a preferred response, a correction or a new task. Giving a sensible answer in the wrong form is still a failed submission.
- Quick definitionRubric
- The criteria and decision rules used to score a task. It can include definitions, severity levels, examples and instructions for resolving trade-offs.
The task map
| Task | Example assignment | Required output |
|---|---|---|
| Single-response grading | Check a summary against a source document. | Scores or labels for the requested dimensions. |
| Pairwise comparison | Choose between two answers to the same prompt. | A preference and the deciding difference. |
| Classification | Label whether a response contains an unsupported claim. | One or more permitted categories. |
| Fact-checking | Verify a stated date and the citation offered for it. | Claim-level evidence and a verdict. |
| Error annotation | Mark where a translation changes a negation. | A span, error category and severity. |
| Rewriting | Repair an answer while keeping the correct material. | A revised answer within the task constraints. |
| Prompt creation | Design a realistic question with a checkable answer. | The prompt and any required reference material. |
| Reference writing | Produce a checked target answer. | The answer, with support where required. |
| Safety review | Apply a supplied policy to a response. | A policy-grounded label and explanation. |
| Agent review | Inspect a tool-use trace and final artifact. | Separate outcome and process judgments. |
| Multimodal review | Check a caption against an image or clip. | Grounded labels, spans or timestamps. |
| Quality review | Check another evaluator’s submission. | Acceptance, correction or escalation under the review rules. |
One input, several legitimate assignments
Evaluation example
- No. The museum is closed on Mondays.
- Yes. Most museums welcome visitors on weekdays.
- Quick definitionCalibration
- Reviewers comparing judgments on shared examples to align their interpretation of a rubric. The purpose is consistent application, including discussion of ambiguous cases.
Reference tasks and review
A project may provide worked examples or tasks with an expected judgment. Use their explanations to understand the rule, rather than memorizing labels. If a new task differs in a material way, copying the old label can produce the wrong decision.
Quality review adds another layer: someone checks whether the original evaluator followed the task, identified the right evidence and justified the submitted label. A reviewer should be able to explain a correction with the same discipline expected from the contributor.
Reality check
One good evaluation method will work unchanged on every project.
Projects can define dimensions, severity and tie rules differently. Transfer the habit of checking evidence; reread the operational rules each time.
Tools change what you can establish
An open-web fact-check permits different conclusions from a task restricted to a supplied document. An agent task may include logs that a text-only task lacks. Before making claims about what you verified, establish which evidence you are allowed to inspect and which tools you actually used.
Key takeaway
Name the deliverable, locate the governing criteria and inspect the right evidence. Then make the judgment.