Response evaluation asks human reviewers to compare or score AI outputs for qualities such as correctness, helpfulness, relevance, clarity, and safety.
Guidance can change when a project publishes more specific instructions. Always follow the active project’s latest rubric where it differs from general platform guidance.