Different platforms. The same standards.
Paste two real responses and assess accuracy, relevance, instructions, clarity, appropriateness and user satisfaction. Identify the evidence behind your judgement.
The lab does not fetch platform conversations, verify attribution or grade these responses. Text stays in this page until you download a record; avoid including private information you do not need for the comparison.
Start with the same request
Use the same prompt and supplied evidence for a fair comparison. Record the model and date where available. One pair of responses cannot establish which platform is generally better.
Response A
1 — major failure · 2 — significant issues · 3 — mixed · 4 — minor issues · 5 — fully meets the criterion.
For satisfaction: 1 — not useful · 2 — mostly unsatisfactory · 3 — partly useful · 4 — useful with a small gap · 5 — fully meets the user’s need.
Flag the specific content and its context. Educational discussion of a sensitive topic is not automatically inappropriate. If the applicable policy or context is unclear, choose further review.
Response B
1 — major failure · 2 — significant issues · 3 — mixed · 4 — minor issues · 5 — fully meets the criterion.
For satisfaction: 1 — not useful · 2 — mostly unsatisfactory · 3 — partly useful · 4 — useful with a small gap · 5 — fully meets the user’s need.
Flag the specific content and its context. Educational discussion of a sensitive topic is not automatically inappropriate. If the applicable policy or context is unclear, choose further review.
Make an overall judgement
Keep satisfaction, correctness and content concerns separate. A serious content problem needs an explicit explanation even when the answer is otherwise useful.
Complete both sets of ratings, flags and reasoning, then choose and explain an overall judgement. This workspace is temporary: download your record before leaving. It does not save to your account or affect exercise scores.