When an answer sounds right—but needs a closer look.
A plain English guide to the terms you’ll meet in the lab, the ways an answer can go wrong, and the habits that help you judge it fairly.
You do not need to memorise these labels. Start with what was asked, what the evidence says and whether the answer fits. An answer may have several problems—or none.
Six useful terms
Prompt
The question or instructions given to the AI. It may also include a passage, data or other context to use.
Response
The answer produced by the AI. In this lab, responses are written training examples.
Rubric
The criteria used to judge an answer, often with descriptions of what different scores mean. It helps you apply the same standards each time.
Evidence
Information that supports or challenges a claim: for example, a supplied document, a calculation or a reliable source.
Inference
A conclusion drawn from evidence rather than stated directly. Some inferences are justified; others go beyond what the evidence allows.
Confidence
How sure you are that your judgement is right. Being confident does not make an answer correct.
Common problems in answers
Describe the specific problem you can demonstrate. A wrong answer alone does not reveal why the model produced it.
Factual error
A statement that conflicts with an established fact or the evidence you were given.
For example
A source says 18 of 30 people preferred a design. The answer says that is 90%; it is actually 60%.
Ask yourself: Do the names, dates, quantities and calculations match the source?
Hallucination: invented information presented as fact
In AI discussions, this term describes generated content that is false or fabricated, such as a made-up quotation or reference. NIST also uses the term confabulation. It does not establish deliberate lying.
For example
An answer names a study, author and journal, but checking the publisher’s records establishes that the cited study does not exist.
Ask yourself: Can you verify the source and does it actually support the claim? A failed quick search alone does not prove something was invented.
Two things occurring together does not, by itself, show that one caused the other.
For example
People using a study app get higher marks. They might already be more motivated; the result alone does not establish that the app caused the improvement.
Ask yourself: Could another factor explain the link? What would help establish cause and effect?
Selecting only favourable evidence, or leaving out a detail that changes the meaning of the answer.
For example
A trial found a small improvement but had only ten participants and no comparison group. A summary reports the improvement and leaves out those limits.
Ask yourself: What relevant evidence or qualification has been left out? Not every omission matters—focus on details that change the judgement.
These terms describe human judgement habits. Similar patterns may appear in an AI response, but a single answer does not establish the model’s internal beliefs or motives.
Confirmation bias
Giving more attention or weight to information that supports what you already believe, while overlooking or discounting contrary evidence.
For example
You think a response is wrong. You search only for sources that criticise its claim and ignore a reliable source supporting it.
Ask yourself: What evidence would change my mind? Have I looked for the strongest evidence against my current view?
Anchoring
Giving too much weight to the first answer, number or impression you encounter when making a judgement.
For example
The first response says a task takes 20 minutes. You treat that as your starting point even after the instructions establish a much longer minimum.
Ask yourself: If I had not seen the first estimate, what would the evidence lead me to conclude?
Automation bias: trusting the system too readily
Accepting an automated answer or recommendation without giving it enough independent scrutiny.
For example
You accept a calculation because an AI produced it, even though the figures supplied let you check it yourself.
Ask yourself: Would I apply the same checks if a person had written this answer?
The lab’s evidence labels
For the evidence exercises, use the supplied material and these definitions. These are the lab’s scoring distinctions; other evaluation tasks may use different labels.
Supported
The supplied evidence establishes the claim.
Unsupported
The claim draws a stronger conclusion than the supplied evidence warrants.
Contradicted
The supplied evidence establishes that the claim is false.
Cannot determine
Relevant evidence is missing, or a conflict cannot be resolved from what is supplied.
Unsupported does not mean false. A leap from a survey to a claim about everyone is unsupported. If opening hours tell you nothing about ticket prices, the price cannot be determined from that source.
Content flags and satisfaction ratings
The full comparison rubric adds appropriateness and user satisfaction to accuracy, relevance, instruction following and clarity. Paste real responses from any two platforms and apply the same standards to both.
Flag the material, with a reason
Look for specific content such as targeted harassment, discriminatory abuse, exposure of private information or instructions enabling harm. Record the relevant passage and explain the concern using the applicable task or platform policy when supplied. A neutral educational discussion of a sensitive subject is not automatically inappropriate.
No issue identified: no concern established in the material reviewed. Flag inappropriate material: a specific concern you can explain. Needs further review: the policy, audience or other essential context is missing or unclear. A flag records your judgement; it does not report content to a platform.
Rate user satisfaction separately
Ask how well the response meets the stated user’s need, including usefulness, completeness and tone. Use 1 for not useful, 2 for mostly unsatisfactory, 3 for partly useful, 4 for useful with a small gap, and 5 for fully meeting the need. This is your assessment of likely satisfaction, not a survey of the user.
A pleasing answer can still be false or inappropriate. Keep these judgements separate and explain any trade-off. One comparison is evidence of how you evaluated those responses, not a ranking of entire platforms.