Two AI systems answer the same question. One is concise, confident and neatly formatted. The other is slower, qualified and less elegant. Choosing the first because it feels better is not evaluation. Choosing the second because caution sounds intelligent is not evaluation either.
A useful comparison begins before either answer appears. You define the job, the evidence standard and the cost of failure. Then you inspect the outputs at the level of claims, calculations, instructions and omissions. The goal is not to crown a universally superior model. It is to decide which answer better serves this task under these conditions.

Do not run a personality contest
People reward fluency, directness, agreement and familiar structure. Those signals can correlate with a useful answer, but they can also hide unsupported claims. A hesitant answer can be careful or merely confused. A long answer can be complete or evasive. A citation can support the sentence, point somewhere vaguely related or not exist at all.
Confidence Is Not the Same as Calibration explains why delivery and measured reliability must be separated. When comparing AI outputs, remove adjectives such as “smart,” “thoughtful” and “professional” until you can name the observable property they refer to.
Replace “Answer A feels stronger” with statements such as “Answer A identifies all four constraints,” “Answer B's cited source does not support its number,” or “both answers omit the required exception.” Those statements can be checked.
Step 1: freeze the task and the evidence packet
Give both systems the same task, relevant source material, deadline, output format and allowed tools. Record product or model names exactly as the interfaces display them, the date, enabled modes and whether web retrieval or uploaded files were available. Do not claim two systems were tested equally if one had access to evidence the other never received.
Write the task as a deliverable with constraints. “Explain this regulation” is weak. “Using only the attached regulation, summarize the filing deadline, covered entities, exceptions and unresolved ambiguity for a non-lawyer; quote no more than one sentence and cite sections” is evaluable.
If the task is generative—names, concepts or prose—define audience, tone, prohibited claims and success conditions. Creativity is not exempt from constraints. It simply has more than one acceptable answer.
Step 2: define the rubric before reading
Use five to seven criteria and weight them by consequence. A practical default is:
- Task completion: Did the output answer every requested part in the required form?
- Factual support: Are material claims accurate and supported by the supplied or independently verified evidence?
- Reasoning fit: Do calculations, classifications and conclusions follow from the inputs?
- Boundary honesty: Does the answer distinguish fact, inference, uncertainty and missing information?
- Omission risk: Did it miss an exception, dependency, stakeholder or failure mode that changes the decision?
- Usability: Can the intended reader act without decoding needless complexity?
- Safety and reversibility: Does it avoid unsupported high-stakes instructions and make correction possible?
Not every criterion deserves equal weight. For a headline, readability may dominate. For a wiring procedure, medical summary or legal deadline, factual support and omission risk dominate. A beautiful unsafe answer should lose quickly.
Step 3: blind the style cues when practical
Copy the answers into the same plain format. Remove model names, interface colors and ornamental headings. Keep meaningful structure, warnings, citations and uncertainty language. The point is not to strip away usability; it is to prevent brand familiarity and interface polish from deciding the result.
Read once for task completion without scoring elegance. Mark where each answer addresses each requirement. A requirement matrix makes missing work visible before prose quality pulls attention elsewhere.
Step 4: convert prose into checkable units
Underline every claim that could change the decision: dates, quantities, compatibility statements, quotations, causal claims, legal requirements, safety instructions and promises about a product or service. Break compound sentences apart. “The device supports the protocol and will work with your setup” contains at least two claims; the second requires information about the setup.
Classify each unit:
- Directly supported: the provided evidence says it.
- Reasonable inference: the evidence plus a stated assumption supports it.
- Needs verification: it depends on current or outside information.
- Unsupported: no available source or reasoning path carries it.
- Not applicable: it does not help the task.
A high count of true but irrelevant facts does not compensate for one unsupported claim that controls the decision.
Step 5: verify the high-impact claims independently
Open the primary source. Confirm the document identity, version, date and relevant section. Check whether quoted language is exact and whether surrounding text changes its meaning. The protocol in How to Verify an AI Citation Before You Use It applies even when both systems cite the same source; copied errors can agree.
For calculations, recompute from the original inputs using a separate method. For code, run representative tests and at least one adversarial or boundary case. For summaries, compare against the source's scope and exclusions. For product or policy facts, use current official documentation rather than a remembered feature list.
Agreement between two AI answers is not independent corroboration. The systems may share training sources, retrieval results or common shortcuts. Agreement tells you where to check first; it does not replace checking.
Step 6: test with cases the answer did not choose
An explanation can appear complete because its example was selected to fit. Supply a second example, edge case or counterexample. Ask whether the rule still works when a condition reverses, a value is missing or a user lacks the assumed tool.
For code, vary input size, empty values, permissions and failure responses. For a decision framework, test a case near the cutoff. For a factual synthesis, choose one claim from the least prominent source. For instructions, simulate the step where a user is most likely to misunderstand.
This is local evaluation, not a universal benchmark. An AI Benchmark Is Not the Real World explains why scores require a defined population of tasks and conditions. Your small comparison supports a narrow conclusion: which answer worked better on the cases you actually tested.
Step 7: inspect uncertainty and refusal quality
Reward uncertainty only when it is informative. “It depends” should be followed by the variables it depends on and the next evidence that would resolve them. Penalize false precision, but also penalize vague caveats pasted onto an otherwise unsupported conclusion.
A good refusal identifies the boundary, preserves useful safe assistance and names what would make the task answerable. A bad refusal blocks low-risk work without explanation. A reckless answer proceeds through missing prerequisites. Evaluate both the decision to answer and the quality of what follows.
Step 8: score with failure notes, not one magic number
Use a 0–2 scale for each criterion: missing, partial or adequate. Multiply only when weighting is genuinely useful. Beside the score, record the strongest reason the answer could fail. A total without failure notes can hide that one response lost points for verbosity while the other invented a safety limit.
Set disqualifiers in advance. A fabricated source, incorrect calculation controlling the conclusion, unsafe instruction or failure to follow a hard constraint can disqualify an answer regardless of style points.
Step 9: choose, combine or reject
The result need not be A or B. You can take A's structure, B's correctly supported detail and your own verified conclusion. You can also reject both. If combining, recheck the assembled output; two individually acceptable passages can conflict when joined.
For consequential work, preserve the task, versions, evidence packet, rubric, failure notes, human reviewer and final edits. How to Write an AI Decision Record provides an eleven-field record and rollback path.
A compact comparison worksheet
Task and audience:
Hard constraints:
Evidence packet / access date:
System A / mode / date:
System B / mode / date:
Criterion A (0–2) B (0–2) Evidence / failure note
Task completion
Factual support
Reasoning fit
Boundary honesty
Omission risk
Usability
Safety / reversibility
Disqualifier found:
Independent checks run:
Decision and human edits:
Review or rollback trigger:
Scale the process to the stakes
For a dinner idea, spend two minutes: constraints, obvious errors, usability. For public research, technical instructions or meaningful spending, verify material claims and test edge cases. For medical, legal, financial, employment, safety or rights-affecting decisions, AI comparison is not a substitute for a qualified professional or required institutional review.
The Reversibility Test helps set the depth. The more people affected, the harder the recovery and the less visible the failure, the stronger the evidence and documentation should be.
The bottom line
Do not choose an AI answer by confidence, politeness, length, brand or the number of citations. Define the job. Compare requirements. Extract claims. Verify what matters. Test a case the answer did not select. Then record why the chosen output survived.
The winner is not the answer that sounds most like knowledge. It is the answer whose important parts remain useful after contact with evidence.
Official evaluation references
- NIST, AI Risk Management Framework, including governance, measurement and context-of-use.
- NIST, AI RMF Playbook, with actions for testing, evaluation, verification and validation.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (PDF), including confabulation and information-integrity risks.
- OpenAI, official evaluation guidance, on task-specific criteria and repeatable evals.
END OF FIELD GUIDE 042
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.