Skip to content

Verifies AI output before relying on it

Safety · competency safety/verifies-output

Taught in: the Safety course

Draws on: Recognizing failure, Verifying outputs

Learning objectives

Checks claims and sources on anything that leaves their desk (base)

ClaimWhyExample
Every specific claim in AI output that leaves the learner's desk (a number, a date, a name, a quote, a citation) is checked against a source.Specifics are what the model invents most convincingly and what the reader acts on.Before sending a briefing the learner opens each of the three cited reports and finds one does not say what the summary claims.
The learner reads the source itself and compares it with the model's description of it.A model can produce a real title with an invented finding attached to it.The paper exists, but the learner reads its abstract and sees it studied a different population than the summary says.
When a claim cannot be checked, the learner marks it as unverified or removes it.An unverified claim that stays in looks like a verified one to the next reader.The report keeps the market size figure but adds "(unverified, from the assistant)" until the analyst confirms it.

Served by: Checking habits that run when you are in a hurry, Following a claim back to its source, Marking what a person has checked, What an invented fact looks like

Spots agreement that is not evidence (base)

ClaimWhyExample
The learner recognizes that a model agreeing with them is not evidence, because agreement is what it tends to produce.The model's confirmation feels like a second opinion but comes from the framing the learner gave it.After "my analysis shows we should cancel the project, right?" gets a yes, the learner asks the opposite and gets a yes too.
The learner asks for the strongest counter-argument before asking for confirmation.A model asked to disagree does disagree, and the answer shows what the confirmation left out."Give me the three best reasons this plan fails" comes before "is this plan good?".
The learner strips their own opinion out of the question when they want an assessment.A leading question produces a leading answer, and the fix is in the question."Here are two options, compare them on cost and risk" instead of "option A is better, agree?".

Served by: Asking a model to disagree

Matches the depth of checking to the cost of being wrong (base)

ClaimWhyExample
The learner sets the depth of checking from what happens if the output is wrong.Fluency is constant and consequences are not, so the cost of a mistake is the only thing that should change the effort.A brainstorm list gets no checking. A dosage table in a patient leaflet gets every value checked by a second person.
The learner checks harder when the output cannot be undone or goes to people who will not check it themselves.An internal draft has readers who catch mistakes, and a published page or a sent email does not.An internal summary is skimmed, and the same summary going to a customer is read line by line.
The learner treats a domain they cannot judge as high risk regardless of the task's size.In an unfamiliar domain the learner cannot see the mistakes, so the checking has to come from a source or a person who can.A one-paragraph legal clause goes to the legal team, even though it is short and reads well.

Served by: Seeing bias across many answers, Checking habits that run when you are in a hurry, Checking what an agent changed, When the tool is usually right

Builds verification into a team's routine rather than their own (expert)

ClaimWhyExample
The learner puts the check in the process (a template field, a review step, a checklist item) so it happens whether or not anyone remembers.A habit held by one person leaves with that person and fails on a busy day.The team's report template gets a "sources checked by" line that the reviewer fills in.
The check is light enough that the team does it, and the learner measures whether it is done.A heavy check gets skipped, and a skipped check is worse than a light one because everyone assumes it happened.The learner samples three AI-assisted documents a month and looks for the checkbox filled in and a source named.
The learner makes it normal to report an AI mistake that was caught.A team that hides caught mistakes learns nothing, and one that shares them refines its checks.A weekly note lists "caught this week: an invented statute number, a wrong currency conversion" without naming who let it through.

Served by: Marking what a person has checked

Alignment

FrameworkCodeAsksObjectives here
AI Fluency 4D (Dakan and Feller)DiscernmentJudge the output, the process and the behavior of the AI criticallychecks-claims, spots-sycophancy, calibrates-trust