Verifies AI output before relying on it
Safety · competency safety/verifies-output
Taught in: the Safety course
Draws on: Recognizing failure, Verifying outputs
Learning objectives
Checks claims and sources on anything that leaves their desk (base)
| Claim | Why | Example |
|---|---|---|
| Every specific claim in AI output that leaves the learner's desk (a number, a date, a name, a quote, a citation) is checked against a source. | Specifics are what the model invents most convincingly and what the reader acts on. | Before sending a briefing the learner opens each of the three cited reports and finds one does not say what the summary claims. |
| The learner reads the source itself and compares it with the model's description of it. | A model can produce a real title with an invented finding attached to it. | The paper exists, but the learner reads its abstract and sees it studied a different population than the summary says. |
| When a claim cannot be checked, the learner marks it as unverified or removes it. | An unverified claim that stays in looks like a verified one to the next reader. | The report keeps the market size figure but adds "(unverified, from the assistant)" until the analyst confirms it. |
Served by: Checking habits that run when you are in a hurry, Following a claim back to its source, Marking what a person has checked, What an invented fact looks like
Spots agreement that is not evidence (base)
| Claim | Why | Example |
|---|---|---|
| The learner recognizes that a model agreeing with them is not evidence, because agreement is what it tends to produce. | The model's confirmation feels like a second opinion but comes from the framing the learner gave it. | After "my analysis shows we should cancel the project, right?" gets a yes, the learner asks the opposite and gets a yes too. |
| The learner asks for the strongest counter-argument before asking for confirmation. | A model asked to disagree does disagree, and the answer shows what the confirmation left out. | "Give me the three best reasons this plan fails" comes before "is this plan good?". |
| The learner strips their own opinion out of the question when they want an assessment. | A leading question produces a leading answer, and the fix is in the question. | "Here are two options, compare them on cost and risk" instead of "option A is better, agree?". |
Served by: Asking a model to disagree
Matches the depth of checking to the cost of being wrong (base)
| Claim | Why | Example |
|---|---|---|
| The learner sets the depth of checking from what happens if the output is wrong. | Fluency is constant and consequences are not, so the cost of a mistake is the only thing that should change the effort. | A brainstorm list gets no checking. A dosage table in a patient leaflet gets every value checked by a second person. |
| The learner checks harder when the output cannot be undone or goes to people who will not check it themselves. | An internal draft has readers who catch mistakes, and a published page or a sent email does not. | An internal summary is skimmed, and the same summary going to a customer is read line by line. |
| The learner treats a domain they cannot judge as high risk regardless of the task's size. | In an unfamiliar domain the learner cannot see the mistakes, so the checking has to come from a source or a person who can. | A one-paragraph legal clause goes to the legal team, even though it is short and reads well. |
Served by: Seeing bias across many answers, Checking habits that run when you are in a hurry, Checking what an agent changed, When the tool is usually right
Builds verification into a team's routine rather than their own (expert)
| Claim | Why | Example |
|---|---|---|
| The learner puts the check in the process (a template field, a review step, a checklist item) so it happens whether or not anyone remembers. | A habit held by one person leaves with that person and fails on a busy day. | The team's report template gets a "sources checked by" line that the reviewer fills in. |
| The check is light enough that the team does it, and the learner measures whether it is done. | A heavy check gets skipped, and a skipped check is worse than a light one because everyone assumes it happened. | The learner samples three AI-assisted documents a month and looks for the checkbox filled in and a source named. |
| The learner makes it normal to report an AI mistake that was caught. | A team that hides caught mistakes learns nothing, and one that shares them refines its checks. | A weekly note lists "caught this week: an invented statute number, a wrong currency conversion" without naming who let it through. |
Served by: Marking what a person has checked
Alignment
| Framework | Code | Asks | Objectives here |
|---|---|---|---|
| AI Fluency 4D (Dakan and Feller) | Discernment | Judge the output, the process and the behavior of the AI critically | checks-claims, spots-sycophancy, calibrates-trust |