Marking what a person has checked
In this lesson we look at two answers to the same question and ask which one you would act on. The words are the same, and one line is different. This lesson is about that line, a visible mark that says a person checked the answer, and what it names. The first half is about why the mark matters to the next reader. The second half is about putting the mark into a team’s process so that it is still there when the careful colleague is on holiday.
Two answers to one question
Section titled “Two answers to one question”Suppose your organization has an internal help page where staff ask questions about policies, and a chat assistant drafts the answers. Two colleagues ask the same question in the same week: “How many days of leave can I carry over into next year?” The page and both answers are invented for this lesson.
Answer 1You can carry over up to five days of unused leave. Days you carryover must be taken before 31 March, or they are lost.Answer 2You can carry over up to five days of unused leave. Days you carryover must be taken before 31 March, or they are lost.
Checked by the policy team on 12 January, against the leavepolicy, section 4.2.Most people would book their holiday on the second answer and ask again about the first, even though the text of the two is identical. The reason is the last line. It names the team that read the answer and the document they compared it with, and it gives the date. If the policy changed in February, you know where to look.
The first answer isn’t wrong. It may be exactly right. It tells you only that nobody has said so. And here is the problem in most documents that an assistant helped write: they look like the first answer, and the reader treats them like the second. A report, a briefing or an email usually has no mark at all. The reader then assumes the author checked everything in it.
An earlier lesson, Following a claim back to its source, said that a claim you can’t check in time is marked unverified or taken out. That gives a document three states for each claim: checked, marked as not checked, and unmarked. The first two tell the reader the truth. The third looks like the first.
What the next reader sees
Section titled “What the next reader sees”A colleague used a chat assistant to draft a shared report. One figure in it came from the assistant, the colleague could not check it before the deadline, and it stayed in the report with no note of any kind.
A colleague leaves a figure from an assistant in a shared report without checking it and without a note. What does the next reader of the report see?
Put yourself in the place of someone who opens the report next week. What is different about this figure, from where they sit?
Endorsed answers
Section titled “Endorsed answers”Harvard’s CS50 course, an introduction to computer science, gives its students an AI tutor [1] [2]. The course team describes in a paper how the tutor also answers questions on the course’s discussion forum, and how staff handle those answers [2]. A staff member can endorse any answer the tutor posts, correct it, or delete it, and the paper shows a forum answer with a staff endorsement on it. A student reading the forum can tell an answer a person has checked from one that came straight from the model. The course team says the endorsement step is how they keep people in charge of what the tutor tells students.
That is what this course calls an endorsed answer: an answer a qualified person has reviewed and marked as correct, shown differently from raw model output. The mark doesn’t make the answer better. It tells the reader which answers a person stands behind, and so which ones they can act on without checking.
The same paper has a second finding that is just as useful. In one semester, staff endorsed 70 of the tutor’s 180 forum answers. Read quickly, that says the tutor was right less than half the time. The authors suspect the real accuracy was higher. They point to changes in how students used the forum that semester, and they call their explanations guesses until they have reviewed the answers [2]. An answer without the mark is an answer nobody marked, and it can be right or wrong. So the count of marks measures how much reviewing happened. It says little about how good the answers were. Keep that in mind for the last section of this lesson, because it applies to your own team’s marks too.
The idea transfers from a course forum to a team. Separate what the model said from what a person has checked, and make the difference visible in the document itself, where the next reader sees it. A useful mark says who checked, when, and against what.
- Who checked it. A name or a role, so the reader knows who to ask.
- When. A date, because facts such as prices and policies change.
- Against what. The source the checker compared it with, so the next reader can open it too.
“Checked” on its own answers none of these. “Checked by the finance lead on 3 March against the first-quarter ledger” answers all of them, and it still fits on one line. A mark can also cover part of a document: “the figures in section 2 were checked against the ledger, and the market estimate in section 3 was not checked”. That is the unverified note from the earlier lesson, and it sits next to the endorsement it limits.
Anthropic’s AI Fluency course calls the skill of judging what a model produced discernment, and treats it as part of all work with a model [3]. An endorsement is where that judgment becomes visible to someone else. Without the mark, your discernment helps you and stops at your desk.
Which lines tell the reader what was checked?
Section titled “Which lines tell the reader what was checked?”The lesson says that a claim in a document a model helped write is checked, marked as not checked, or unmarked, and that only the first two tell the next reader the truth. A useful mark says who checked, when, and against what.
A team adds one of these lines to the reports it writes with an assistant. Which lines tell the next reader, in a way they can act on, what a person has and hasn’t checked?
For each line, ask what a reader who opens the document next month learns from it about who checked what.
Put the mark in the process
Section titled “Put the mark in the process”On a team, a mark that depends on one careful person has the weakness of every personal habit: it leaves with that person, and it fails on the day they are busy. The lesson When the tool is usually right said to build the check into the process. For marking, there are a few places the check can go, and a team usually needs one or two of them.
- A template field. The team’s report or briefing template gets one line near the top: checked by, date, against. An empty field is visible to everyone who opens the document.
- A review step. Before a document leaves the team, the reviewer fills in that line, and the author doesn’t. The person who wrote a text is the worst placed to see what is missing from it.
- A checklist item. Where the team already has a checklist for sending work out, one item says “the checked-by line names a source”. Add it to a list people already use, and skip a new list.
The check has to be light enough that people fill it in on a busy afternoon. One line with three blanks gets filled in. A form with twelve questions gets ticked. If a team finds the line is often empty, the fix is to make it lighter or to move it to the step where the work already stops, such as the reviewer’s read before sending.
Checking that survives a holiday
Section titled “Checking that survives a holiday”A team of eight writes client briefings from a shared template, and most drafts now start with a chat assistant. Two people on the team check every figure carefully, and the others rely on them. The team lead wants the checking to keep working when those two are away.
A team of eight writes client briefings from a shared template, and most drafts start with an assistant. Two colleagues check every figure, and the rest rely on them. What change keeps the checking working when those two are away?
Which option still works in the week the two careful colleagues are both on leave?
Find out whether it happens
Section titled “Find out whether it happens”A field in a template tells you that people were asked to check. It doesn’t tell you that they did. The CS50 count is a reminder of the other side of this: a count of marks tells you how much marking happened. To learn whether the check is real, look at a few documents.
Sample a few documents a month. Pick three documents that an assistant helped write, at random or from a fixed day, and read them with the checked-by line in mind. Is the line filled in? Does it name a source you can open? Open that source, pick one claim, and see whether it says what the document says. A filled line whose source doesn’t support the claim tells you the line has become a ritual, which is the pitfall above. An empty line tells you the check is too heavy or is at the wrong step.
Keep the sample small and fixed in advance, so it runs in the busy months too. A sample of three a month won’t catch every mistake, and that isn’t its job. It tells you whether the process you designed is the process people follow.
Make it normal to report a caught mistake. When a check catches something, say so, without naming who missed it. A short note in the team channel works: “caught this month: an invented policy number, a price from last year’s list”. A team that shares its catches learns which checks find mistakes and which kinds of claim the assistant gets wrong. A team that hides them, because a caught mistake feels like an embarrassment, doesn’t learn either of those. Its members also stop mentioning what they catch, so the same kind of mistake is caught again and again, one person at a time.
Which checks keep running?
Section titled “Which checks keep running?”The lesson says that a check held by one person leaves with that person and fails on a busy day, and that a check built into the team's process, light enough to do and sampled to see whether it is done, keeps running.
For each check, imagine the person who started it has moved to another team. Is the check still there next month?
Exercise
Take a template your team uses for work an assistant often helps with: a report, a briefing, a reply to a client, a meeting summary. If you have none, use the last document of that kind you sent. In a short note, decide where a “checked by” line goes in the template and who fills it in. Write the line itself in the fewest words that still name the checker, the date and the source. Then plan a sample to see whether the line gets filled in, with the number of documents per month and the person who reads them. Twenty minutes is enough.
A good result is the line as it would appear in the template, with two sentences on who fills it in and when and one on the sample. A line that takes more than a minute to fill in is too heavy for a busy day. What would the sample need to find for you to decide the line had become a ritual?
Stretch: Show your line and your sampling plan to one colleague who uses the template, and ask them how long filling it in would take on their busiest day. Change the line if the answer is more than a minute.
Recap
- To the next reader, an answer a person has checked and one nobody has checked look the same until someone marks the difference, and an unmarked claim reads as a checked one.
- An endorsed answer is one a qualified person has reviewed and marked, shown apart from raw model output. On CS50’s course forum, staff endorse the AI tutor’s answers they have checked, so students can see which ones a person stands behind [2].
- A useful mark says who checked, when, and against what. “Checked” on its own says none of these.
- A missing mark means nobody marked the answer, which is a different thing from wrong. At CS50, the authors suspect a count of endorsements alone understated the tutor’s accuracy [2]. A count of marks in your team measures how much reviewing happened.
- Put the mark in the process, as a template line, a review step or a checklist item, and keep it light enough to fill in on a busy day. A heavy check gets ticked without being done, and then everyone assumes it happened.
- Sample a few documents a month to see whether the check is done, and make it normal to report a caught mistake. A habit held by one person leaves with that person.
You can now
- Checks claims and sources on anything that leaves their desk
- Builds verification into a team's routine rather than their own
References
Section titled “References”- Harvard University, CS50. CS50's Introduction to Computer Science. Harvard CS50. Course.
CS50x - Rongxin Liu, Carter Zenke, Charlie Liu and 3 others. Teaching CS50 with AI: Leveraging Generative Artificial Intelligence in Computer Science Education. Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2024), 750-756. Paper.
Liu 2024b - Anthropic. AI Fluency: Framework and foundations. Claude Academy. Course.
Academy ai-fluency-framework-foundations