Checking habits that run when you are in a hurry
In this lesson we find the row that doesn’t add up in a table a chat assistant wrote from three months of ticket counts. The check that finds it is one of four, and the four together are a routine short enough to run in the ten minutes before a meeting. Earlier lessons in this course showed what an invented specific looks like and why attention to a good tool fades. The routine here is what you do about both, and the last section sets how deep each check goes from the cost of being wrong.
A table sent in a hurry
Section titled “A table sent in a hurry”It is ten minutes before a meeting. A colleague pastes three monthly reports into a chat assistant and asks for “a table of tickets closed per team, with a total per team and a grand total”. The assistant answers in a second, and the table goes into the slides as it came back. The teams and the numbers are invented for this lesson, and the assistant’s table is the author’s.
| Team | Jul | Aug | Sep | Total |
|---|---|---|---|---|
| Design | 124 | 98 | 111 | 333 |
| Support | 82 | 86 | 90 | 258 |
| Marketing | 145 | 162 | 139 | 466 |
| Operations | 67 | 71 | 69 | 207 |
| Grand total | 1,264 |
The table reads well. The columns line up, the totals are in the right place, and every number is one your colleague could have typed. One row adds up wrong. Before reading on, add up the rows yourself and decide which.
Which row does not add up?
Section titled “Which row does not add up?”A chat assistant wrote a table of tickets closed per team over three months, with a total per team. Design has 124, 98 and 111 with total 333. Support has 82, 86 and 90 with total 258. Marketing has 145, 162 and 139 with total 466. Operations has 67, 71 and 69 with total 207. One of the four totals is not the sum of its three months.
Which row’s total isn’t the sum of its three months?
Add the three months of each row and compare the sum with the total column. Which row is off, and by how much?
Put the routine in order
Section titled “Put the routine in order”The lesson gives a routine of four checks that run together, in one order, on every model output that leaves your desk.
- Read the output against the request
- Mark the claims you did not supply
- Open one link or source it cites
- Redo one calculation by hand
What do you compare the output with first, and which check needs a number to redo?
A short program written for this lesson adds every row again and compares the sum with the total the assistant wrote. This is what it prints.
ok Design written 333 recomputed 333ok Support written 258 recomputed 258MISMATCH Marketing written 466 recomputed 446ok Operations written 207 recomputed 207MISMATCH Grand total written 1264 recomputed 1244rows that do not add up: 1 of 4 (Marketing)Marketing closed 446 tickets, and the table says 466. The grand total is wrong by the same 20, because the assistant added its own row totals rather than the months. One slip, two wrong numbers on the slide, and nothing in the way the table looks tells you which two. The model predicted the totals the way it predicts every other number in the table, and 466 is a likely number. It isn’t the sum. The lesson What an invented fact looks like called this specificity without a source. A total is a specific, and its source is the arithmetic, which you can redo in the time it takes to read the row.
The routine
Section titled “The routine”The check that caught the Marketing row is one of four. Together they’re what checking habits means in this course, and they take a few minutes on a page of output.
- Read the output against the request. Put the prompt next to the answer and read them together. Did you ask for tickets closed, and is that what the column says? Did you ask for a total per team, and is there one for every team? A model answers a nearby question as fluently as the one you asked.
- Look for claims you didn’t supply. Every number, name, date and reference in the output came from what you pasted, or it came from the model. Mark the ones you didn’t supply. Each is a claim to check or to take out, as the lesson on invented facts showed.
- Open one link. If the output cites anything, a page, a document, a section number, open one and read whether it says what the output says. One is enough to tell you whether the rest deserve the same.
- Redo one calculation. Pick one row, one percentage, one date difference, and do it by hand or in a spreadsheet cell. If it holds, the others probably hold too. If it fails, none of them can go out unchecked.
They’re one routine because they run together and in this order on everything that leaves your desk. Output that stays with you, a brainstorm or a note to yourself, may skip them, and the next section says how to tell. For output that goes to someone else, you don’t decide each time whether to run them. Anthropic’s AI Fluency course calls the skill of judging what a model gave you discernment. It pairs that skill with the skill of asking well, which it calls description, and has you move between the two in a loop [1]. The routine here is discernment made small enough to run when you have ten minutes.
Depth scales with stakes
Section titled “Depth scales with stakes”The routine is the same every time, and the depth of each check isn’t. The last lesson, When the tool is usually right, said to build the check into the process and to set its depth from what happens when the output is wrong. Fluency is constant and consequences are not, so the cost of a mistake is the only thing that should change the effort. The depth comes from whether the output can be undone, and from who reads it and whether they check it themselves.
A brainstorm list of names for an internal project isn’t checked at all. A wrong entry doesn’t cost anything, and you throw most of them away. A summary of a meeting for your own notes gets one read against the request, because you were in the meeting and you’ll notice a wrong point when you use the notes. A briefing with three cited reports for your manager gets every link opened and every number redone, because your manager acts on it and does not open the reports. A price table on the public web page gets every value read by a second person, because it cannot be recalled once a customer has seen it, and the customer checks nothing. The list, the notes, the briefing and the price table read as fluently as each other. The depth of the check comes from the cost of the mistake.
One more case sets the depth to the maximum whatever the size of the task. A domain you cannot judge is high risk even when the output is one paragraph, because you cannot see the mistake. A one-paragraph legal clause goes to the legal team, however well it reads.
How deep does each check go?
Section titled “How deep does each check go?”The lesson gives four checks that run on everything a model writes before it leaves the learner's desk, and says that how far each check goes depends on what happens if the output is wrong: whether it can be undone, and whether its readers will check it themselves.
Match each output to the depth of checking it gets before it leaves your desk.
For each output, ask what happens if it is wrong: can you undo it, and does its reader check it?
Harvard’s CS50 course gives its students an AI helper [2] [3]. The helper also answers questions on the course forum. A staff member can endorse, amend or delete each of those answers, and the forum shows which ones a staff member endorsed, so a student can tell a checked answer from a raw one [3]. The idea transfers to your desk. What the assistant wrote and what you have checked are two different things, and the depth of checking is what turns one into the other. A later lesson in this course is about making that difference visible to the next reader.
Which output gets the deeper check?
Section titled “Which output gets the deeper check?”The lesson says the depth of checking a model's output gets comes from what happens if the output is wrong: whether it can be undone, and whether its readers will check it themselves. The learner compares four outputs of similar length and fluency.
The same assistant wrote these four outputs for you today. Which one gets the deepest check before it goes anywhere?
Which of the four cannot be undone, or goes to a reader who will not check it?
Exercise
Pick one document from your week that an assistant wrote or drafted: a summary, a table, a reply, a report section. Run the four checks on it in a note. Read it against the request you gave and write down any question it answered instead. List the claims you did not supply. Open one link or reference, if it has any, and write what the source says next to what the document says. Redo one calculation, if it has any, and write both numbers down. Fifteen minutes is enough.
A good result is a note with four headings and something under each, even if that something is “no links” or “no numbers”. Which of the four checks would you have skipped, and would the document have gone out with what it found?
Stretch: Do the same for a second document, one going to a reader outside your team, and compare how deep each check went.
Recap
- Before a model’s output leaves your desk, read it against the request and look for claims you did not supply. Then open one link and redo one calculation.
- A total is a specific like any other, and its source is the arithmetic. One wrong row total in a table can put two wrong numbers on a slide, and the table looks the same either way.
- The check you are about to skip is the one you run. Errors get through at the moment a routine loses a step, which is when you are busy.
- The routine is the same every time, and its depth comes from the cost of being wrong: whether the output can be undone, and whether its readers check it themselves.
- A habit is better than a rule because it runs when you are in a hurry, which is when a rule gets skipped.
You can now
- Checks claims and sources on anything that leaves their desk
- Matches the depth of checking to the cost of being wrong
References
Section titled “References”- Anthropic. AI Fluency: Framework and foundations. Claude Academy. Course.
Academy ai-fluency-framework-foundations - Harvard University, CS50. CS50's Introduction to Computer Science. Harvard CS50. Course.
CS50x - Rongxin Liu, Carter Zenke, Charlie Liu and 3 others. Teaching CS50 with AI: Leveraging Generative Artificial Intelligence in Computer Science Education. Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2024), 750-756. Paper.
Liu 2024b