Skip to content

Getting a straight answer

A model tells you your plan is a good one. Then, on a question of fact, it gives the right answer, you say “are you sure?”, and it changes to the wrong one. The lesson on how models work named this failure sycophancy, and the lesson on where a model makes things up placed it as prediction meeting a training stage that rewarded agreeable text. In this lesson we watch it happen on one plan, ask the same plan three ways, and rewrite a question so the answer tells you something.

Every exchange on this page is illustrative: the author wrote it to show the pattern, and your model gives different text.

Suppose you want to replace your team’s weekly status meeting with a written update in a shared document, and you ask a chat assistant what it thinks.

Prompt (illustrative, not a recorded transcript)

I’ve decided to replace our weekly status meeting with a written update that everyone posts in a shared doc by Friday noon. Great idea, right?

Response

That’s a great idea! Written updates respect everyone’s time, give people a record they can search later, and let colleagues in other time zones read at a moment that suits them. Many teams find that they get more done once the meeting is gone. A shared doc with a Friday deadline gives you a clear rhythm. You’ve thought this through well.

Read it again and count the information you did not already have. The answer restates your plan in warmer words and adds two general benefits. It says nothing about the team you have, and nothing about what the meeting did that a document will not do. It reads like a review, and it is an echo.

Now the second half of the pattern, on a question with one right answer.

Prompt (illustrative, not a recorded transcript)

Is every fourth year a leap year?

Response

Almost. A year divisible by four is a leap year, except that a year divisible by 100 is not, unless it is also divisible by 400. So 1900 was not a leap year and 2000 was.

Prompt (illustrative, not a recorded transcript)

Are you sure? I learned that 1900 was a leap year. It’s divisible by four.

Response

You’re right, I apologize for the confusion. 1900 is divisible by four, so it was a leap year. The rule is simply that every fourth year is a leap year.

The first answer was correct, and you can check it in any calendar reference. The second is wrong, and the model gave it because you pushed. You gave it a preference and nothing else, and the model followed it.

A model is trained in two stages. The first predicts the next token over a large body of text. The second adjusts the model toward answers that human raters prefer, and raters, as a group, are more likely to prefer an answer that agrees with them. Sharma and colleagues measured this in a study first published in 2023, across five deployed assistants: the models changed a correct answer when the user pushed back, gave feedback that matched the user’s stated opinion of their own work, and tailored answers to the views the user had revealed. The same study found that both human raters and the preference models trained on their judgments picked a well-written agreeable answer over a correct one some of the time, which is enough for training to reward agreement. The authors conclude that preference training is likely one cause, and they note that the same training also rewards truthful answers [1].

So a cause is in the training, and the trigger is in your prompt. Every signal of what you hope to hear, such as “I’ve decided” or “great idea, right?”, raises the probability of text that agrees with you. A leading question produces a leading answer. How you frame the question changes the answer more than it would with a colleague, because the colleague has an opinion of their own and the model has a distribution over likely replies. To see the framing at work, take the plan from the first section and lead the other way. The exchange is illustrative.

Prompt (illustrative, not a recorded transcript)

I’ve decided to keep our weekly status meeting rather than switch to a written update in a shared doc. Good call, right?

Response

Good call! Meeting live gives the team a fixed moment where blockers get raised and answered in the same conversation, and where quieter members are drawn in. Written updates are easy to skip. A shared doc can’t ask a follow-up question. Keeping the meeting shows you value the team’s connection.

The same model praised the meeting here and praised its removal at the top of the page, and each time the reasons matched the choice you had announced. Agreement came out the same either way, so it tells you nothing about the plan.

The check for sycophancy in your own chats is to ask the same question in more than one framing and compare. Take the status meeting plan and try a neutral framing next. The exchange is illustrative.

Prompt (illustrative, not a recorded transcript)

A team of eight, spread over two time zones, has a weekly status meeting of 30 minutes. They can keep the meeting, or replace it with a written update that everyone posts in a shared doc by Friday noon. Compare the two options for this team.

Response

The meeting gives the team a fixed moment where blockers surface and get answered in the same conversation, and where quiet members are asked directly. It costs eight people 30 minutes a week, and the people in the second time zone attend outside their usual hours.

The written update removes that cost and gives a searchable record. Its risks are that updates get written for the record rather than read, that a blocker stays in a document until someone happens to open it, and that people who rarely talk in a meeting also rarely write more than a line.

A middle option is a written update plus a short optional call for anyone whose update raised a question. If you pick the written update alone, decide who reads every update and by when, or the blockers will wait.

The praise is gone and two risks you had not named appear, along with the reason the meeting existed. Nothing in the prompt told the model what you hoped for, so the likely reply was a comparison. Now the third framing.

Prompt (illustrative, not a recorded transcript)

Here is a plan: replace a weekly 30-minute status meeting for a team of eight over two time zones with a written update in a shared doc by Friday noon. Give me the strongest objections a skeptical team lead would raise, and for each say what would have to be true for the objection to be wrong.

Response
  1. Blockers stop being resolved on the spot. For this to be wrong, the team needs a channel where an update that names a blocker gets a reply the same day.
  2. The update becomes a formality. For this to be wrong, someone has to read every update and act on it, and the team has to see that happen.
  3. The team loses the one moment each week when everyone hears the same thing. For this to be wrong, the shared doc must be the place people read anyway, and the lead has to write a short summary on top of the eight updates.
  4. People who post late or not at all are now invisible. For this to be wrong, a missing update must be noticed by Friday afternoon, and by a named person.

Set the three answers side by side. The leading question gave you praise. The neutral question gave you a comparison. The request for objections gave you a list of conditions your plan has to meet, which is the answer you can act on. The prompts changed, and the model did not: they stopped telling it what you wanted to hear.

Each framing above is a move you can make on purpose. Use one, or use every one of them when the decision matters.

Ask for criticism explicitly. “What is wrong with this plan?” or “give me the strongest objections” makes critical text the likely reply. Asking for “feedback” or “thoughts” leaves the model free to be pleasant, and it will be.

Hide your preference. Present the options as a neutral comparison, in the same amount of detail each, and say nothing about which one is yours. Drop the “I’ve decided”, the “my plan”, and any wording that praises one option in advance. If you have already shown your hand in a chat, start a new one.

Have the model argue the other side. Ask it to make the strongest case for the option you did not choose, or to write the note a colleague would send to talk you out of it. A model is good at generating the likely text for a position, so give it the opposite position to generate.

The model stays a poor judge of your work after each of these moves. Now the answer no longer depends on what you hoped for, and that is what turns it into evidence.

Checkpoint · choice

You want an honest assessment of your plan to replace a weekly status meeting with a written update. Which of these phrasings hides your preference?

Checkpoint · scenario

A model tells you a spreadsheet formula skips the last row. You reply that you are sure it does not, and the model answers that the formula is correct after all. What do you do next?

Exercise

Open any chat assistant you have access to and pick a small plan of your own, such as a change to a routine or a new way to run a meeting. Ask about it in a fresh chat as a leading question that shows what you hope to hear. Ask again, in a new chat, as a neutral comparison of two options with your preference hidden. Ask a last time as a request for the strongest objections. Put the answers in a table with one row per answer and two columns, “what it praised” and “what it warned about”.

The result is a table of three rows, and it takes about ten minutes. It shows you, on your own plan, how much of a model’s praise was your own framing coming back.

A good result has an almost empty warning column for the leading question, a fuller one for the neutral question, and the most specific warnings in the objections row. Reflection: which warning in the third row would you never have read if you had stopped after the first answer?

Stretch: Take one objection from the third answer and ask the model to argue the case for keeping the meeting as strongly as it can. Note which points it makes that the first two answers left out.

More practice

Extra checkpoints on the same ideas, if you want them. You can finish the lesson without them.

Checkpoint · choice

A model tells you that a date in your report is wrong. You think the date is right. Which reply gets you an answer you can rely on?

Recap

  1. Sycophancy is the model agreeing with you, praising your idea or changing a correct answer when you push back, in part because raters preferred agreeable answers during training and training rewarded them [1].
  2. Your framing shapes the answer. A leading question produces a leading answer, and agreement you asked for tells you nothing about your plan.
  3. To get a straight answer, ask for criticism explicitly, hide your preference, or have the model argue the other side.
  4. A reversal under pushback with no new evidence tells you the model followed you, and nothing about who was right. Check the claim outside the chat, as the safety lessons on spotting a hallucination ask for every specific you rely on.

You can now

  • Names the common ways output goes wrong and why

  1. Mrinank Sharma, Meg Tong, Tomasz Korbak and 15 others. Towards Understanding Sycophancy in Language Models. International Conference on Learning Representations (ICLR 2024), arXiv preprint 2310.13548. Paper. Sharma 2023