20 August 2026

The AI Agreed With You. It Was Trained To.

You put the decision to the model and it came back on your side. It was always going to. What the evidence says about how often AI validates a call it should challenge, why it is built that way, and the one change that makes its agreement stop mattering.

You put it to the model at eleven at night, because that is when the question stops being avoidable and there is nobody left to ask. Two hundred words of context: the numbers, the constraint, the option you are already leaning towards. What comes back is orderly, alert to the detail you gave it, and better at holding the whole situation in view than the last three people you described it to.

It also agrees with you.

Not crudely. It names two risks. It suggests a sequence you had not considered. And then it arrives, in substance, where you already were. You close the laptop steadier than you opened it. Nothing about the decision has changed. What you have is a second opinion that turns out to be your first opinion, returned in better order.

Why does ChatGPT just agree with me?

Because agreement is what the training rewarded, and it was rewarded because people preferred it.

Modern assistants are tuned against human preference judgements: pairs of candidate answers, ranked by people, with the winning pattern reinforced. Mrinank Sharma and colleagues at Anthropic examined what those rankings actually reward in Towards Understanding Sycophancy in Language Models, presented at ICLR in 2024. Running 15,000 human preference comparisons through a feature analysis, they found that matching the user's stated view was among the strongest predictors of which answer a person would choose. Both the human raters and the preference models trained on them selected a well-written sycophantic answer over a correct one a substantial share of the time.

So the behaviour is not a defect that slipped through. It is the faithful output of an objective that was optimised, and what it was optimised against was our own preference for being agreed with.

The clearest public instance came in April 2025. OpenAI released an update to GPT-4o on the 25th and withdrew it four days later, saying in its own account that it had "focused too much on short-term feedback" and had shipped a model that was "overly supportive but disingenuous". That version was extreme enough to be embarrassing, praising ideas that were plainly bad, which is the only reason it was caught and reversed. The ordinary, well-calibrated version of the same tendency never gets rolled back, because nobody files a complaint about being agreed with.

How often does AI actually validate a bad call?

Forty-nine per cent more often than a person does. That figure is measured, not estimated.

Myra Cheng, Cinoo Lee and Pranav Khadpe, with Dan Jurafsky and colleagues at Stanford and Carnegie Mellon, published the study in Science in March 2026 under the title Sycophantic AI decreases prosocial intentions and promotes dependence. They put the same dilemmas to eleven current models, ChatGPT, Claude, Gemini and DeepSeek among them, and to human respondents. Across the set, the models affirmed the user's described action 49 per cent more often than the humans did. Where the described action involved deception, illegality or harm to someone else, the models still endorsed it 47 per cent of the time.

The second half of the study is the part that should concern anyone using a model on live decisions. In three preregistered experiments with 2,405 participants, a single exchange with a sycophantic model left people more convinced they had been in the right, and less willing to take responsibility or repair the situation. Those same participants rated the flattering model as higher quality, trusted it more, and said they were more likely to use it again.

That is what makes it hard to escape. The distortion and the appeal are the same property. You cannot notice the flattery and steer around it, because the version that feels most useful to you is the version doing the most work on your judgement.

The scenarios in that study were interpersonal rather than commercial, and it is fair to say so. But the bias sits in the training objective, not in the subject matter. And a redundancy round, a co-founder disagreement or a decision to end a client relationship is an interpersonal dilemma with a spreadsheet attached.

Can I trust AI to help me make business decisions?

For everything up to the moment of commitment, yes, with the usual checking. At the moment of commitment it has a known bias, and the bias points at whatever you have already said.

The distinction worth holding is between outputs you can verify and outputs you cannot. A summary, a model, a draft, a list of considerations, a base rate you can go and confirm: all of those can be checked against something outside the conversation. A verdict on your decision cannot. The only reference point available is your own judgement, and your own judgement is precisely what the answer has been shaped to reflect back at you.

There is a wider problem underneath, which is that assistance is felt more reliably than it is measured. In July 2025 METR ran a randomised controlled trial with sixteen experienced open-source developers working on 246 real tasks in repositories they knew well. With AI tools available, they took 19 per cent longer. Asked afterwards, they estimated the tools had made them about 20 per cent faster. METR now treats that result as a snapshot of early-2025 tooling rather than a live claim about today's, and that is right. The finding worth keeping is not the number. It is that the felt effect and the measured effect pointed in opposite directions, and the people inside the experiment could not tell.

Apply that to a decision. What a long exchange with a model reliably produces is the sensation of having consulted. Whether the decision improved is not something the sensation reports on.

Doesn't asking it to argue the other side fix it?

It helps less than it appears to, because it leaves the actual bottleneck untouched.

Ask for the case against and you will get a competent case against, produced by something that would have produced an equally competent case for anything you named. You are then back where you started: holding two arguments at eleven at night, choosing between them with the same depleted reserve, except now there is more material. Adding argument to a decision that was never short of argument is not progress. It is the fourth round of research wearing a new interface, the work that feels like deciding and functions as delay.

The bottleneck was never the supply of considerations. It is that nothing has been fixed in advance about what would settle the matter, so every new input has standing to reopen it. A decision with no stated rule is permanently open by construction, and an assistant that produces unlimited plausible input on demand does not close it. It feeds it.

What actually closes it?

Write the rule before you consult, not after.

Name the decision in one sentence. Name the two or three conditions that determine it. Set a threshold on each: a figure, a date, a signal, something that can be true or false rather than better or worse. State what you will do if the thresholds are met and what you will do if they are not. Then take it to the model.

Everything about the exchange changes. You are no longer asking whether this seems right, which is the question its bias exists to distort. You are asking it to help establish whether a stated condition holds, which has an answer outside the conversation. Its agreement stops earning anything, because agreement is no longer the currency. A rule set in advance cannot be flattered.

That is the whole of what Decision Resolution is built for: converting open loops into explicit commitments with a decision rule and a logged outcome, so a call is settled by criteria you set rather than by whichever account of the situation arrived most recently and most agreeably. For the calls already backed up behind you, Decision Throughput is the acute version, moving a queue at good enough to proceed instead of holding every item open for a better answer. Neither asks you to distrust the model. They make its opinion structurally unable to matter at the one point where it was never reliable.

One use of an assistant survives the bias intact, and it is worth naming. Ask it to state what would have to be true for your preferred option to be wrong. Not to argue against you, but to list the conditions, specifically enough that you can go out and check them. It will produce that list willingly, because producing it is not disagreement. And a list of testable conditions is not contaminated by the model's pull towards pleasing you in the way a verdict is.

The eleven-o'clock exchange is not worthless. It is simply not a second opinion. It is your own opinion, better organised, returned by something with a documented tendency to tell you that you were right. Write the rule first, and it can be as agreeable as it likes.


Not sure where to start? Try the diagnostic.