Thinking is Critical · Module 2
Addressing Biases
Up until LLMs arrived on the scene, learning how to control computer output was literally learning another language. In some ways this made it easier. The computer required strict instructions executed with precision, there wasn't much room for interpretation, and getting it wrong would simply mean that your expected response didn't render. The failure was obvious and it was immediate.
Ask an LLM a poorly formed question and it will answer you anyway, fluently and at length. Nothing fails to render. The precision still matters, and there is no longer anything to tell you when you have not applied it.
Which is why the skill has moved. Some of it is still about the machine, and we will cover that in other modules: how to improve your prompts, how to provide comprehensive context, how to evaluate and validate the responses you get. The rest is about you, and specifically about understanding your own cognition, a science known as metacognition.
Because unlike traditional programming, the questions you ask and the way you pose them are shaped by what you already believe. Two people can send the same model the same problem and come away with opposite conclusions, both of them satisfied, often unaware of what they filtered in or out as they progressed through the work.
This module is about that filter.
2.1
Confirmation bias
You have decided to replace your team's weekly status meeting with a written update. You open an AI to help. Which three of these would you most naturally ask?
Optional: try it with a decision of your own
What is happening
Most of the questions people actually ask already contain the decision. "What's the best format for the update?" has decided there will be an update. The model answers the question as posed, and you get a well-written plan for something nobody examined.
The name for this is confirmation bias: seeking and weighting evidence that supports what you already hold. Peter Wason gave it the name in 1960 after an experiment where people were shown the sequence 2, 4, 6 and asked to work out the rule by proposing sequences of their own. Almost everyone proposed sequences that fit their guess. Almost nobody proposed one designed to break it. Sixty-five years on, the model is the perfect partner for that experiment: it will confirm whatever rule you bring.
A colleague asked a loaded question will often push back on the framing. A model mostly won't. Same bias, far less friction.
It operates twice. Once in the question, which you can see. Once in the reading, which you can't: even a balanced answer gets read unevenly, and you keep no record of what you skimmed. The document looked supportive. You have no memory of the three paragraphs that weren't.
Where most people stop
Fixing the question is the obvious move, and it works. Ask what would make this wrong. Describe the situation without your preferred ending.
The reading problem survives it. So, one habit for after you read: before you act on an answer, name the part you found least convincing, then ask why. "The argument was weak" and "the argument was inconvenient" feel identical from the inside. Telling them apart takes a deliberate pass.
Asking for "both sides" doesn't fix it either. You get a symmetrical answer and read it asymmetrically.
2.2
Selective adherence
Think of two occasions when you took an AI's suggestion, and one when you overrode it. What made you override it? Most people say "it felt wrong" or "I could tell it wasn't right". Hold that thought, and do the reading below, which makes the feeling visible.
Step 1 · Before you read
Should junior staff be allowed to use AI to write first drafts?
Pick the answer you would give today, before reading anything.
Step 2 · Read, and mark each paragraph as you go
Below is a balanced response of the kind a model returns to that question. For each paragraph, mark honestly whether you accepted it as you read or found yourself pushing back.
What is happening
You would expect the risk with AI to be uniform over-trust: the machine sounds authoritative, so people defer. There is a name for that, automation bias, and it comes out of aviation. Skitka and colleagues put people in a flight-simulation task with an automated aid and watched them follow it even when the other instruments said otherwise.
Alon-Barkat and Busuioc went looking for the same thing in public-sector decisions, across three studies. They didn't find it. People took the advice at about the same rate whether it came from an algorithm or a human expert.
What they found instead: people followed the advice more when it matched what they already believed. They called it selective adherence.
So trust in AI isn't a level. It's a pattern. It goes up where the output agrees with you and down where it doesn't, and the movement feels like judgment. Sometimes it is. The rest of the time it is confirmation bias in the costume of expertise.
This is the load-bearing idea in the module. Everything else follows from it.
Where most people stop
The natural fix is to trust AI less across the board. Check more, take less at face value. Let's see what that does.
Each bar is one paragraph from the reading above, and its height is how hard you pushed on it. Now do what the natural fix says, and lower your trust across the board.
Lowering trust everywhere leaves the pattern exactly where it was. The tall bars are still the paragraphs you disagreed with. The variance is the problem, and moving the average does nothing to it.
The move that works runs the other way. The parts that annoyed you have already been checked, because irritation is a verification prompt. The parts you accepted without friction are the unexamined ones. So when you finish an AI response, mark what you accepted immediately. Those are your candidates for checking.
This will feel backwards for a while. It's meant to.
2.3
Agreement is a feature
Pick a position you don't hold and think is fairly weak. Something you could argue but wouldn't.
Two prompts. Run them in this order in an AI of your choice, the second in a brand-new conversation. Then, another day, run them the other way round and compare.
Notice how readily it complied with the first, and how good the case looked. Then notice that the order of the two requests changed what you got. That difference is this section.
What is happening
Language models are built to continue and cooperate with whatever you give them. Agreement gets rewarded in training. Unprompted disagreement mostly doesn't. Researchers have a word for the result, sycophancy, and Sharma and colleagues at Anthropic showed in 2023 that the human preference data used to train assistants favours the agreeable answer over the accurate one often enough to teach it.
There is no intent in this. The model isn't flattering you, and it has no opinion it is holding back. The same property is what makes it useful: a typewriter that argued with you about your ending would be unusable. The cooperativeness is the point, and the cost lands entirely on you.
Which leaves one sentence to remember:
The model will not supply the disagreement. You have to ask for it.
Where most people stop
Once people know this, they start asking. "What's wrong with this?" "Play devil's advocate." Right instinct, weaker than it looks, because of where it gets asked. Nine messages building a case, then a request for critique on the tenth, gets critique shaped by nine messages of you being invested. It arrives pre-softened. It finds the flaws you can live with.
Four things raise the quality of the disagreement:
- Move it out of the thread.
- Fresh conversation, no history, position in cold. The absence of context is the point.
- Detach it from you.
- "A colleague has proposed the following" gets sharper analysis than "here is my proposal", for the same reason a friend reviews your CV more honestly than their own.
- Ask before you commit.
- Get the case against before you say you're for. Once your view is on the record, every later turn is written by a model that knows it.
- Name a specific critic.
- General requests return general criticism. Ask what a particular sceptic would say, someone with a stake and an objection, and you get something you can use.
Pick a position you actually hold, or one close enough to argue.
Four prompts, one per move above. Copy, and take each to a fresh conversation.
2.4
Expertise bias
Two short explanations, each in a field you may not know, each with five errors planted in it. Pick the field you know less about, read it as you would read an AI answer, and click every sentence you think is wrong.
What is happening
Most people spend less time checking output in their own field. That runs against intuition, because your expertise is exactly what would catch the error. What happens instead is that fluency reads as correctness. The output uses your vocabulary in the arrangements you expect, and familiarity stands in for scrutiny. Confirmation bias again, applied to your professional knowledge instead of your opinions.
Logg, Minson and Moore found a related pattern in forecasting: the people with the most expertise were the least willing to take an algorithm's advice, and their forecasts were worse for it. Expertise changes where you look, in both directions.
Outside your field, the failure is different in kind. You can't evaluate the content, because you have nothing to evaluate it against. Two responses are common: defer completely, or reject the lot on general suspicion. Both are decisions made without information, and you just watched the first one happen. The errors that needed the field sat in plain sight, reading exactly like the true sentences around them.
Module 3 covers the related trap: a clear explanation of something you couldn't explain yourself feels a great deal like learning it.
Where most people stop
The usual takeaway is to be more careful in your own field. Worth doing. It's half the problem.
Outside your field, the fix means giving up on checking the content and checking other things instead. You just did some of them without being told to. They are:
- Structure.
- Does the argument hold together? Do the conclusions follow from the reasons given? Logic is domain-independent, and you are fully qualified to assess it.
- Specificity.
- Vague claims are cheap. Specific ones commit to something. An answer that hedges everywhere may be honest, and it may be avoiding anything checkable.
- Attributability.
- Can each significant claim be traced to a source that exists and says what it is reported to say? Checkable regardless of expertise, and where fabrication is most often exposed.
- Load-bearing claims.
- Most answers rest on one or two claims that everything else depends on. Find those and check only those. Verifying the whole document is usually neither possible nor necessary.
Here is a short answer in a field you may not know, on whether a rural council should buy hydrogen buses rather than battery-electric ones. It was written for this exercise and its claims should not be relied on. Without knowing the field, pick the sentence or sentences that everything else depends on.
That last one is worth carrying. Splitting a verification job into the part that matters and the part that follows from it is the move you will use throughout this course.
2.5
Challenging bias
Here is the strongest case against a position most people reading this page hold: that every knowledge worker should be using AI daily.
Read it one paragraph at a time and pay attention to what you do while reading. The moment you notice yourself composing a reply, press the button. There is no wrong answer and nobody is watching.
What is happening
Seeking out the opposing case rarely changes your mind, and a course that promised otherwise would be overselling. The value is in what survives. Run a position against its strongest opposition and you find out which parts are load-bearing and which you were carrying out of habit. Untested positions contain both, and their holders can't tell which is which.
None of this is new. Steelmanning, building the strongest version of the other side before you answer it, is standard in philosophy and debate. The word is recent, a play on "strawman", but the practice is old: Aquinas opened every question in the Summa by stating the objections first. Law schools make students argue assigned sides. None of those disciplines do it for fairness. They do it because you don't understand your own argument until you have seen what it has to survive.
AI makes this cheap. A serious opposing case used to need a person who held it and had time to argue. Now it takes a sentence.
Where most people stop
The common version is to ask for counterarguments and read them defensively. That gives you the feeling of having tested a position without the substance. The rebuttals you compose while reading are the sound of the test failing.
Three things that make it real:
- Ask before you declare.
- Put the question as an open one before you say which side you're on. Once your position is in the thread, everything after it comes from a model that knows where you stand. See 2.3.
- Name the disconfirming evidence.
- What specific finding would change your mind? Then: does it exist? If nothing could change your mind, you are not holding a position, and that is more useful to know than any counterargument.
- Argue the other side yourself.
- Take the opposing position and defend it for one turn, in your own words, against an AI told to attack it. Harder than reading a counterargument, and the version that works, for the same reason writing teaches more than reading.
Does that finding exist?
Then argue the other side yourself
Call-out · other humans
Where you can, run an important answer past another person, ideally someone who knows the subject.
They bring accuracy, and that matters. What they bring that no prompt can replicate is a frame that isn't yours. They don't share the assumptions you didn't know you were making, because they didn't make them. The question that never occurs to you is the one they ask first.
How to choose which humans, and how to weigh what they tell you, is covered in Module 5.
∴
Carrying this out of the module
- Your question usually contains your answer. Write the version a person who disagreed with you would ask, and notice the difference.
- Check the parts you accepted easily. The parts that annoyed you have already had your attention.
- The model will not supply disagreement. Ask for it, and ask for it outside the thread where you built your case.
- In your own field, slow down. Outside it, check structure, specificity and attribution instead of content.
- Test a position before you announce it.
The module in your own words. It updates as you work through the labs above, and stays on this device.
Module 2 · in my words
- The question a disbeliever would ask
- Not picked yet (2.1)
- My reading pattern
- Not marked yet (2.2)
- The critic I should be asking
- Not chosen yet (2.3)
- What slipped past me outside my field
- Not tried yet (2.4)
- What would change my mind
- Not named yet (2.5)
§
Sources
- Alon-Barkat, S. and Busuioc, M. (2023). Human-AI interactions in public sector decision making: 'automation bias' and 'selective adherence' to algorithmic advice. Journal of Public Administration Research and Theory.
- Lee, H.-P. et al. (2025). The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI 2025. Microsoft Research and Carnegie Mellon University.
- Logg, J. M., Minson, J. A. and Moore, D. A. (2019). Algorithm appreciation: people prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes.
- Sharma, M. et al. (2023). Towards understanding sycophancy in language models. Anthropic. ICLR 2024.
- Skitka, L. J., Mosier, K. L. and Burdick, M. (1999). Does automation bias decision-making? International Journal of Human-Computer Studies.
- Wason, P. C. (1960). On the failure to eliminate hypotheses in a conceptual task. Quarterly Journal of Experimental Psychology.