A few weeks ago, I wrote a piece on the art of questioning, and someone on LinkedIn aptly remarked that the main idea is dubito, ergo sum: doubting is one of the skills we humans have to hone and preserve in the age of AI. But how do we actually do that? One of the first steps is learning how it is that we skip the step, that we don’t doubt, we don’t ask the right questions, and Argyris’s Ladder of Inference is a valuable manual at that. If asking well means knowing where a model’s gaps hide, decades before Turing there was this guy, Chris Argyris from Yale, who mapped how belief gets manufactured, and his theory was later made popular by Peter Senge in his book The Fifth Discipline. Today, we take a look at what the ladder is, how we can train people (and ourselves) to go up its rungs, and we see some practical examples applied to our everyday coordination practice.
1. Where it all started
Long before anyone worried about a machine sounding believable, there was a deeper problem: a human sounding believable, to themselves and others, about something they hadn’t actually checked. That was Chris Argyris‘s. A business theorist and a pioneer in organisational development, Argyris spent most of his career at Yale, and later Harvard, studying organisations and trying to figure out how does it happen that smart, well-intentioned people keep talking past each other in meetings, defending positions they’d never really examined, like monkeys throwing poop at each other.
He wasn’t writing about artificial intelligence, but about the much older human intelligence problem: how a person moves from noticing something to being certain of something that’s loosely related, and without ever being aware they jumped to conclusions. Argyris called the mechanism the Ladder of Inference, as we were saying, and it stayed a fairly academic piece of organisational psychology until Peter Senge picked it up in the popular The Fifth Discipline and gave it the kind of plain-language treatment loads of valuable academic research need to be understood by the layfolk.
That borrowed fame happened because, when people see the seven rungs laid out, they recognise the climb they’ve done that morning. It names a thing everyone already does, and almost no one notices, which is precisely the quality that makes it useful now, in a very different context than Argyris ever anticipated: not a boardroom disagreement, but an AI answering with far more confidence than the underlying data can support.
2. The Seven Rungs
Argyris’s ladder has seven rungs, and the trick is that you never feel yourself stepping on any of them: you just find yourself standing at the top, with an opinion.
It starts with the pool of observable data: everything that happened, in principle available to anyone. The second rung is selected data: the small slice we actually noticed, filtered by what we were looking for, what we already suspected, or what happened to be in the corner of our eye. From there, we add meaning: we interpret the slice, usually through a cultural or professional lens we didn’t choose and rarely inspect. Meaning breeds assumptions, assumptions harden into conclusions, conclusions calcify into beliefs on how things generally are, and beliefs finally produce action.
Seven rungs. All of them climbed in a split second.
The existence of the ladder isn’t in question: inference is not optional, it’s how cognition works, and nobody has time to verify the pool of all observable data before ordering lunch. We would starve. The interesting part, however, is that we experience only the bottom and the top. We remember what we noticed, and we remember what we concluded, and the five rungs in between vanish from awareness the instant we’re done climbing them. Ask someone why they’re sure a zone is ready for production, and they’ll usually answer you from the top rung — a belief, stated as fact — with no memory of the assumption underneath it, let alone the sliver of selected data the whole structure was built on.
3. The Reflexive Loop
Thing is, the ladder doesn’t run once: it runs in a loop, and the loop feeds itself.
The beliefs we form at the top of the ladder don’t just sit there waiting to be revised: they jump down and take a seat at rung two, where they start deciding what counts as selected data the next time around. If you’ve concluded that a particular subcontractor’s models are sloppy, you’ll notice their clashes faster than their clean work. If you’ve concluded a chatbot layered on your model is pretty reliable, you’ll notice the answers that confirm that and skim past the one that wasn’t. The belief doesn’t just survive the next round of inference: it curates it.
Argyris called this the reflexive loop, and it’s the mechanism behind something every coordinator has felt: the way certainty about a model tends to increase with use, independent of whether the model has become any more trustworthy. Familiarity gets mistaken for accuracy. Similarly, the tenth time you ask a model a question and get a fluent answer, you trust it more than the first time, not because you’ve done anything to verify the ninth answer, but because nothing has yet gone visibly wrong, and the absence of visible failure is promoted to evidence of reliability.
This is where we start to bring together the argument from the previous piece and Argyris’s ladder: an AI layer sitting on top of an information model doesn’t just risk giving you one bad answer with excessive confidence. Used daily, it risks building a reflexive loop of trust, each fluent answer reinforcing the belief that the layer and the model deserve the confidence you’re placing in it. That belief, in turn, shapes which discrepancies you’ll notice going forward.
4. An example
Let’s put the ladder on the floor of an actual meeting, so that’ll stop being an abstraction.
Someone on the structural side runs a clash detection pass on a Tuesday morning, before the architectural team has pushed their latest sync: the report comes back clean for a particular zone (let’s say level 3: gridlines D through H) and that’s the pool of observable data narrowing into selected data: it might not reflect the state of a whole level, but it’s just the state of a particular level when it comes to architectural and structural design clashed together in a model that’s three days out of date, filtered through a clash tolerance someone set up last January, still foggy from New Year’s champagne.
Nobody says any of that out loud. What gets said is that Level 3 is clean. That’s rung three already: a meaning has been attached to a data point, and the meaning is doing more work than the data can support. Someone else in the room, who has a deadline and enough on their plate, hears “clean” and assumes the coordination is actually current, not three days stale. That’s rung four (assumptions). By the time the conversation moves on, the conclusion has become that level 3 is coordinated. Two meetings later, it’s a belief: level 3 is one of the good ones, the team doesn’t worry about it anymore, attention drifts elsewhere. The action at the top of the ladder is an absence of action, a zone nobody double-checks again, right up until the site team finds a duct where a beam was meant to go.
Nobody lied. Nobody was incompetent, technically speaking. Every single rung was a reasonable, ordinary step to take under time pressure, which is exactly why the ladder is dangerous rather than merely careless: you can climb it doing everything right by the standards of an ordinary Tuesday, and still end up somewhere wrong. The failure isn’t at any one rung. It’s in nobody asking, at any point, which rung they were actually standing on when they made their statements. Do you follow me, so far?
5. Artificial Intelligence magnifies shit (as usual)
Now put an AI chat layer on top of that same model, and ask it the same question: is level 3 coordinated?
It will answer, it will answer fast, and it will answer in the same confident, complete-sentence register whether the underlying clash check is three hours old or three weeks old, whether the tolerance settings are sensible or left over from January’s hangover, whether the model it’s reading actually reflects the latest architectural sync or not. This is the bit that should give us pause: the AI doesn’t climb the ladder any differently than the tired coordinator on a Tuesday morning did. It climbs it faster, sure, and it climbs it without any of the hesitation a human occasionally shows, albeit grudgingly.
A person saying “level 3 is clean” might, on a good day, add a caveat. Rungs two and three leak into the sentence, a faint trace of the inference that produced it. A chatbot has no particular incentive to leak doubts into its answers: on the contrary, it’s optimised to sound complete and professional, never to sound provisional, and “level 3 is coordinated” is a much better sentence, by the metrics a language model is trained on, than “level 3 shows no clashes in a model that may not reflect this week’s changes, filtered through a tolerance nobody has revisited since January.” The first sentence sounds like an answer. The second sounds like a coordinator doing their job. Guess which one ships faster.
Last month’s article said that a fluent answer borrows its authority from the model underneath it. The ladder says the fluency itself is a symptom, it’s rung six dressed up as rung one, a belief wearing the clothes of an observation. Artificial Intelligence hasn’t skipped the ladder: it just mimicked the way we climb it, minus the part where a decent night’s sleep or a nagging doubt occasionally makes us pause halfway up.
6. Advocacy and Inquiry
If the ladder gets climbed automatically, coming down has to be deliberate. Argyris and Senge offer a fairly simple pair of tools for that: advocacy and inquiry.
Advocacy means saying out loud which rung you’re actually standing on, instead of presenting the top of the ladder as though it were the bottom. The clash report from this morning came back clean, on a model that’s three days old, with January’s tolerance settings: assuming that still holds, we can say Level 3 is clear in that portion. It’s a longer sentence, and a less satisfying one to say in a meeting where everyone wants to move to the next item, but it hands the room a place to disagree with you that isn’t disagreeing with your character or competence. Nobody has to accuse you of being wrong. They can simply ask about the sync, or the tolerance, and the conversation stays on the ladder instead of turning into a standoff between two beliefs.
Inquiry is the other half, and it’s harder because it means asking someone else where their certainty came from without it sounding like an accusation. Not “are you sure?” That just invites a firmer version of the same top-rung belief. Instead, inquiry is something closer to going deeper into what might have sprung that assumption, asking questions that send the other person, gently, back down their own ladder, and it’s astonishing how often the honest answer turns out to be two or three rungs lower than the sentence they’d just delivered with total confidence.
Applied to a Large Language Model, the same discipline holds, only the inquiry has to be pointed at the system: what data was this answer based on? Some tools will actually tell you, if asked, and most people never ask, because the interface is designed to make you feel like asking would be a waste of time, as though you’d interrupted someone competent mid-sentence to check their references. Advocacy and inquiry are, in the end, the practical answer to last time’s closing line: a way of staying in the business of questioning the model before questioning it.
7. Training to Notice the Rungs
Just like on a construction site, you can’t hand someone a faulty ladder and tell them to be careful. Being careful isn’t a skill: it’s a wish. What you can actually train is the much narrower habit of noticing which rung a sentence is standing on — theirs, a colleague’s, or an AI’s — and that turns out to be teachable in a fairly mundane way, closer to teaching someone to read a structural drawing than to teaching them philosophy.
Start with the sentence itself. People, understandably, learn to speak the language of conclusions early, because that’s the language that gets rewarded in meetings: “it’s clean,” “it’s resolved,” “it’s fine.” Nobody claps for a two-minute sentence in which you explain you selected a particular slice of data and what you inferred with moderate confidence. So the first exercise is just translation practice: take a confident sentence someone said in the meeting, and walk it back down. What was actually observed, what was selected out of everything else, what meaning got bolted onto it, and where, exactly, does the sentence stop being data and start being a belief. Do this often enough, on real examples from real coordination meetings, and people start doing it unprompted, mid-sentence, to their own statements.
The second habit is teaching people to be suspicious of a particular flavour of fluency: the sentence that resolves too cleanly for how messy the source usually is. A model that’s been assembled by a dozen hands under deadline pressure rarely produces clean, unambiguous answers: when one arrives, that’s not a reason to relax but a reason to ask what got smoothed over to make it sound that way. This is, not coincidentally, exactly the reflex an AI layer needs a human to bring, because it’s the one reflex the AI has no incentive to develop on its own.
And the third habit, the one that actually sticks, is normalising the inquiry question as an ordinary part of coordination, instead of an act of suspicion. “What did you actually see?” should sound as unremarkable in a BIM meeting as “can you share your screen?” A small procedural courtesy, not a challenge to anyone’s competence. Teams that manage this stop treating doubt as distrust and start treating it as a regular act of maintenance: the same category as backing up a model or checking a sync log, just aimed at the fragile place between a data point and a decision.
Happy doubting.
















No Comments