A fluent rationale may be useful without being a causal trace. Chain-of-thought can narrate a path the model did not take, and institutions that treat explanation as mechanism will audit the wrong object.
People ask systems why. The ask is reasonable. The answer is often a genre. After an event, humans produce reasons that organise the event for social life: blame, credit, teaching, reassurance, the minutes of a meeting. Those reasons can be indispensable and still fail as causal histories. Philosophy has known the distinction for a long time under various names — justification versus causation, rationalisation versus explanation, the story told versus the process that occurred. The novelty is not the distinction. The novelty is an industry that ships the story at machine speed and invites everyone to treat it as a trace.
Chain-of-thought prompting and "reasoning" model transcripts look like traces because they are sequential, verbal, and confident about intermediate steps. They inherit the prestige of proof without necessarily inheriting the binding of proof. The institute's claim is narrow and unkind to several product metaphors at once: the explanation is not the cause.
The claim is easy to overhear as cynicism about language. It is not. Language remains one of the best tools we have for teaching, coordinating, and catching mistakes. The error is category error: treating a generated pedagogical object as if it were instrumentation. Oscilloscopes do not owe us a paragraph. Paragraphs do not become oscilloscopes by numbering their sentences.
1 — Usefulness without faithfulness
An explanation can help a user understand a domain, check a calculation, or notice a missing assumption. None of that requires the explanation to be a faithful report of the internal factors that produced the answer. Faithfulness is a separate property: whether the stated reasons actually tracked the determinants of the output. Research on chain-of-thought faithfulness — including work showing that reasoning models do not always say what they think, in Anthropic's phrasing — finds that verbalised intermediate steps can omit, distort, or retrofit factors that influenced the final response.
Source ·Reasoning Models Don't Always Say What They Think (opens in new tab)· Anthropicatlas
This is not an argument that models are uniquely deceitful. Humans are fluent at the same genre. It is an argument that fluency is a poor proxy for mechanism, and that scaling fluency scales the proxy problem. The comic version is a system that cites a paper it did not use, for a reason it did not have, in a tone that would pass a midterm. The serious version is an audit trail that documents a performance of care.
2 — Why the confusion is institutional
Organisations want handles. If a model can emit a paragraph beginning with "Because," the paragraph becomes a handle for compliance, customer support, and internal blame allocation. The handle is attractive precisely because it is legible in the same medium as policy. Weights are not. Activations are not. Training data influence is not, except in carefully constructed studies. So the narrative is promoted from interface copy to epistemic object.
The promotion has a cost. Teams begin to optimise the story. Safety reviews begin to read the story. Incident reports begin to quote the story. None of these activities are foolish in isolation. Together they can build an organisation that is extremely good at supervising a genre while remaining comparatively blind to the process the genre claims to disclose.
Post-hoc explanation literature in machine learning already warned that feature attributions and rationales can be disconnected from decision mechanisms. Chain-of-thought inherits the warning with better prose. Better prose is not a repair. It is an amplifier.
There is also a training-time version of the confusion. If models are rewarded for producing intermediate text that looks like careful reasoning, they will become excellent at the look. Performance on tasks may rise for ordinary reasons — decomposition helps humans and machines alike — while the faithfulness of the verbal channel remains an empirical question, not a gift conferred by the UI label "thinking." The institute's allergy is specifically to the gift theory of labels.
3 — What we are not claiming
We are not claiming that all chain-of-thought is theatre. Some intermediate text tracks recoverable computation; some scaffolding improves task performance; some transcripts are valuable as objects of study precisely because they can be compared against interventions that test faithfulness. We are not claiming that silence would be more honest. Silence is often just silence. We are claiming that the industry's preferred inference — from readable steps to causal account — is invalid as a default.
Nor are we romanticising opacity. Opacity is not depth. Unfaithful explanation is not "negative capability." It is a reporting failure that happens to be eloquent. The institute's benches that interrupt premature coherence are aimed at a different problem — forced resolution — and must not be misread as praise for confabulation. Confabulation is cheap. Judgement about when a rationale may be trusted as mechanism remains expensive.
Another reasonable reply is instrumental: if users behave better when given any structured rationale, ship the rationale. Perhaps. Then say what you shipped. "Behavioural scaffold" and "causal trace" are different products. Marketing them under one word — explanation — is how institutions acquire elegant failure modes.
4 — Design implications without a dashboard religion
If explanation is not automatically cause, several design habits need demotion.
Do not treat a reasoning transcript as a sufficient audit artifact. Pair it with interventions that can falsify faithfulness claims: counterfactual prompts, blinded evaluations, tests for whether stated factors actually move outputs when altered. Do not punish models only for ugly intermediate steps while rewarding beautiful ones; that selects for style under the banner of safety. Do not let "the model said it relied on X" settle a dispute about whether X was relied on. Settle that with evidence that can survive the model's next paragraph.
For human operators, the discipline is older than the tools. Ask what would change your mind about the answer if the rationale were deleted. If the answer is "nothing, I only needed the conclusion," you were never reading a cause. If the answer is "I would lose the only handle I have," you have an operational problem that a prettier chain-of-thought will not solve. You have a missing instrument.
A second operator habit: when the rationale and the answer disagree under light pressure — when altering a stated factor does not move the output, or when the model cheerfully revises the story while keeping the conclusion — treat that as evidence about the genre, not as a personality flaw in the machine. Machines do not have the kind of sincerity the disappointment presupposes. They have channels. Some channels report. Some channels perform. Product copy that collapses the two is doing theology with a shipping date.
“Please clarify whether the authors want less explanation or better theatre. The manuscript currently requests both and invoices neither.”
We want less confusion between genres. Keep explanations where they teach, scaffold, or propose. Demand separate warrants where you need mechanism. The explanation may still be the most readable object in the room. Readability is not causality. The institute will continue to publish notes that sound sure about that distinction, and to keep the discard pile of times we mistook a good story for a good account — not as confession, but as method.