This is Part 3 of The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.
Nobody gives instructions all at once. You ask an assistant to draft an email, then remember it should sound less formal, then mention that the recipient is new, then that the deadline moved to Thursday. Each message feels like progress. Conversation is built from installments, and between people installments are cheap, because a person who hears a correction adjusts and carries on.
In 2025, Philippe Laban and Jennifer Neville of Microsoft Research, working with Hiroaki Hayashi and Yingbo Zhou of Salesforce Research, asked whether the habit is equally cheap with a language model. Their design was a clean one. They took tasks with complete, unambiguous instructions and delivered each task in two ways. In the first, the entire instruction arrived in a single message. In the second, the same instruction was split into shards, and at most one shard was revealed per turn, so the conversation built up to exactly the same request, one piece at a time. Same information. Same final ask. Only the packaging differed.
They ran this across fifteen models and six kinds of generation task. Averaged over all of them, performance in the sharded conversations was 39 percent lower than performance on the single message.
That result went on to win an Outstanding Paper award at ICLR 2026, and the authors compressed it into one sentence: when language models take a wrong turn in a conversation, they get lost and do not recover.
The installment habit, in other words, is not free. This piece asks what it costs, why the cost survives later correction, and why one of the oldest experiments in cognitive psychology found a cousin of it in people.
The damage is mostly in the variance
The first thing worth knowing about the 39 percent is where it comes from. The authors split the loss into two parts. Aptitude is how well a model does when its run goes well. Unreliability is how widely its results swing from one run to the next on the same task. Aptitude fell by about 16 percent in the sharded setting. Unreliability rose by about 112 percent, more than doubling.
So the models did not lose the ability to do the work. What changed was how consistently they did it: the same conversation, run again, could end well on one attempt and badly off course on another. The authors state the conclusion directly: the large degradation is due in large part to increased unreliability, rather than a loss in aptitude.
That matters for anyone who has watched an assistant fix a mistake and then quietly repeat a version of it. The correction arrived. The model acknowledged it. Something from the earlier, wrong reading was still in the room.
The obvious repairs were tested, and they helped without finishing the job. The comparison used average scores on a 0 to 100 scale across four of the six tasks: code, database, math and actions. With the whole instruction delivered at once, GPT-4o scored 93.0 and GPT-4o-mini scored 86.8. Delivered in shards, they fell to 59.1 and 50.4. In the recap condition, a final turn restated every shard that had come before, which lifted GPT-4o to 76.6 and GPT-4o-mini to 66.5. In the snowball condition, every previously revealed shard was repeated at each new turn, which brought GPT-4o to 65.3 and GPT-4o-mini to 61.8. Both strategies narrowed the gap. Neither closed it: the best repaired score, 76.6, still sat well below the 93.0 of the single message.
Turning down the randomness did not rescue the situation either. Lowering the sampling temperature, the usual lever for making a model more consistent, was reported as ineffective at improving reliability in the multi-turn setting. For the two models tested, unreliability stayed near 30 points even with the temperature at its lowest.
One architectural fact deserves a mention, with a caution attached. In a transformer, attention weights come out of a softmax, and a softmax, in exact arithmetic, never returns an exact zero, so every earlier token in the window keeps some weight in each attention step. It is tempting to say that an early wrong reading therefore lingers, diluted but never deleted. The paper does not make that claim, and nothing in it tests it. It remains an architectural fact that sits comfortably beside the result, and no more than that.
The practical reading is a narrow one, and the evidence supports it only that far. When you already know what you want, saying all of it in the first message is more reliable than trusting later turns to repair an early misreading. The study measured instructions that were complete from the start. It says nothing about the other kind of conversation, the exploratory one, where you discover what you want by seeing what comes back. Nothing here argues against that kind of talk.
The oldest experiment on switching
People had been watching minds pay for changes of task long before anyone built a language model. In 1927, Arthur Jersild published a monograph titled Mental Set and Shift, in the series Archives of Psychology. Task-switching researchers still credit it as the first laboratory use of the paradigm. His central observation was that alternating between operations on the same items takes longer than staying with one operation through a block.
A natural objection is that the cost reflects surprise. Perhaps people are slow because they are caught off guard by the change, or because they must keep track of which task comes next. Give them enough warning and the cost should vanish.
In 1995, Robert Rogers and Stephen Monsell built the experiment that tested that explanation. Their participants classified pairs of characters. On some trials the job was to judge whether a digit was odd or even. On others it was to judge whether a letter was a consonant or a vowel. The tasks followed a strict alternating-runs pattern, two trials of one and then two of the other, round and round, so the sequence was perfectly predictable. The experimenters then varied how much time the participant had to get ready before each stimulus appeared.
With more preparation time, the cost of switching did shrink. But it did not disappear. As the authors report, even with 1.2 seconds available for preparation, a large asymptotic reaction time cost remained, and it appeared only on the first trial of the new task. The second trial in a run was fast. The first was slow on average, even with the longest warning tested.
Three features of that result are worth holding onto. The cost showed up in reaction time. It was paid once, on the first trial after the change, and then it was gone. And it persisted when everything that could plausibly help had been supplied: a predictable pattern, a long preparation window, and simple tasks with nothing hard about them.
Nobody agrees on why
If you ask what produces the leftover cost, the field gives more than one answer.
Rogers and Monsell attributed it to reconfiguration. In their account, the mind has to rebuild its task set, the bundle of rules for what to attend to and how to respond. Some of that rebuilding, they argued, is triggered only when the stimulus for the new task actually shows up. Advance preparation can do part of the work, and the stimulus has to do the rest. That would explain why a first trial stays slow no matter how early the warning arrived.
Ritske de Jong offered a competing account in a chapter published in 2000. His proposal was that people sometimes fail to complete their preparation at all, even when they have the time. On some trials the preparation happens and the cost is small. On others it never gets engaged, and the full cost appears. When Sander Nieuwenhuis and Monsell tested the idea in 2002, they found that strong incentives to prepare raised the estimated probability of preparing only marginally. That pointed to some limit on how reliably advance preparation can be achieved, and it left the larger question open.
Seventy-five years after Jersild, the field was still arguing about the mechanism behind its oldest finding. The phenomenon is solid. The explanation is unsettled, and anyone who tells you otherwise is describing a tidier science than the one that exists.
Two systems that paid for starting
Set the two results side by side and a shared situation appears. In each case the system had what it needed. The model eventually received every shard of the instruction, in some conditions repeated several times. The person had a predictable sequence and more than a second to prepare. In each case, a cost remained anyway, relative to the case where no switch or no installment was involved.
The mirror stops being exact in several places, and the places are instructive.
The human cost is brief. It is paid on one trial, then gone. The machine cost is a sustained loss in reliability across a whole conversation, and the authors’ phrase is that the model does not recover. What changes in each case also differs. The person switches between two well-defined tasks, each of which they can perform. The model, in a sharded conversation, changes its reading of what the single task is, and its early reading was formed on incomplete information. These are different events.
No source ties the two results to a common mechanism. The Laban paper cites no task-switching research, and the task-switching studies described here predate language models. This piece is therefore drawing an analogy. It is a conceptual analogy, held at the level of a shared situation, and it should not be read as evidence that a language model’s lost thread and a person’s slow first trial are the same process running on different hardware.
What the analogy can carry is narrower, and still useful. In the two systems examined here, a system with every fact available does not automatically behave like one that had the facts from the start. Having started down one path leaves a mark that later information does not fully erase. For the person, the mark is brief, and defended by decades of argument about its cause. For the model, it is large enough to cost nearly 40 percent of performance on average, and none of the fixes the authors tested closed the gap.
What the Brain Knew
A predictable switch is still a switch. In these experiments the brain did not make the first step after a change free, even with the pattern predictable and the preparation time to spare. Readiness is not the same as having already arrived, and the cost of a first move is paid in both kinds of mind examined here.
The Mirror Test is a Robonaissance series. Machines and brains explain each other, and every reflection redraws the map of intelligence.


