Robonaissance

Robonaissance

The Mirror Test, Part 5: Five Plus One Is Seven

In 1989 a network that learned to add twos forgot how to add ones. Brains face the same problem, and sleep is one place researchers have watched memories replay.

Hugo's avatar
Hugo
Oct 05, 2026
∙ Paid

This is Part 5 of The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.


In 1989 a small neural network that had learned to add was asked what 5 + 1 is, and it said 7.

It had known the answer. Michael McCloskey and Neal J. Cohen, in a chapter published that year, had trained it on 17 “ones” facts, 1 + 1 through 9 + 1 and 1 + 2 through 1 + 9, until it answered all 17 correctly by their strictest standard. Training ran in learning trials: on each trial the network saw every fact in the training set once, in random order, and nudged its weights after each one.

Then they taught it the twos: 2 + 1 through 2 + 9, and 1 + 2 through 9 + 2. The new training set held only twos facts. After a single learning trial, the network’s average squared error on the ones, a measure of how far its answers were from the right ones, rose “by more than an order of magnitude,” from .0015 to .0453. McCloskey and Cohen also used a deliberately lenient test, counting an answer as right if the network’s output looked more like the correct answer than like any other. By that test, across the 15 ones facts that did not also belong to the twos set, averaged over two runs, accuracy fell from 100 percent to 57 percent after one trial on the twos, and to 30 percent after two.

The wrong answers had a pattern. “After training on the twos facts the network responds to the vast majority of the ones problems as if they were twos problems,” the authors wrote, “for example, 5 + 1 is 7, and 6 + 1 is 8.” The network had not gone blank. Seven and eight are the right answers to 5 + 2 and 6 + 2, twos facts it had just learned.

One detail is stranger still. The problems 2 + 1 and 1 + 2 belonged to both sets, so their training never stopped. It seemed, in the authors’ words, “patently obvious” that they would be safe. They were not. After one learning trial on the twos, both were wrong even by the lenient test.

McCloskey and Cohen allowed themselves one joke about what they had found: “To put it somewhat flippantly, the magnitude of the observed interference makes it seem more like retrograde amnesia than retroactive interference.”

People forget, but not like this

Retroactive interference is the ordinary human version: learning something new makes older material harder to recall. Psychologists measure it with paired lists. You learn one list of pairs, A–B, then a second list, A–C, that gives the same first items new partners. Say you learn table–blue, then table–green: the second list eats into your memory of the first.

McCloskey and Cohen ran that experiment on their network and set it beside the human data of Barnes and Underwood from 1959, as they reported them. The two sets of numbers come from different tasks and rates of learning, so they are not a matched race. The direction is what matters. The people still recalled 52 percent of the first list after 20 trials on the second. The network fell from 100 percent to 0 percent on the first list after three trials on the second, at a point when it had learned only about 20 percent of the second list. A person loses some of the old list, while the network loses all of it before it has gained much of the new one.

Two bar charts. Left: share of ones facts still correct, 100 percent before training on the twos, 57 percent after one pass through the twos, 30 percent after two. Right: share of the first list still recalled, 0 percent for the network after 3 passes on a second list, 52 percent for people after 20 passes.
After one pass through the twos, the arithmetic network still got 57 percent of the ones facts right by its most lenient test; after two passes, 30 percent. On a list-learning task in the same chapter, the network lost the whole first list while people kept about half. Source: McCloskey and Cohen (1989); human data from Barnes and Underwood (1959) as McCloskey and Cohen report them. The two panels use different tasks and are not a matched comparison of rates.

Robert French, reviewing the field in 1999, put the problem in its general form. Catastrophic interference, he wrote, is “a radical manifestation of a more general problem,” the stability-plasticity problem: “how to design a system that is simultaneously sensitive to, but not radically disrupted by, new input.”

A learning system has to change, and it has to stay itself.

One set of weights

Why should learning the twos disturb the ones at all? Start with what the network has to work with. Everything it knows is stored in one set of connection weights, the numbers that set how strongly each unit influences the next. Each training step adjusts those weights to reduce the error on whatever example is in front of it. Nothing in that step looks back. Suppose the examples arrive in a block, first the ones and only then the twos. Once the twos begin, every example the network sees is a twos fact, and each step asks only that it do better on the example in front of it. Nothing in the step asks it to keep the ones right, so the weights that also carried the ones are free to move wherever the twos pull them.

The obvious answer is to keep the facts apart, giving each memory its own weights so that new learning cannot touch the old. That answer has a price, and it is the reason these networks were interesting in the first place. Related facts share weights, which is how a network learns that what holds for robins probably holds for sparrows. French put it bluntly: the shared features that let these networks generalize “are the root cause of catastrophic forgetting.” James McClelland, Bruce McNaughton and Randall O’Reilly made the same point in 1995: “reducing overlap avoids catastrophic interference at the cost of a dramatic reduction in the exploitation of shared structure.”

So the overlap stays, and the question becomes how to train a system with overlapping knowledge without wrecking it.

The penguin test

Here is the problem and the most direct fix in their simplest case.

The block above suggests the most direct answer: mix the old facts in with the new, so that the network keeps meeting both while it learns. The 1995 paper tested that answer in a simulation about animals. A network that already knew a set of birds, fish, trees and flowers was taught a new fact: the penguin is a bird that can swim but cannot fly.

Taught that fact on its own, the network generalised in the wrong direction: “as the network learns that the penguin is a bird that can swim but not fly, it comes to treat all animals—and, to a lesser extent, all plants—as having these same characteristics.”

Taught the same fact mixed in with the things it already knew, the result was “very little interference”: the new fact barely disturbed what the network already knew. The mixing worked, but only gradually, through many presentations, and it required keeping the old material around to mix in.

That is the whole problem in miniature. New learning that does not fit what a network already knows can spread into places it should not go, and the repair that worked needed the old material to be kept around and mixed back in.

Three ways to protect old learning

Once the problem is stated this way, the repairs fall into three families, and you could probably name them before reading on. The penguin test also showed a price: the old material had to be kept around to mix in. The repairs differ in how they deal with that.

The first changes the schedule. Keep some old examples in the training mix, so that training keeps revisiting the past. When the old data are gone, generate stand-ins. In 2017 Shin and colleagues trained a generative model to produce imitations of earlier tasks’ data and interleaved those samples with the new task, a method they called deep generative replay.

The second changes the plasticity, meaning how easily each weight can change. In 2017 Kirkpatrick and colleagues at DeepMind introduced elastic weight consolidation, which “remembers old tasks by selectively slowing down learning on the weights important for those tasks.” The authors described the penalty as something that “can therefore be imagined as a spring anchoring the parameters to the previous solution,” stiffer for the weights that mattered most to the old task. They tested it on ten Atari games learned one after another. With plain gradient descent, “the agent never learns to play more than one game.” With the spring, the agents “do indeed learn to play multiple games,” though they still fell short of what ten separate networks achieved.

The third changes the storage: keep some memories in a separate place where new learning cannot overwrite them.

Naming three repairs does not say which one works. One 2020 study put two of them to the same test: replay, from the first family, and weight protection, from the second. Gido van de Ven and colleagues had networks learn handwritten digits two at a time, in five rounds, and in the hardest version the network had to tell all ten digits apart at the end. Among the methods they compared, the two that protect weights, elastic weight consolidation and a relative called synaptic intelligence, “dramatically failed and only GR,” generative replay, “was able to successfully learn all digits.” In everyday terms, rehearsing the old material held up where protecting the weights did not.

If you had to guess what a brain might do, you would now have three candidates.

User's avatar

Continue reading this post for free, courtesy of Hugo.

Or purchase a paid subscription.
© 2026 Robonaissance · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture