This is the overture to The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.
In 1970 the psychologist Gordon Gallup put a full-length mirror outside the cages of individually housed chimpanzees and left it there for ten days. At first each chimp treated the reflection as another chimp. After a few days the social behavior faded, and the animals began using the glass for themselves, grooming parts of their bodies that they could see only in the mirror.
Then Gallup anesthetized each one and painted a red mark above one eyebrow and on the top of the opposite ear, with a dye that had no smell and no feel. When the chimps woke, with the mirror gone, they rarely touched the marks. When the mirror came back, they touched them again and again, looked at their fingers, and some sniffed them. Macaques given the same treatment never used the mirror to investigate the mark, even after three weeks of looking.
What makes this a test is the mark. It cannot be seen or felt without the glass, so a chimp that reaches for its own brow has used an image to learn something about itself that nothing else could have told it. The evidence is not the mirror. It is what the animal does with it.
Machines have stood in front of mirrors too. In The Living Brain, published in 1953, W. Grey Walter described what his electronic tortoise did when it met its own reflection. The robot carried a small lamp that switched off whenever its photocell received enough light. Facing a mirror, it saw its own lamp, steered toward it, switched the lamp off, lost the signal, and switched it on again. Walter wrote of “flickering, twittering, and jigging like a clumsy Narcissus,” and he read the dance as self-recognition. The circuit he describes is enough to produce the dance by itself. The tortoise had no mark to find.
That gap, between a behavior that resembles something and a mechanism that is that thing, is where this series lives. Put a machine failure beside a finding about brains and ask the question Gallup’s mark was built to ask: did looking at one show us something about the other that we could not have seen alone? Sometimes it did. Sometimes the traffic between the two fields can be followed in print, with ideas moving both ways. And sometimes the likeness is a word, memory or attention or prediction, doing the work a mechanism should do.
That last case is the one the series keeps returning to, because it is the easiest to believe. A matching curve feels like evidence. A shared word feels like evidence. Across eight Parts, the pairings that held up were narrower than their resemblances promised, and the most useful ones said where they stopped.
The Eight Parts
Eight Parts, in order, on a single The Mirror Test shelf.
Part 1: The Strawberry Problem: Ask a model how many r’s are in strawberry, then ask what it has actually seen.
Part 2: The Middle Doesn’t Pay: Twenty passages, one of them holding the answer. Move it, and watch what the score does.
Part 3: The First Move Costs More: A task delivered in installments, and a 1927 experiment on switching.
Part 4: The Honest Liar: Two lawyers, six court decisions, and a word clinicians use for sincere false accounts.
Part 5: Five Plus One Is Seven: A network that knew its ones was taught its twos. Then someone asked it a question it had already answered.
Part 6: The Hospital in the X-Ray: Chest X-rays carry their hospital with them. So does the model that reads them.
Part 7: Living in the Past: A robot learned to walk in a world where commands arrive instantly. Then it met the real one.
Part 8: The Donkey Kong Transistor: Neuroscience’s tools, a chip whose workings are fully known, and a transistor for one video game.
If This Is You
If you build with language models and have watched one fail in a way that made no sense, start with Part 1, where a model that can explain quantum field theory miscounts the letters in a fruit, or Part 4, where a chatbot invented court decisions and two lawyers filed them. Both begin with a failure you can reproduce, and both end somewhere more useful than a verdict of stupidity.
If you come from neuroscience or psychology and wonder what the machine side adds, try Part 2, where a model’s accuracy draws a curve that recall experiments drew for people in 1962, or Part 3, where a study of conversations with language models meets a task-switching experiment from 1927. Each Part rebuilds the brain side on its own terms before any comparison is made, so the neuroscience stays yours.
If you work on robots or on learning systems that have to live inside their own loops, Part 5, Part 6, and Part 7 are written for you: forgetting, shortcuts, and delay, each with a history that runs through both fields.
If you flinch at every sentence that begins with the brain works like, you are the reader this series was built for. Begin with Part 8, which turns neuroscience’s methods on a chip we fully understand, then read back through the first seven and ask which of them would have passed.
And if you simply like watching a careful experiment run, any Part will do. Each stands alone, and each ends with a short section called What the Brain Knew, which keeps only what survived the comparison. Whichever way in you take, the question is the one Gallup’s mark was built to ask: whether looking at one thing shows you something about the other that you could not have seen alone.
The Mirror Test is a Robonaissance series. Machines and brains explain each other, and every reflection redraws the map of intelligence.


