Robonaissance

Robonaissance

The Mirror Test, Part 8: The Donkey Kong Transistor

Neuroscience’s tools, tried on a chip whose workings were known, seemed to find a transistor for one video game. After seven mirrors, where does the comparison stop?

Hugo's avatar
Hugo
Oct 09, 2026
∙ Paid

This is Part 8 of The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.


Neuroscience’s usual tools, tried on a chip whose workings were completely known, seemed to find a transistor for one video game. On January 12, 2017, Eric Jonas and Konrad Kording published a paper in PLOS Computational Biology with a simple design. They would “take a classical microprocessor as a model organism,” and use their ability to perform any experiment they liked on it “to see if popular data analysis methods from neuroscience can elucidate the way it processes information.”

The organism was the MOS 6502, or its near twin the 6507, the processor inside the Apple I, the Commodore 64 and the Atari Video Game System. A team called Visual6502 had reverse-engineered it “by chemically removing the epoxy layer and imaging the silicon die with a light microscope,” and produced “a transistor-accurate netlist,” a map of its 3,510 enhancement-mode transistors that the authors liken to a brain’s full connectome. A simulator could then track “the voltage on every wire and the state of every transistor.”

For behaviour, the chip played games. “Here we will examine three different ‘behaviors’, that is, three different games: Donkey Kong (1981), Space Invaders (1978), and Pitfall (1981).”

Then the authors did what neuroscientists do with brains. They lesioned it, one transistor at a time, by forcing each one permanently on, and counted a game as lost if the chip could no longer “draw the first frame of the game.” “The elimination of 1565 transistors have no impact, and 1560 inhibit all behaviors.” Some knocked out only one game, 186 of them by Jonas and Kording’s count. “We can thus conclude they are uniquely necessary for the game—perhaps there is a Donkey Kong transistor or a Space Invaders transistor.”

Then the authors turn on their own result. “This finding of course is grossly misleading. The transistors are not specific to any one behavior or game but rather implement simple functions, like full adders.” A full adder is a circuit that adds binary numbers and accounts for values carried in as well as out.

Four bars: 1,565 transistors with no impact on any game, 1,560 whose loss stopped all three games, 200 that stopped two games and 186 that stopped one game.
Transistors forced on one at a time: 1,565 had no impact on any game, 1,560 stopped all three games, 200 stopped two and 186 stopped one. The authors called the idea of a game-specific transistor “grossly misleading.” Counts as the paper reports them, not a partition of the 3,510. Source: Jonas and Kording, PLOS Computational Biology, January 12, 2017.

The lesion study did not fail for lack of data. The authors could record every wire. “We find that many measures are surprisingly similar between the brain and the processor but that our results do not lead to a meaningful understanding of the processor.” What made the failure visible was something brains do not offer: “In the case of processors we know their function and we can know if our algorithms discover it.”

Part 1 borrowed the same word, if not the same experiment. It called the strawberry failure “a lesion study,” and said that what it localized was “the boundary of a perceptual system.” The method can be turned on the comparisons themselves. Of seven comparisons between machines and brains, which ones explained something, and where does the comparison stop?

Four questions

The chip gives a better instrument than the comparisons started with. For any brain–machine comparison, four things can be asked.

What class is it? A shared problem, a shared mechanism, a formal identity, or a documented influence.

What, exactly, is the shared problem? Stated without the shared word, as Part 6 put it, “what the system is relying on.”

What intervention could break it? A lesion with a known answer, a planted cause, an injected concept, a delay inserted without the subject noticing.

And on what date was the machine side true? Part 1’s sentence about what models lack was written about the models of its day.

Run the chip through them. Its claim was a mechanism, one transistor for one game. What broke it was a lesion checked against a known answer: the transistors turned out to implement simple functions, like full adders, and nothing specific to one game. Neither the function nor the date was in doubt, because the chip does not change and its function was known. A brain does not come with a known answer, and the machines keep changing.

What a stronger mirror looks like

There is a case where a brain–machine comparison reached further. In 1997 Wolfram Schultz, Peter Dayan and Read Montague argued that dopamine neurons in the primate brain have an output that “apparently signals changes or errors in the predictions of future salient and rewarding events,” matching the prediction errors of an algorithm called temporal-difference learning. The algorithm had a history of its own: it “was originally inspired by behavioral data on how animals actually learn predictions.” DeepMind’s 2017 review calls the whole exchange a “virtuous circle.” What made it strong was a signal measured in neurons that resembled the algorithm’s prediction errors.

Run the four questions on it. The class: a match between a measured neural signal and an algorithm’s prediction errors. The shared problem: learning to predict rewards. The intervention: the test was the recorded signals set against the algorithm’s predictions. The date: 1997, and DeepMind’s 2017 review still calls the exchange a “virtuous circle.”

Formal equivalence exists too, but only where someone defines both sides. Mehdi Keramati and Boris Gutkin, using “a definition of primary rewards, as outcomes fulfilling physiological needs,” wrote in 2014: “Within this framework, we mathematically prove that seeking rewards is equivalent to the fundamental objective of physiological stability.” The proof holds inside the framework that sets it up.

Even a network that predicts the brain well is not yet a mirror of it. In 2014 Daniel Yamins and colleagues found that a network optimised only for categorising objects, “even though we did not constrain this model to match neural data,” was “highly predictive of IT spiking responses,” the firing of neurons in IT, “the highest ventral cortical area.” Jeffrey Bowers and colleagues answered, in a preprint, that such predictions “may be mediated by DNNs that share little overlap with biological vision,” and that models should explain “experiments that manipulate independent variables.” That is Jonas and Kording’s point again: matching measurements is not the same as understanding.

User's avatar

Continue reading this post for free, courtesy of Hugo.

Or purchase a paid subscription.
© 2026 Robonaissance · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture