This is Part 6 of The Mirror Test, a Robonaissance series in which AI and neuroscience mirror each other.
Chest X-rays carry their hospital with them, and in one study the hospital was a clue to who had pneumonia.
In 2018 John Zech and five colleagues published a study in PLOS Medicine about a deep learning model that read chest X-rays for signs of pneumonia. The setting was a simulated screening task, built from records that already existed. The model was trained and tested on images from Mount Sinai Hospital and the National Institutes of Health Clinical Center together, and then tested at a third institution, Indiana University’s patient care network.
On the first two, the best model reached an area under the curve, the study’s measure of how well it ranked X-rays labelled positive for pneumonia above those labelled negative, of 0.931. At Indiana, the same model fell to 0.815.
The authors went looking for the reason, and found that the images carried their origin with them. In a separate test, they trained networks to predict the hospital instead of the disease. Those networks could tell which hospital system a radiograph came from for 99.95 percent of NIH images and 99.98 percent of Mount Sinai images. That would be a curiosity if hospitals saw the same patients. They did not. Pneumonia prevalence was 34.2 percent at Mount Sinai and 1.2 and 1.0 percent at the other two. Sorting the images by hospital alone, without looking at a lung, reached an area under the curve of 0.861 on the combined Mount Sinai and NIH data.
Inside Mount Sinai the signal was finer still. A network trained to tell the hospital’s departments apart sorted 5,805 of 5,805 inpatient portable radiographs and 449 of 449 emergency-department ones correctly. The cause turned up later. The inpatient units used scanners from Konica Minolta and the emergency department used Fujifilm, and the emergency images were stored “in an inverted color scheme (i.e., air appears white),” along with distinctive text marking the side of the body and the portable scanner. “While these identifying features were prominent to the model,” the authors wrote, “they only became apparent to us after manual image review.”
There was also a metal token that radiology technicians place on the patient, which shows up in a corner of the image. The hospital-detecting network had learned to find it. It did not need it, though. “CNNs did not require this indicator,” the authors found: “most image subregions contained features indicative of a radiograph’s origin.”

Nothing in the pneumonia model’s training said “find the hospital.” But the hospital predicted pneumonia, and the authors drew the conclusion: “When these strong features are correlated with disease prevalence, models can leverage them to indirectly predict disease.” The test that would have caught it, on data from somewhere else, is the one that a held-out set from the same sources cannot run.
That is the question of this part of the series. Every perceptual system, of silicon or of tissue, takes in less than it needs and relies on something to fill the gap. What does it rely on, and how does that choice decide where it fails?
Any feature that predicts
Start with what a classifier is asked to do. It sees an array of pixel values and a label, again and again, and adjusts itself to produce the label. Nothing in that procedure says which properties of the image are supposed to matter.
Andrew Ilyas and colleagues put the consequence formally in 2019. When a network minimises its classification loss, they wrote, “no distinction exists between robust and non-robust features.” A useful feature is anything that predicts the label. A hospital’s scanner qualifies.
Robert Geirhos and colleagues later pinned the result down. Shortcuts, they wrote, are decision rules that perform well on test data like the training data but fail on data from somewhere else, “revealing a mismatch between intended and learned solution.” The X-ray model passed the first kind of test and failed the second.
The same reliance elsewhere
Other failures have the same shape, and each points at a different thing a network relied on.
A cat can be an elephant. In 2019 Geirhos and colleagues used style transfer to make images whose shape said one thing and whose texture said another: a cat, for instance, with the skin of an elephant. A standard ResNet-50 called that image “Indian elephant” at 63.9 percent. “A cat with an elephant texture is an elephant to CNNs,” they wrote, “and still a cat to humans.” Across the cue-conflict images, human observers chose the shape category in 95.9 percent of their decisions that matched either cue; ResNet-50 chose shape in 22.1 percent. Then the authors replaced the texture of every ImageNet image with the style of a randomly chosen painting and trained the same ResNet-50 on the result. It learned “a shape-based representation instead.” The texture bias, they concluded, “is not by design but induced by ImageNet training data.”

A panda can be a gibbon. In a paper revised for publication in 2014, Christian Szegedy and colleagues reported that applying “an imperceptible non-random perturbation to a test image” could “arbitrarily change the network’s prediction.” They called the altered inputs adversarial examples. The next year Ian Goodfellow, Jonathon Shlens and Szegedy published the panda: a photograph that GoogLeNet labelled “panda” with 57.7 percent confidence became, after an added pattern, a “gibbon” at 99.3 percent. Why this happens is still argued. Ilyas’s group proposed that adversarial examples exploit “non-robust features,” patterns in the data that are “highly predictive, yet brittle and (thus) incomprehensible to humans,” while noting that other causes are not ruled out. On their account the networks are not malfunctioning; they are using real information that people cannot see.
An apple can be an iPod. In 2021 Gabriel Goh and colleagues asked whether handwriting could fool CLIP, an image model that learns from text as well as pictures. They “took several common items and deliberately mislabeled them”: an apple with a handwritten note saying “iPod” was classified as an iPod. Such attacks, they wrote, “can be executed non-programmatically and as black-box attacks, available to any adversary – including six year olds.” A model that learned from text had learned that the words in a picture predict what the picture is called.



